质量改进方法

Risk-based improvement planning for AI-supported assessment

质量改进方法

Sets out risk-based improvement planning as an evidence-led approach to AI-supported assessment, covering responsibility, outcome evidence and sustained effect.

Against the background of the rapid adoption of generative AI tools, education authorities and providers should review how AI-supported assessment is defined, implemented and evidenced. Improvement of AI-supported assessment should begin with a defined problem, a credible account of its causes and a measure capable of showing whether the response has worked. Proportionality should be assessed against effects on access, learning, fair treatment and the accuracy of learner information.

For AI-supported assessment, the intended substantive result should remain the starting point for review. Assessment should provide valid and sufficiently consistent evidence that the stated learning outcomes have been achieved by the learner receiving the result. The existence of an approved measure or completed activity is not evidence of educational effect. Authorities and providers require evidence of operation and effect, with a route to identify and correct unequal or unintended consequences.

Improvement objective and baseline

Assurance of AI-supported assessment should draw on more than one form of evidence. Useful records include authorship and identity controls proportionate to risk, analysis of results and differential outcomes, appeal and correction records, marking criteria and calibrated judgement, and approval and change-control records. System-wide assurance cannot be inferred from a favourable case chosen after the event.

The relevant context is provided by rapid adoption of generative AI tools. Its relevance to the relevant practice should be assessed against the affected jurisdiction, learner population and form of provision.

In work concerning AI-supported assessment, materiality should be judged by the possible effect on learning, safety, rights, recognition, public resources and the reliability of a consequential decision.

Risk assessment of the corrective action should give particular attention to tasks that do not assess the stated outcome, weak assurance of authorship or performance, and results used beyond the evidence they support. A provider should also consider uncontrolled changes to assessment and inconsistent judgement between markers or locations. For AI-supported assessment, the control response should reflect whether an affected learner can identify the error and obtain an effective remedy in time.

Controls and accountable action

As regards AI-supported assessment, the applicable expectation should be capable of consistent application. Effectiveness should be judged against an agreed outcome and reference period, not against completion of activities alone. Within the scope under review, operational definitions should be precise enough to support consistent consequential decisions and explain justified variation.

Accountability for AI-supported assessment should follow decision-making authority. Relevant evidence should reach the body authorised to commit resources, amend policy or accept residual risk, and its judgement should be recorded. Delegation of delivery does not remove the need for a named authority to oversee material learner impact.

When examining AI-supported assessment, decisions concerning the corrective action should remain traceable to the information available for the stated reference period. Without this distinction, a reporting change may be mistaken for improvement or deterioration in educational practice.

  • Control changes.
  • Align tasks and criteria with learning outcomes.
  • Moderate material variation.
  • Review differential and anomalous results.
  • Retain evidence sufficient for review.

Evidence of effect

The review method for AI-supported assessment should be reproducible. For the corrective action, the reviewer should define escalation thresholds before reviewing cases, consider severity, reach, duration, recurrence and detectability, and record the reason for the final classification. Working papers should allow another competent reviewer to understand the evidence, judgement and treatment of material exceptions.

The improvement record for the corrective action should contain the verified problem, affected scope, immediate containment, causal analysis, selected intervention, accountable owner, resources, milestones and effectiveness measure. When examining AI-supported assessment, the action record should separate administrative completion from verification of the intended change. The oversight record should preserve both outstanding action and the risk that continues during implementation.

The basis and limits of any conclusion concerning corrective action should be explicit. For decisions concerning AI-supported assessment, reliability without validity produces consistent but potentially irrelevant results. Validity without adequate consistency may expose learners to unequal judgement. A short-term increase in activity may not represent sustained improvement. Measures should remain in place long enough to detect recurrence and unintended effects. Within the scope under review, limitations should be prominent wherever the finding may influence a consequential decision.

Neither one indicator nor one control can establish the complete position on the relevant practice. A conclusion concerning AI-supported assessment should be revised when stronger evidence materially changes the assessment of implementation, outcome or risk.