Quality improvement method

Risk-based improvement planning for AI-supported assessment

Quality Improvement Methods

Examines how improvement in AI-supported assessment should be designed, implemented and tested against the intended educational outcome.

Against the background of the rapid adoption of generative AI tools, education authorities and providers should review how AI-supported assessment is defined, implemented and evidenced. The analysis of the intervention proceeds on the basis that improvement should begin with a defined problem, a credible account of its causes and a measure capable of showing whether the response has worked. Proportionality should be assessed against effects on access, learning, fair treatment and the accuracy of learner information.

The intended substantive result should remain the starting point for review. A decision concerning the intervention should recognise that assessment should provide valid and sufficiently consistent evidence that the stated learning outcomes have been achieved by the learner receiving the result. The existence of an approved measure or completed activity is not evidence of educational effect. Authorities and providers require evidence of operation and effect, with a route to identify and correct unequal or unintended consequences.

The present position

Assurance of AI-supported assessment should draw on more than one form of evidence. Useful records include authorship and identity controls proportionate to risk, analysis of results and differential outcomes, appeal and correction records, marking criteria and calibrated judgement, and approval and change-control records. Documentary conformity alone is insufficient where operation or learner experience indicates a material difference. System-wide assurance cannot be inferred from a favourable case chosen after the event.

The historical reference basis is the rapid adoption of generative AI tools. Its relevance to the affected practice should be assessed against the affected jurisdiction, learner population and form of provision. International developments provide context; decisions affecting learners require evidence that is current and representative of the setting concerned.

A focused examination of the affected practice requires a clear analytical discipline. Oversight of the affected practice should reflect the principle that materiality should be judged by the possible effect on learning, safety, rights, recognition, public resources and the reliability of a consequential decision. Frequency is relevant, but a rare event may still be material where the effect is serious or irreversible. The distinction matters because evidence may appear sufficient while addressing a different population, period or outcome.

Risk assessment of the intervention should give particular attention to tasks that do not assess the stated outcome, weak assurance of authorship or performance, and results used beyond the evidence they support. A provider should also consider uncontrolled changes to assessment and inconsistent judgement between markers or locations. The control response should reflect whether an affected learner can identify the error and obtain an effective remedy in time.

Operational significance

The governing expectation for AI-supported assessment should be capable of consistent application. The analysis of the improvement priority proceeds on the basis that effectiveness should be judged against an agreed outcome and reference period, not against completion of activities alone. Operational definitions should be precise enough to support consistent consequential decisions and explain justified variation.

Accountability for the improvement priority should follow decision-making authority. Relevant evidence should reach the body authorised to commit resources, amend policy or accept residual risk, and its judgement should be recorded. Delegation of delivery does not remove the need for a named authority to oversee material learner impact.

Decisions concerning the intervention should remain traceable to the information available for the stated reference period. Any revised finding should identify precisely what has changed and why the earlier conclusion no longer applies. Without this distinction, a reporting change may be mistaken for improvement or deterioration in educational practice.

  • Control changes and retain evidence sufficient for independent review.
  • Align tasks and criteria with learning outcomes, including material exceptions and unequal effects.
  • Moderate material variation, including material exceptions and unequal effects.
  • Review differential and anomalous results, recording who is responsible and which provision or learners are affected.
  • Retain evidence sufficient for review, including material exceptions and unequal effects.

What should be examined

The review method for AI-supported assessment should be reproducible. For the intervention, the reviewer should define escalation thresholds before reviewing cases, consider severity, reach, duration, recurrence and detectability, and record the reason for the final classification. Reassess materiality when new evidence changes the likely scope or consequence. Working papers should allow another competent reviewer to understand the evidence, judgement and treatment of material exceptions.

The improvement record for the intervention should contain the verified problem, affected scope, immediate containment, causal analysis, selected intervention, accountable owner, resources, milestones and effectiveness measure. The action record should separate administrative completion from verification of the intended change. The oversight record should preserve both outstanding action and the risk that continues during implementation.

The basis and limits of any conclusion concerning the corrective programme should be explicit. A decision concerning the corrective programme should recognise that reliability without validity produces consistent but potentially irrelevant results. Validity without adequate consistency may expose learners to unequal judgement. Oversight of the affected practice should reflect the principle that a short-term increase in activity may not represent sustained improvement. Measures should remain in place long enough to detect recurrence and unintended effects. Limitations should be prominent wherever the finding may influence a consequential decision.

Neither one indicator nor one control can establish the complete position on the affected practice. A conclusion should be revised when stronger evidence materially changes the assessment of implementation, outcome or risk.