Quality improvement method

AI-supported assessment: an evidence-led improvement method

Quality Improvement Methods

Considers the controls required to improve AI-supported assessment and to distinguish completed activity from demonstrated change.

Current consideration of AI-supported assessment is informed by the rapid adoption of generative AI tools, with consequences for governance, evidence and the treatment of affected learners. In reviewing The corrective programme, improvement should begin with a defined problem, a credible account of its causes and a measure capable of showing whether the response has worked. The unit of review should correspond to the full reach of the decision, including significant differences in provision and population. Evidence of formal policy should not be treated as evidence of uniform implementation.

Purpose and present context

The stated reference—the rapid adoption of generative AI tools—establishes the contemporaneous context. Any conclusion about AI-supported assessment still requires evidence from the setting concerned. The decision basis should identify what is evidenced, what reflects policy and what depends on authorised discretion. Later review should not obscure whether the earlier position rested on fact, policy or judgement.

The governing expectation for the affected practice should be capable of consistent application. In reviewing The improvement priority, follow-up should determine whether the change is embedded in ordinary operations and whether it has created new risks or unequal effects. Criteria affecting learners should not permit materially different interpretation without an evidenced reason.

The principal risks in relation to the corrective programme are reasonable adjustment altering the assessed outcome, tasks that do not assess the stated outcome, uncontrolled changes to assessment, and results used beyond the evidence they support. The risks are interdependent; failure of one control may conceal or disable another. The evidential trail should be examined from initial decision to outcome, including transfers of responsibility.

The analysis of the intervention should make its decision rule explicit. In reviewing The matter under review, the subject should be examined as a connected system of policy, people, resources, decisions and evidence. Gaps may emerge when authority, records or action pass between responsible bodies. A stated decision rule enables comparable examination and limits retrospective explanations of adverse evidence.

Evidence should be selected against a clearly defined question. For The corrective programme, the most relevant material is likely to include approval and change-control records, moderation and exception records, marking criteria and calibrated judgement, and assessment maps to learning outcomes. Independent records should be reconciled, with disagreement and uncertainty reported alongside the finding.

Responsibilities and material risks

The analysis of AI-supported assessment should remain within the limits of the evidence. Oversight of the matter under review should reflect the principle that a short-term increase in activity may not represent sustained improvement. Measures should remain in place long enough to detect recurrence and unintended effects. For The corrective programme, reliability without validity produces consistent but potentially irrelevant results. Validity without adequate consistency may expose learners to unequal judgement. Material uncertainty should result in further enquiry or an expressly limited finding.

The evidential trail should allow an affected decision to be identified, examined and corrected. For The matter under review, the responsible body should be able to identify the evidence considered, the judgement made, the person or body authorised to make it and the action that followed. Historical decisions should be assessed against the information then available, with later amendments separately dated and explained.

  • Align tasks and criteria with learning outcomes, including material exceptions and unequal effects.
  • Calibrate assessors and retain evidence sufficient for independent review.
  • Control changes, with responsibility, scope and timing recorded.
  • Retain evidence sufficient for review within a defined period and review the result.
  • Moderate material variation, recording who is responsible and which provision or learners are affected.

Information required for oversight

A proportionate method is available for AI-supported assessment. The method for the matter under review is to map the complete process, identify the intended result and responsible authority at each stage, and test normal cases together with exceptions. The review should determine whether correction of an individual case is sufficient or broader action is required. Contrary evidence should not be removed merely because aggregate performance appears acceptable.

A decision to close improvement work on the intervention should be made by a person with authority and sufficient independence from implementation. The closure evidence should cover the relevant period and scope, include adverse cases and show whether the change is sustained. Recurrence or unequal effect should trigger renewed analysis rather than automatic repetition of the same intervention.

  • Who controls each stage?
  • What outcome is intended?
  • Where do exceptions occur?
  • What action is required by the finding?
  • Which evidence establishes operation?

Proportionality and exceptions

Public reporting on AI-supported assessment should distinguish established fact, analytical judgement and planned action. Material revisions should be traceable to their reason and effective date. Changes to definitions or evidence should be recorded separately from changes in educational performance.

The central objective should not be obscured by the form of the administrative response. In reviewing The improvement priority, assessment should provide valid and sufficiently consistent evidence that the stated learning outcomes have been achieved by the learner receiving the result. Inputs and formal commitments should be distinguished from demonstrated operation and outcome. The operating record should enable responsible bodies to detect unintended effects and act where outcomes are unequal.

Authorities and providers should use the current development to test whether the improvement priority connects public commitment with effective operation and evidence of result. Improvement should be supported by evidence and an accountable decision record capable of public scrutiny.