Quality improvement method

Improving consistency in AI-supported assessment

Quality Improvement Methods

Provides a proportionate method for addressing AI-supported assessment, with clear responsibility, measurable outcomes and follow-up of residual risk.

Current consideration of consistency in AI-supported assessment is informed by the rapid adoption of generative AI tools, with consequences for governance, evidence and the treatment of affected learners. For the corrective programme, the purpose of an improvement method is not to produce an action plan; it is to change a material condition and verify that the change is sustained. The scope should include every materially affected setting, with differences in location, programme, delivery mode and learner population kept visible. Central policy alone does not establish consistent operation across the declared scope.

The historical reference basis is the rapid adoption of generative AI tools. Its relevance to the improvement priority should be assessed against the affected jurisdiction, learner population and form of provision. The international development warrants attention, but a consequential conclusion still requires current, attributable and representative evidence for the affected scope.

The system and institutional dimensions of the corrective programme should be considered together. In reviewing the intervention, assessment should provide valid and sufficiently consistent evidence that the stated learning outcomes have been achieved by the learner receiving the result. Authorities and providers hold different responsibilities, both of which must be discharged for the arrangement to operate reliably. The allocation of responsibility should prevent gaps between system oversight and institutional operation.

Purpose and present context

The analysis of consistency in AI-supported assessment should make its decision rule explicit. A decision concerning the improvement priority should recognise that consistency does not require identical decisions regardless of context. It requires comparable matters to be treated on the same principles, with material differences explained by relevant evidence and recorded criteria. The method should prevent an unfavourable result from being dismissed through an unrecorded change in interpretation.

A proper review of the intervention should establish the intended outcome before selecting controls or indicators. A decision concerning the intervention should recognise that a complete improvement record should define the baseline, affected scope, causal hypothesis, responsible owner, resources, milestones and measures of effectiveness. A chosen approach should be justified against its context, with departures and review points under documented control.

Application in practice

A narrow control over consistency in AI-supported assessment may create false assurance. In the present context, results used beyond the evidence they support, inconsistent judgement between markers or locations and uncontrolled changes to assessment may produce acceptable aggregate reporting while individual learners remain exposed to material disadvantage. Testing should include exceptions and adverse cases, not only routine or successful operation.

Readily available material should not define the enquiry if it cannot answer the relevant decision question. For the intervention, the most relevant material is likely to include approval and change-control records, marking criteria and calibrated judgement, moderation and exception records, and appeal and correction records. Confidence is strengthened by corroboration, not by the volume of records drawn from the same underlying source.

A proportionate method is available for the intervention. Review of the affected practice should use common definitions and decision criteria, calibrate responsible staff, review outliers and compare outcomes across locations and groups. Where variation is justified, retain the reason and verify that it is applied without arbitrary disadvantage. Adverse cases and unresolved contradictions should be retained because they may reveal limitations concealed by an average result.

Testing implementation and effect

A decision to close improvement work on consistency in AI-supported assessment should be made by a person with authority and sufficient independence from implementation. The closure evidence should cover the relevant period and scope, include adverse cases and show whether the change is sustained. Recurrence or unequal effect should trigger renewed analysis rather than automatic repetition of the same intervention.

Interpretation of the improvement priority should avoid two errors: treating a formal commitment as proof of effect, and treating one adverse case as proof that every part of the system has failed. A decision concerning the intervention should recognise that reliability without validity produces consistent but potentially irrelevant results. Validity without adequate consistency may expose learners to unequal judgement. The analysis of the intervention proceeds on the basis that methods should be proportionate to the significance and recurrence of the problem; low-risk local issues and systemic learner-protection failures require different levels of control.

Decisions concerning the matter under review should remain traceable to the information available for the stated reference period. Any revised finding should identify precisely what has changed and why the earlier conclusion no longer applies. Users should not be left to infer a change in performance where the observed movement results from revised reporting.

For the corrective programme, governing bodies should receive a concise account of the intended result, affected scope, principal risks, evidence limitations and unresolved exceptions. Material action requires a named responsible function and a defined completion point. Closure requires evidence that the condition has changed; completion of planned activity is not sufficient.

The measure of progress on the intervention is not the amount of policy or documentation produced. A credible measure shows whether the intended result is present across the affected scope and what action follows when it is not.