Data and research analysis

AI-supported assessment: evidence, coverage and limitations

Data Research

Considers what the available data can establish about AI-supported assessment and identifies the limitations that should accompany any public conclusion.

In 2023, consideration of AI-supported assessment must take account of the rapid adoption of generative AI tools and the responsibilities it places before education systems. Oversight of the reported measure should reflect the principle that comparable indicators can support public decision-making, but they do not remove the need to examine variation within systems and institutions. The scope should include every materially affected setting, with differences in location, programme, delivery mode and learner population kept visible. A policy approved at the centre is insufficient where local implementation has not been tested.

The historical reference basis is the rapid adoption of generative AI tools. Its relevance to the evidence under review should be assessed against the affected jurisdiction, learner population and form of provision. International developments provide context; decisions affecting learners require evidence that is current and representative of the setting concerned.

Responsibility for the comparison should be visible at the point where consequential decisions are made. A decision concerning the reported measure should recognise that trend claims require comparable observations over time and a documented account of revisions, breaks in series and changes in coverage. A decision should not be closed at the operating level where material impact, conflict or a significant evidential gap remains unresolved.

The present position

The quality significance of AI-supported assessment follows from a basic distinction between availability and effective provision. A decision concerning the evidence under review should recognise that assessment should provide valid and sufficiently consistent evidence that the stated learning outcomes have been achieved by the learner receiving the result. A single entry control or reported outcome cannot demonstrate consistent operation across the learner journey.

In practical terms, The comparison should be reviewed against a stated method rather than general assurance. Oversight of the analytical question should reflect the principle that data quality comprises accuracy, completeness, timeliness, consistency and traceability. Strength in one dimension does not compensate automatically for weakness in another, particularly where the information informs a consequential learner decision. Decision-makers should receive an intelligible account of how the result was reached and where it should not be applied.

A narrow control over the evidence under review may create false assurance. In the present context, tasks that do not assess the stated outcome, uncontrolled changes to assessment and inconsistent judgement between markers or locations may produce acceptable aggregate reporting while individual learners remain exposed to material disadvantage. Adverse cases should form part of the sample wherever they may reveal a material control weakness.

The evidential record for the comparison should permit a reviewer to trace the matter from decision to outcome. This may require approval and change-control records, authorship and identity controls proportionate to risk, marking criteria and calibrated judgement, and appeal and correction records, supported by assessment maps to learning outcomes and analysis of results and differential outcomes. Sampling remains insufficient where it excludes a material group or cannot resolve contradictory evidence or recurrence.

Application in practice

The review method for AI-supported assessment should be reproducible. Review of the matter examined should trace selected records to source, reconcile totals across systems, quantify missing and late submissions, review manual adjustments and retain a revision history. Escalate discrepancies that could alter a published conclusion or individual outcome. The retained analysis should be reproducible from the selected evidence, decision rule and recorded reasons for accepted exceptions.

Publication of findings on the matter examined should distinguish observed values, estimates and interpretation. Revisions, breaks in series and changes in classification should be visible. Where disaggregation creates small or unstable groups, confidentiality and uncertainty should be managed without concealing a material disparity that requires further investigation.

Decisions concerning the analytical question should remain traceable to the information available for the stated reference period. Changes in condition, evidence, method and interpretation should be recorded separately when a conclusion is revised. Without this distinction, a reporting change may be mistaken for improvement or deterioration in educational practice.

Testing implementation and effect

The analysis of AI-supported assessment should remain within the limits of the evidence. In reviewing The matter examined, a single indicator rarely provides an adequate account of quality. Quantitative evidence should be considered with implementation records and the experience of affected learners. For The comparison, reliability without validity produces consistent but potentially irrelevant results. Validity without adequate consistency may expose learners to unequal judgement. If uncertainty could change a consequential decision, additional evidence or a narrower conclusion is required.

Where The comparison involves partners, suppliers or several public bodies, responsibility should be mapped across the complete service. Contractual or inter-agency arrangements should identify who holds records, informs learners and acts on incidents. Multiple delivery partners do not justify fragmented accountability or remedy.

Assurance concerning the evidence under review requires corroborating evidence across the material scope. Assurance should be based on the combined legal or policy basis, operating evidence and learner effect, not on one element alone.