数据与研究分析

AI-supported assessment: evidence, coverage and limitations

数据研究

Examines AI-supported assessment, addressing evidence, coverage and limitations and the evidential limits relevant to responsible interpretation and decision-making.

In 2023, consideration of AI-supported assessment must take account of the rapid adoption of generative AI tools and the responsibilities it places before education systems. Comparable indicators can support public decision-making, but they do not remove the need to examine variation within systems and institutions.

Rapid adoption of generative AI tools provides the reference point for this analysis. Its relevance to the available evidence should be assessed against the affected jurisdiction, learner population and form of provision.

For AI-supported assessment, responsibility should be identifiable at the point where consequential decisions are made. Trend claims require comparable observations over time and a documented account of revisions, breaks in series and changes in coverage. A decision should not be closed at the operating level where material impact, conflict or a significant evidential gap remains unresolved.

Evidence and method

AI-supported assessment should provide valid and sufficiently consistent evidence that the stated learning outcomes have been achieved by the learner receiving the result.

When examining AI-supported assessment, review of the comparison should be based on a stated method rather than general assurance. Data quality comprises accuracy, completeness, timeliness, consistency and traceability. Strength in one dimension does not compensate automatically for weakness in another, particularly where the information informs a consequential learner decision. Decision-makers should receive an intelligible account of how the result was reached and where it should not be applied.

A narrow control over the available evidence may create false assurance. In the present context, tasks that do not assess the stated outcome, uncontrolled changes to assessment and inconsistent judgement between markers or locations may produce acceptable aggregate reporting while individual learners remain exposed to material disadvantage. In work concerning AI-supported assessment, adverse cases should form part of the sample wherever they may reveal a material control weakness.

The evidential record for the comparison should permit a reviewer to trace the matter from decision to outcome. This may require approval and change-control records, authorship and identity controls proportionate to risk, marking criteria and calibrated judgement, and appeal and correction records, supported by assessment maps to learning outcomes and analysis of results and differential outcomes. As regards AI-supported assessment, sampling remains insufficient where it excludes a material group or cannot resolve contradictory evidence or recurrence.

Patterns requiring examination

The review method for AI-supported assessment should be reproducible. The review should trace selected records to source, reconcile totals across systems, quantify missing and late submissions, review manual adjustments and retain a revision history. Escalate discrepancies that could alter a published conclusion or individual outcome. Within the scope under review, the retained analysis should be reproducible from the selected evidence, decision rule and recorded reasons for accepted exceptions.

Publication of findings on AI-supported assessment should distinguish observed values, estimates and interpretation.

Decisions concerning the analysis should remain traceable to the information available for the stated reference period. For AI-supported assessment, changes in condition, evidence, method and interpretation should be recorded separately when a conclusion is revised. Without this distinction, a reporting change may be mistaken for improvement or deterioration in educational practice.

Implications for decision-makers

The analysis of AI-supported assessment should remain within the limits of the evidence. A single indicator rarely provides an adequate account of quality. Quantitative evidence should be considered with implementation records and the experience of affected learners. For comparative analysis, reliability without validity produces consistent but potentially irrelevant results. Validity without adequate consistency may expose learners to unequal judgement. If uncertainty could change a consequential decision, additional evidence or a narrower conclusion is required.

For AI-supported assessment, where responsibilities for delivery are shared with partners, suppliers or several public bodies, responsibility should be mapped across the complete service. Contractual or inter-agency arrangements should identify who holds records, informs learners and acts on incidents. Within the scope under review, multiple delivery partners do not justify fragmented accountability or remedy.

A reliable conclusion requires corroboration across the material scope.