Examines AI-supported assessment, addressing the available evidence, source definitions, coverage, comparability, uncertainty and limits on inference.
The present attention to AI-supported assessment follows the rapid adoption of generative AI tools and requires a careful distinction between public commitment, institutional practice and demonstrated result. Evidence concerning AI-supported assessment should inform action without implying a level of precision, coverage or causal certainty that the underlying data cannot support. Learner effect, institutional duty and proper resource use should inform the judgement. System context should determine the appropriate administrative arrangement within the governing requirements.
The relevant context is provided by rapid adoption of generative AI tools. Its relevance to the analysis should be assessed against the affected jurisdiction, learner population and form of provision.
For AI-supported assessment, implementation of the comparison should be organised around a decision that can be tested. A sound interpretation should identify the unit of analysis, reference period, denominator, exclusions, missing values and any change in definition or collection practice. The implementation record should link purpose, authority, resources, operation and reported result.
Analytical scope
Any indicator used in relation to the measure should distinguish description from causal explanation. In work concerning AI-supported assessment, interpretation should retain uncertainty, distributional differences and limits on generalisation.
For the measure, an average may improve while a material group experiences no improvement or a worse outcome. For decisions concerning AI-supported assessment, disaggregation should follow a defined public-interest question and should protect confidentiality where small numbers could identify individuals. The judgement should state its supporting evidence and any condition limiting application to the declared scope.
The principal risks associated with the issue should be assessed as connected conditions. A provider should also consider tasks that do not assess the stated outcome and reasonable adjustment altering the assessed outcome. In the context of AI-supported assessment, the control response should reflect whether an affected learner can identify the error and obtain an effective remedy in time.
Evidence collection should be designed around the decision question rather than administrative convenience. For the analysis, the most relevant material is likely to include assessment maps to learning outcomes, appeal and correction records, approval and change-control records, and authorship and identity controls proportionate to risk. In work concerning AI-supported assessment, each source has limitations; confidence depends on corroboration between independent records and transparent treatment of uncertainty.
Definitions and data coverage
Implementation of AI-supported assessment can be tested without imposing unnecessary reporting. For the analysis, the reviewer should examine results by relevant learner, programme, location and delivery characteristics; compare both levels and rates of change; and test whether observed gaps persist after differences in coverage and prior conditions are considered. Reuse of existing information is appropriate only where its purpose, scope and reliability correspond to the decision under review.
The analytical record for AI-supported assessment should state the research question, data source, unit of analysis, reference period, coverage, exclusions, treatment of missing values and principal limitations.
Records relating to the analysis should preserve both the conclusion and its limits. For AI-supported assessment, the correction record should state what the new evidence changes and which earlier conclusions or decisions require review. Where reliance has occurred, correction may require review of affected decisions as well as amendment of published information.
Use of the findings
Proportionality in relation to AI-supported assessment does not mean reduced protection for learners exposed to greater risk. Reliability without validity produces consistent but potentially irrelevant results. Validity without adequate consistency may expose learners to unequal judgement. Missing or delayed information may be patterned rather than random. No exception should continue without a documented basis, accountable approval and scheduled review.
Accountability for AI-supported assessment should follow decision-making authority. Where work is delegated, the record should continue to identify who is accountable for material consequences to learners.
Within the scope under review, neither one indicator nor one control can establish the complete position on the issue.
The reference period and version of the source identified in Rapid adoption of generative AI tools should remain traceable. A later correction, expanded dataset or revised classification may justify a new conclusion, but it should not reconstruct the earlier record without stating what changed and how the change affects comparability.