Examines how improvement in artificial intelligence in education policy should be designed, implemented and tested against the intended educational outcome.
The Beijing Consensus adopted in May 2019 provides the immediate context for the causes of underperformance in artificial intelligence in education policy. Corrective action concerning assessing the improvement priority should address the identified cause, assign responsibility and set a review period. Residual risk should remain open until sustained improvement is demonstrated. The scope should include every materially affected setting, with differences in location, programme, delivery mode and learner population kept visible. Evidence of formal policy should not be treated as evidence of uniform implementation.
The Beijing Consensus on Artificial Intelligence and Education was adopted in May 2019. It addresses policy planning, management, teaching, learning, skills, lifelong learning, inclusion, gender equality, data and research. It promotes human-centred and equitable use rather than technology adoption as an end in itself. Authorities and providers should therefore connect each proposed use to an educational purpose, governance responsibility and evidence of benefit and risk.
The present position
The formal status of the Beijing Consensus adopted in May 2019 should be preserved in any public account. Adoption records an agreed instrument or policy position; it does not necessarily make every provision directly enforceable in every jurisdiction. For the causes of underperformance in artificial intelligence in education policy, the instrument should be used to identify the intended direction, the actors addressed and the implementation measures that remain necessary. Domestic law and authorised guidance continue to determine specific legal duties.
The position on assessing the intervention should be established through proportionate evidence and should remain open to correction when material new information becomes available. Responsibility for the corrective programme should be identifiable at each consequential decision point. Delegation should identify both the operating role and the body retaining oversight of learner impact. System-level policy does not displace provider responsibility for the quality, integrity and lawful operation of its provision. The allocation of responsibility should prevent gaps between system oversight and institutional operation.
The review method for the intervention should connect the question under examination to suitable evidence and a conclusion no broader than the tested scope. The evidential basis for the affected practice should identify source, period, coverage and material limitations. Corroboration is required where a single record cannot support the decision. End-to-end assurance is required because individual functions may operate as designed while the combined process fails. The distinction matters because evidence may appear sufficient while addressing a different population, period or outcome.
Governance of the improvement priority requires a clear allocation of authority, information and follow-through. Material matters should be referred to the body authorised to act or accept residual risk. Improvement work on the corrective programme should begin with a verified problem, defined baseline and measurable outcome. Completion should depend on evidence of effect rather than completion of planned activity. A decision should not be closed at the operating level where material impact, conflict or a significant evidential gap remains unresolved.
Assurance concerning the affected practice should state the scope examined, evidence relied upon and any condition preventing a complete conclusion. Unsupported elements should remain open. In the present context, unequal performance across learner groups, unclear responsibility between providers and suppliers and unverified outputs entering teaching or assessment may produce acceptable aggregate reporting while individual learners remain exposed to material disadvantage. Adverse cases should form part of the sample wherever they may reveal a material control weakness.
The record for the affected practice should identify the responsible function, decision authority and escalation route. Gaps between public oversight and provider control should not remain implicit. Evidence outside the relevant period or scope should be identified and given no more weight than its limitations permit. An unresolved contradiction is a limitation on the conclusion and should be reported as such.
Responsibilities and material risks
A proportionate method is available for the causes of underperformance in artificial intelligence in education policy. The principal risks associated with assessing the matter under review should be assessed as connected conditions. A failed safeguard may conceal another weakness or prevent timely correction. The review should determine whether correction of an individual case is sufficient or broader action is required. Adverse cases and unresolved contradictions should be retained because they may reveal limitations concealed by an average result.
Data used for the matter under review should be interpreted against stable definitions and an identifiable population. Changes in method, definition or series should remain separate from changes in the underlying result. Reporting should distinguish work performed from the outcome demonstrated after implementation. Closure reporting should not obscure unresolved action or risk retained by the responsible authority.
The assurance record for the improvement priority should permit another competent reviewer to understand the evidence, method, judgement and treatment of material exceptions. A conclusion on the intervention should extend no further than the available evidence permits. Missing populations, inconsistent records and unresolved exceptions should be reported with the finding. Accuracy measured in one setting may not transfer to another population, language, curriculum or decision context. Any indicator used in relation to assessing the causes of underperformance in the intervention should distinguish description from causal explanation. Interpretation should retain uncertainty, distributional differences and limits on generalisation. Data used for the improvement priority should be interpreted against stable definitions and an identifiable population. A revision or break in series should not be reported as a change in performance. A finding should not be separated from limitations capable of changing how it is understood or applied.
Governance of the matter under review requires a clear allocation of authority, information and follow-through. Traceable source and version information allow genuine improvement to be distinguished from administrative revision. The evidential history should preserve conclusions that were operative when a material decision was made.
The record for the intervention should identify the responsible function, decision authority and escalation route. Relevant evidence should reach the body authorised to commit resources, amend policy or accept residual risk, and its judgement should be recorded. The operating function may change, but responsibility for oversight and learner protection should remain clear.
For the corrective programme, the implementation record should distinguish binding duties, policy expectations and institutional choices, including any transition or jurisdictional limitation. Neither administrative activity nor general assurance should obscure the intended result or its effect on learners. Where evidence cannot support assurance, the limitation should be reported and corrective work should remain open.