ICEQC-R-2018-05 — Evidence Standards for Claims of System-Level Learning Improvement cover

Rapport de recherche thématique

ICEQC-R-2018-05 — Evidence Standards for Claims of System-Level Learning Improvement

A global standards interpretation of baselines, trend comparability, causal contribution, distribution and public proof

Date de publication
Catégorie de recherche
Interprétation des normes
Modèle de rapport
Étude d'interprétation des normes
Portée géographique
Global
Date limite de soumission des preuves
Organisme responsable
Direction de la recherche et des politiques de l'ICEQC
ICEQC-R-2018-05 — Evidence Standards for Claims of System-Level Learning Improvement cover

Publication record

This is the controlled English edition. Evidence and institutional status are stated as at the evidence cut-off date.

Executive summary

Key findings

    Scope and method

    Part II

    Purpose and interpretive frame

    1

    What national averages conceal

    What national averages conceal within Part II — Purpose and interpretive frame places this proposition within the evidence required for credible evidence of system-level learning improvement: The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns national participation or attainment. Its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, What national averages conceal within Part II — Purpose and interpretive frame must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is the weighted experience of the population included in the denominator. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, What national averages conceal within Part II — Purpose and interpretive frame must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: distribution within countries, not a league table between them. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, What national averages conceal within Part II — Purpose and interpretive frame must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    What national averages conceal within Part II — Purpose and interpretive frame places this proposition within the evidence required for credible evidence of system-level learning improvement: Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include remote rural learners, urban informal settlements and displaced populations. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, What national averages conceal within Part II — Purpose and interpretive frame must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    The final public statement should identify remaining uncertainty and the next evidentiary step rather than convert a partial result into assurance. The decision record should connect the finding to a responsible authority, available intervention and review date. An indicator is useful when it can alter service location, staffing, language support, accessibility, household assistance or another defined condition. It should not be used to rank schools or communities where differences in population and opportunity to learn are uncontrolled. Monitoring after action must preserve the original baseline and follow both reach and outcome. Improvement in an average does not demonstrate that the intended group benefited; participation, intensity and learner experience should be checked. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, What national averages conceal within Part II — Purpose and interpretive frame must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    2

    Marginalisation as accumulated disadvantage

    For Marginalisation as accumulated disadvantage within Part II — Purpose and interpretive frame, a responsible judgement on credible evidence of system-level learning improvement should address the following point: The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. Marginalisation as accumulated disadvantage should be approached as a defined measurement problem. The substantive interest is distance from a socially secured educational minimum, observed for a population and period that are stated before calculation. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Marginalisation as accumulated disadvantage within Part II — Purpose and interpretive frame must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined. School returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. Reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use the joint effect of exclusion, weak provision and adverse social conditions. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Marginalisation as accumulated disadvantage within Part II — Purpose and interpretive frame must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. Percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. The comparison should identify the reference category but avoid presenting it as a natural norm. Policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that multiple indicators read together over the learner's course. Disparity measures should not replace the underlying distributions. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Marginalisation as accumulated disadvantage within Part II — Purpose and interpretive frame must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    Combining unknown observations with the majority group biases both estimates and obscures the weakness. Where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for children facing poverty, gender disadvantage, disability or minority status. Their circumstances may alter access to enumeration, classification and the service being measured. The review should record non-response, unknown status and excluded locations separately. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Marginalisation as accumulated disadvantage within Part II — Purpose and interpretive frame must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    Targets should specify the population expected to benefit and guard against gains achieved by concentrating on learners nearest a threshold. Accountability lies in changed opportunity, not in the favourable movement of an indicator alone. Policy use should begin with a question that the competent body can answer. The evidence may justify further enquiry, immediate removal of a barrier, redistribution of resources or evaluation of an existing measure. These are different decisions and require different certainty. Urgent protection need not await a perfect estimate where credible evidence shows serious exclusion, but long-term allocation should be reviewed as coverage improves. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Marginalisation as accumulated disadvantage within Part II — Purpose and interpretive frame must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    3

    From monitoring commitment to decision

    Evidence from From monitoring commitment to decision within Part II — Purpose and interpretive frame bears on credible evidence of system-level learning improvement through this institutional requirement: The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. For from monitoring commitment to decision, the first requirement is conceptual clarity. The study seeks evidence on evidence capable of changing resource or service decisions; it does not infer a learner's circumstances from a national or regional mean. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, From monitoring commitment to decision within Part II — Purpose and interpretive frame must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    A defensible statistic would be based on a link between observed disparity, responsible body and remedy. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression. The estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. When several sources exist, consistency is evidence to consider, not proof that common error is absent. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, From monitoring commitment to decision within Part II — Purpose and interpretive frame must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    From monitoring commitment to decision within Part II — Purpose and interpretive frame places this proposition within the evidence required for credible evidence of system-level learning improvement: Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: indicators selected for action rather than visibility. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, From monitoring commitment to decision within Part II — Purpose and interpretive frame must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    An adequate equity account asks whether groups absent from routine plans and budget classifications are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. Missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, From monitoring commitment to decision within Part II — Purpose and interpretive frame must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    Local interpretation is valuable because national classifications cannot capture every barrier; it should operate within common definitions sufficient for aggregation. Communities should be able to question both the category and the conclusion drawn from it. If a measure creates incentives to exclude difficult cases, narrow the denominator or reclassify non-completion, an independent check is required. The report should recognise such behaviour as a measurement risk without assuming misconduct in every discrepancy. A proportionate response records who acts, the condition to be changed and the evidence that will show whether reach was equitable. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, From monitoring commitment to decision within Part II — Purpose and interpretive frame must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    Part III

    Population and denominator

    4

    Defining the population entitled to education

    The educational consequence in Defining the population entitled to education within Part III — Population and denominator gives practical meaning to credible evidence of system-level learning improvement: The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of resident and temporarily absent learners within the relevant age or programme group begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Defining the population entitled to education within Part III — Population and denominator must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    A census figure should disclose enumeration rules. These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is a denominator consistent with the right, level and reference period under review. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Defining the population entitled to education within Part III — Population and denominator must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that census, survey and administrative estimates reconciled openly. A responsible commentary distinguishes observation, calculation and interpretation. It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Defining the population entitled to education within Part III — Population and denominator must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community. Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. For unregistered residents, migrants and displaced children, a single group label may conceal important internal differences. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Defining the population entitled to education within Part III — Population and denominator must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    A policy conclusion should be no broader than that evidence. Subsequent reports should distinguish real change from late reporting, revised population estimates and altered definitions. Where a disparity remains, the responsible body should state whether the obstacle is knowledge, authority, resources or implementation. This makes the indicator a means of scrutiny rather than a decorative measure of concern. Public accountability requires a concise explanation of the result and its boundary. Readers should be able to see who was counted, who was not, what period the figure covers, how large the underlying population is and which comparisons are justified. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Defining the population entitled to education within Part III — Population and denominator must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    5

    Age, grade and programme populations

    Public accountability for Age, grade and programme populations within Part III — Population and denominator requires a reasoned finding about credible evidence of system-level learning improvement: The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of age, grade and programme populations lies in making unequal educational experience observable. Here, the relevant phenomenon is age-specific, grade-specific and programme-specific participation, not the administrative convenience of the available categories. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Age, grade and programme populations within Part III — Population and denominator must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is separate denominators for different educational questions. Numerator, denominator, reference date, unit and exclusions should appear together. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Age, grade and programme populations within Part III — Population and denominator must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: exact age, school age and enrolled population never substituted silently. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Age, grade and programme populations within Part III — Population and denominator must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    The distributional review must deliberately include over-age entrants and learners repeating grades. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Age, grade and programme populations within Part III — Population and denominator must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    6

    Population movement and disrupted residence

    Population movement and disrupted residence within Part III — Population and denominator places this proposition within the evidence required for credible evidence of system-level learning improvement: Its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns education status amid migration, displacement and return. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Population movement and disrupted residence within Part III — Population and denominator must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use dated location and residence rules with sensitivity to mobility. This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined. School returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. Reconciliation should record what each source can and cannot represent. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Population movement and disrupted residence within Part III — Population and denominator must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    Population movement and disrupted residence within Part III — Population and denominator places this proposition within the evidence required for credible evidence of system-level learning improvement: Percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. The comparison should identify the reference category but avoid presenting it as a natural norm. Policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that origin, current location and service responsibility distinguished. Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Population movement and disrupted residence within Part III — Population and denominator must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    Population movement and disrupted residence within Part III — Population and denominator places this proposition within the evidence required for credible evidence of system-level learning improvement: Where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for families affected by conflict, disaster and seasonal movement. Their circumstances may alter access to enumeration, classification and the service being measured. The review should record non-response, unknown status and excluded locations separately. Combining unknown observations with the majority group biases both estimates and obscures the weakness. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Population movement and disrupted residence within Part III — Population and denominator must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-20] [REF-21]

    Part IV

    Access and participation

    7

    Entry at the official starting age

    Entry at the official starting age should be approached as a defined measurement problem. The substantive interest is timely admission to the first grade of primary education, observed for a population and period that are stated before calculation. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Entry at the official starting age within Part IV — Access and participation must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    A defensible statistic would be based on new entrants of official age relative to the corresponding population. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression. The estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. When several sources exist, consistency is evidence to consider, not proof that common error is absent. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Entry at the official starting age within Part IV — Access and participation must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    For Entry at the official starting age within Part IV — Access and participation, a responsible judgement on credible evidence of system-level learning improvement should address the following point: Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: late entry examined alongside non-entry. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Entry at the official starting age within Part IV — Access and participation must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-20] [REF-21]

    Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. Missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. An adequate equity account asks whether children facing fees, distance, disability or documentation barriers are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Entry at the official starting age within Part IV — Access and participation must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    8

    Attendance beyond enrolment

    Attendance beyond enrolment within Part IV — Access and participation places this proposition within the evidence required for credible evidence of system-level learning improvement: The report should state the educational consequence before choosing a gap, ratio, threshold or rank. For attendance beyond enrolment, the first requirement is conceptual clarity. The study seeks evidence on actual participation during a stated recent period; it does not infer a learner's circumstances from a national or regional mean. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Attendance beyond enrolment within Part IV — Access and participation must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    Attendance beyond enrolment within Part IV — Access and participation places this proposition within the evidence required for credible evidence of system-level learning improvement: These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is presence measured independently of registration status. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Attendance beyond enrolment within Part IV — Access and participation must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-20] [REF-21]

    A responsible commentary distinguishes observation, calculation and interpretation. It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference. Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that frequency, season and reasons for absence retained. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Attendance beyond enrolment within Part IV — Access and participation must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    For working children, carers and learners affected by illness, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community. Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Attendance beyond enrolment within Part IV — Access and participation must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    9

    Out-of-school status

    Out-of-school status within Part IV — Access and participation places this proposition within the evidence required for credible evidence of system-level learning improvement: A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of children in the relevant age group not participating at the defined level begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement. The measure should preserve both the observed level and the distribution relevant to the claim. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Out-of-school status within Part IV — Access and participation must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-20] [REF-21]

    A discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is a transparent residual from compatible population and participation concepts. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Out-of-school status within Part IV — Access and participation must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: never-enrolled and formerly enrolled children separated. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Out-of-school status within Part IV — Access and participation must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    For Out-of-school status within Part IV — Access and participation, a responsible judgement on credible evidence of system-level learning improvement should address the following point: Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include children invisible to school registers. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Out-of-school status within Part IV — Access and participation must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    Part V

    Progression and completion

    10

    Repetition and grade survival

    A decision concerning Repetition and grade survival within Part V — Progression and completion should connect credible evidence of system-level learning improvement to the following evidential premise: The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of repetition and grade survival lies in making unequal educational experience observable. Here, the relevant phenomenon is movement through grades without avoidable delay or exit, not the administrative convenience of the available categories. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Repetition and grade survival within Part V — Progression and completion must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined. School returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. Reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use cohort or reconstructed-cohort evidence with explicit assumptions. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Repetition and grade survival within Part V — Progression and completion must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. Percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. The comparison should identify the reference category but avoid presenting it as a natural norm. Policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that repeaters distinguished from re-entrants and transfers. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Repetition and grade survival within Part V — Progression and completion must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    The review should record non-response, unknown status and excluded locations separately. Combining unknown observations with the majority group biases both estimates and obscures the weakness. Where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for learners in overcrowded or intermittently operating schools. Their circumstances may alter access to enumeration, classification and the service being measured. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Repetition and grade survival within Part V — Progression and completion must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    11

    Transition between education levels

    For Transition between education levels within Part V — Progression and completion, a responsible judgement on credible evidence of system-level learning improvement should address the following point: Its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns entry to the next level after completion of the preceding one. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Transition between education levels within Part V — Progression and completion must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    The estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. When several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on matched completion and new-entry populations over coherent periods. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Transition between education levels within Part V — Progression and completion must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: capacity constraints distinguished from learner attainment. Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Transition between education levels within Part V — Progression and completion must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    Missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. An adequate equity account asks whether rural learners and those unable to relocate are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Transition between education levels within Part V — Progression and completion must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    12

    Completion and educational entitlement

    Completion and educational entitlement should be approached as a defined measurement problem. The substantive interest is finishing the final grade or meeting recognised programme requirements, observed for a population and period that are stated before calculation. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Completion and educational entitlement within Part V — Progression and completion must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    For Completion and educational entitlement within Part V — Progression and completion, a responsible judgement on credible evidence of system-level learning improvement should address the following point: These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is completion defined separately from sitting or passing an examination. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Completion and educational entitlement within Part V — Progression and completion must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference. Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that late completion and alternative pathways reported. A responsible commentary distinguishes observation, calculation and interpretation. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Completion and educational entitlement within Part V — Progression and completion must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. For over-age learners and people returning after interruption, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Completion and educational entitlement within Part V — Progression and completion must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    Part VI

    Learning and assessment

    13

    Minimum learning outcomes

    For Minimum learning outcomes within Part VI — Learning and assessment, a responsible judgement on credible evidence of system-level learning improvement should address the following point: The report should state the educational consequence before choosing a gap, ratio, threshold or rank. For minimum learning outcomes, the first requirement is conceptual clarity. The study seeks evidence on demonstrated knowledge or skill against a declared domain; it does not infer a learner's circumstances from a national or regional mean. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Minimum learning outcomes within Part VI — Learning and assessment must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is assessment evidence whose population and conditions are known. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Minimum learning outcomes within Part VI — Learning and assessment must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    Interpretation follows this limitation: results interpreted with opportunity to learn and participation. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Minimum learning outcomes within Part VI — Learning and assessment must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include learners taught in an unfamiliar language. These populations may be missing not only from good outcomes but from the denominator itself. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Minimum learning outcomes within Part VI — Learning and assessment must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    14

    Assessment participation

    For Assessment participation within Part VI — Learning and assessment, a responsible judgement on credible evidence of system-level learning improvement should address the following point: A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of who was eligible, present, absent and excluded from testing begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement. The measure should preserve both the observed level and the distribution relevant to the claim. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Assessment participation within Part VI — Learning and assessment must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    Source coverage must be tested before sources are combined. School returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. Reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use a participation profile accompanying every result distribution. This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Assessment participation within Part VI — Learning and assessment must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    Policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that non-participation never treated as low attainment or ignored. Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. Percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. The comparison should identify the reference category but avoid presenting it as a natural norm. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Assessment participation within Part VI — Learning and assessment must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    Particular scrutiny is required for learners with disabilities and remote candidates. Their circumstances may alter access to enumeration, classification and the service being measured. The review should record non-response, unknown status and excluded locations separately. Combining unknown observations with the majority group biases both estimates and obscures the weakness. Where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Assessment participation within Part VI — Learning and assessment must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    15

    Distribution of achievement

    The population addressed in Distribution of achievement within Part VI — Learning and assessment tests the adequacy of credible evidence of system-level learning improvement in this respect: The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of distribution of achievement lies in making unequal educational experience observable. Here, the relevant phenomenon is variation across the full score or proficiency distribution, not the administrative convenience of the available categories. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Distribution of achievement within Part VI — Learning and assessment must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression. The estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. When several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on percentiles and threshold shares beside a mean. The definition should be fixed for the comparison at hand and deviations recorded. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Distribution of achievement within Part VI — Learning and assessment must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: uncertainty and scale properties stated. Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Distribution of achievement within Part VI — Learning and assessment must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    An adequate equity account asks whether learners concentrated below a minimum proficiency threshold are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. Missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Distribution of achievement within Part VI — Learning and assessment must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    Part VII

    Gender and household resources

    16

    Gender parity and its limits

    Evidence from Gender parity and its limits within Part VII — Gender and household resources bears on credible evidence of system-level learning improvement through this institutional requirement: Its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns differences between girls and boys in access, progression and learning. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Gender parity and its limits within Part VII — Gender and household resources must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    Evidence from Gender parity and its limits within Part VII — Gender and household resources bears on credible evidence of system-level learning improvement through this institutional requirement: These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is female-to-male ratios read with levels and absolute gaps. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Gender parity and its limits within Part VII — Gender and household resources must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    It does not assign cause from a cross-sectional difference. Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that parity not confused with adequacy for either group. A responsible commentary distinguishes observation, calculation and interpretation. It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Gender parity and its limits within Part VII — Gender and household resources must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community. Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. For girls in poor rural households and boys exposed to hazardous work, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Gender parity and its limits within Part VII — Gender and household resources must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-20] [REF-21]

    17

    Household wealth gradients

    Household wealth gradients should be approached as a defined measurement problem. The substantive interest is education outcomes across relative household resource groups, observed for a population and period that are stated before calculation. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Household wealth gradients within Part VII — Gender and household resources must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is a documented asset or consumption classification within each setting. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Household wealth gradients within Part VII — Gender and household resources must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: wealth ranks not assumed equivalent across countries or time. The comparison should show the level for each group as well as any ratio or gap. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Household wealth gradients within Part VII — Gender and household resources must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-20] [REF-21]

    The distributional review must deliberately include children in the poorest quintile and those near classification boundaries. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Household wealth gradients within Part VII — Gender and household resources must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    18

    Costs borne by households

    Evidence from Costs borne by households within Part VII — Gender and household resources bears on credible evidence of system-level learning improvement through this institutional requirement: The report should state the educational consequence before choosing a gap, ratio, threshold or rank. For costs borne by households, the first requirement is conceptual clarity. The study seeks evidence on fees, materials, transport, clothing and foregone labour; it does not infer a learner's circumstances from a national or regional mean. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Costs borne by households within Part VII — Gender and household resources must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    Reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use participation read against direct and indirect education costs. This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined. School returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Costs borne by households within Part VII — Gender and household resources must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-20] [REF-21]

    For Costs borne by households within Part VII — Gender and household resources, a responsible judgement on credible evidence of system-level learning improvement should address the following point: Percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. The comparison should identify the reference category but avoid presenting it as a natural norm. Policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that nominal fee abolition checked against remaining expenditure. Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Costs borne by households within Part VII — Gender and household resources must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    For Costs borne by households within Part VII — Gender and household resources, a responsible judgement on credible evidence of system-level learning improvement should address the following point: Where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for large families and households hit by economic crisis. Their circumstances may alter access to enumeration, classification and the service being measured. The review should record non-response, unknown status and excluded locations separately. Combining unknown observations with the majority group biases both estimates and obscures the weakness. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Costs borne by households within Part VII — Gender and household resources must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    Part VIII

    Place and service geography

    19

    Rural and urban residence

    Evidence from Rural and urban residence within Part VIII — Place and service geography bears on credible evidence of system-level learning improvement through this institutional requirement: A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of education outcomes by a declared settlement classification begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement. The measure should preserve both the observed level and the distribution relevant to the claim. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Rural and urban residence within Part VIII — Place and service geography must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-20] [REF-21]

    A defensible statistic would be based on residence linked to service availability and travel conditions. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression. The estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. When several sources exist, consistency is evidence to consider, not proof that common error is absent. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Rural and urban residence within Part VIII — Place and service geography must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    Evidence from Rural and urban residence within Part VIII — Place and service geography bears on credible evidence of system-level learning improvement through this institutional requirement: Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: national rural definitions preserved and comparison limitations stated. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Rural and urban residence within Part VIII — Place and service geography must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. Missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. An adequate equity account asks whether remote villages, pastoral populations and peri-urban settlements are represented at each stage: population frame, collection, valid response, classification, analysis and publication. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Rural and urban residence within Part VIII — Place and service geography must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    20

    Subnational administrative disparity

    Within Subnational administrative disparity within Part VIII — Place and service geography, the authority should interpret credible evidence of system-level learning improvement against this condition: The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of subnational administrative disparity lies in making unequal educational experience observable. Here, the relevant phenomenon is variation between provinces, districts or comparable areas, not the administrative convenience of the available categories. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Subnational administrative disparity within Part VIII — Place and service geography must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    The educational consequence in Subnational administrative disparity within Part VIII — Place and service geography gives practical meaning to credible evidence of system-level learning improvement: These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is area estimates with population size and precision. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Subnational administrative disparity within Part VIII — Place and service geography must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that administrative rankings not mistaken for causal explanations. A responsible commentary distinguishes observation, calculation and interpretation. It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Subnational administrative disparity within Part VIII — Place and service geography must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    For small districts and areas with incomplete reporting, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community. Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Subnational administrative disparity within Part VIII — Place and service geography must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    21

    Distance, isolation and transport

    The educational consequence in Distance, isolation and transport within Part VIII — Place and service geography gives practical meaning to credible evidence of system-level learning improvement: Its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns physical accessibility of the nearest appropriate service. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Distance, isolation and transport within Part VIII — Place and service geography must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is travel time, route safety and seasonal interruption rather than straight-line distance alone. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Distance, isolation and transport within Part VIII — Place and service geography must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: household reports and facility mapping reconciled. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Distance, isolation and transport within Part VIII — Place and service geography must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    Evidence from Distance, isolation and transport within Part VIII — Place and service geography bears on credible evidence of system-level learning improvement through this institutional requirement: Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include learners with limited mobility and communities cut off seasonally. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Distance, isolation and transport within Part VIII — Place and service geography must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    Part IX

    Disability, language and identity

    22

    Disability-sensitive education data

    Disability-sensitive education data should be approached as a defined measurement problem. The substantive interest is participation and learning by functional difficulty and support requirement, observed for a population and period that are stated before calculation. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Disability-sensitive education data within Part IX — Disability, language and identity must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    Measurement should use questions designed for comparable reporting without defining the child by diagnosis alone. This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined. School returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. Reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Disability-sensitive education data within Part IX — Disability, language and identity must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. Percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. The comparison should identify the reference category but avoid presenting it as a natural norm. Policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that identification, environment and accommodation kept analytically distinct. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Disability-sensitive education data within Part IX — Disability, language and identity must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    The review should record non-response, unknown status and excluded locations separately. Combining unknown observations with the majority group biases both estimates and obscures the weakness. Where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for learners whose impairments are not recorded by schools. Their circumstances may alter access to enumeration, classification and the service being measured. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Disability-sensitive education data within Part IX — Disability, language and identity must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    23

    Language of home and instruction

    The educational consequence in Language of home and instruction within Part IX — Disability, language and identity gives practical meaning to credible evidence of system-level learning improvement: The report should state the educational consequence before choosing a gap, ratio, threshold or rank. For language of home and instruction, the first requirement is conceptual clarity. The study seeks evidence on alignment between learner language, teaching and assessment; it does not infer a learner's circumstances from a national or regional mean. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Language of home and instruction within Part IX — Disability, language and identity must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    The estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. When several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on language categories reflecting local use and instructional practice. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Language of home and instruction within Part IX — Disability, language and identity must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: small groups not erased through broad national labels. Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Language of home and instruction within Part IX — Disability, language and identity must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    Field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. Missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. An adequate equity account asks whether minority-language and multilingual learners are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Language of home and instruction within Part IX — Disability, language and identity must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    24

    Ethnicity, indigeneity and protected identity

    The educational consequence in Ethnicity, indigeneity and protected identity within Part IX — Disability, language and identity gives practical meaning to credible evidence of system-level learning improvement: A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of disparity associated with historically excluded identity groups begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement. The measure should preserve both the observed level and the distribution relevant to the claim. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Ethnicity, indigeneity and protected identity within Part IX — Disability, language and identity must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    Public accountability for Ethnicity, indigeneity and protected identity within Part IX — Disability, language and identity requires a reasoned finding about credible evidence of system-level learning improvement: These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is lawful, voluntary and contextually meaningful classification. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Ethnicity, indigeneity and protected identity within Part IX — Disability, language and identity must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    Use of the indicator is bounded by the principle that self-identification protected and non-response reported. A responsible commentary distinguishes observation, calculation and interpretation. It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference. Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Ethnicity, indigeneity and protected identity within Part IX — Disability, language and identity must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. For communities exposed to discrimination or forced assimilation, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Ethnicity, indigeneity and protected identity within Part IX — Disability, language and identity must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    Part X

    Conflict, disaster and mobility

    25

    Education under conflict and insecurity

    Equity analysis in Education under conflict and insecurity within Part X — Conflict, disaster and mobility qualifies any conclusion about credible evidence of system-level learning improvement: The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of education under conflict and insecurity lies in making unequal educational experience observable. Here, the relevant phenomenon is access, attendance and learning where violence alters service and movement, not the administrative convenience of the available categories. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Education under conflict and insecurity within Part X — Conflict, disaster and mobility must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is location- and time-specific observation with explicit coverage gaps. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Education under conflict and insecurity within Part X — Conflict, disaster and mobility must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: absence caused by insecurity distinguished from ordinary dropout. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Education under conflict and insecurity within Part X — Conflict, disaster and mobility must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include learners in insecure areas and host communities. These populations may be missing not only from good outcomes but from the denominator itself. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Education under conflict and insecurity within Part X — Conflict, disaster and mobility must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    27

    Refugees, displaced persons and migrants

    Refugees, displaced persons and migrants should be approached as a defined measurement problem. The substantive interest is educational participation across changing legal and residential situations, observed for a population and period that are stated before calculation. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Refugees, displaced persons and migrants within Part X — Conflict, disaster and mobility must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression. The estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. When several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on status, origin, current location and service access recorded separately. The definition should be fixed for the comparison at hand and deviations recorded. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Refugees, displaced persons and migrants within Part X — Conflict, disaster and mobility must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: mobility never converted into duplicate enrolment or unexplained disappearance. Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Refugees, displaced persons and migrants within Part X — Conflict, disaster and mobility must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-20] [REF-21]

    Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. Missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. An adequate equity account asks whether undocumented migrants, refugees and internally displaced learners are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Refugees, displaced persons and migrants within Part X — Conflict, disaster and mobility must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    Part XI

    School conditions and teachers

    28

    Teacher availability and distribution

    Public accountability for Teacher availability and distribution within Part XI — School conditions and teachers requires a reasoned finding about credible evidence of system-level learning improvement: The report should state the educational consequence before choosing a gap, ratio, threshold or rank. For teacher availability and distribution, the first requirement is conceptual clarity. The study seeks evidence on access to competent teaching across schools and subjects; it does not infer a learner's circumstances from a national or regional mean. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Teacher availability and distribution within Part XI — School conditions and teachers must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    A decision concerning Teacher availability and distribution within Part XI — School conditions and teachers should connect credible evidence of system-level learning improvement to the following evidential premise: These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is teachers present and assigned relative to learner and curriculum need. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Teacher availability and distribution within Part XI — School conditions and teachers must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-20] [REF-21]

    A responsible commentary distinguishes observation, calculation and interpretation. It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference. Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that payroll totals not substituted for classroom availability. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Teacher availability and distribution within Part XI — School conditions and teachers must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community. Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. For schools serving poor, remote or displaced communities, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Teacher availability and distribution within Part XI — School conditions and teachers must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    29

    Class size, multi-grade teaching and time

    Public accountability for Class size, multi-grade teaching and time within Part XI — School conditions and teachers requires a reasoned finding about credible evidence of system-level learning improvement: A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of the instructional conditions experienced by learners begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement. The measure should preserve both the observed level and the distribution relevant to the claim. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Class size, multi-grade teaching and time within Part XI — School conditions and teachers must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-20] [REF-21]

    The preferred construction is class organisation, scheduled time and delivered time considered together. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Class size, multi-grade teaching and time within Part XI — School conditions and teachers must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: simple pupil-teacher ratios not treated as a complete quality measure. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Class size, multi-grade teaching and time within Part XI — School conditions and teachers must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    The distributional review must deliberately include early grades and mixed-age classes. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Class size, multi-grade teaching and time within Part XI — School conditions and teachers must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    30

    Materials, facilities and basic services

    The responsible body for Materials, facilities and basic services within Part XI — School conditions and teachers must state how credible evidence of system-level learning improvement resolves this question: The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of materials, facilities and basic services lies in making unequal educational experience observable. Here, the relevant phenomenon is usable learning resources and safe, accessible school conditions, not the administrative convenience of the available categories. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Materials, facilities and basic services within Part XI — School conditions and teachers must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    School returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. Reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use availability joined to condition, accessibility and regular use. This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Materials, facilities and basic services within Part XI — School conditions and teachers must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    Evidence from Materials, facilities and basic services within Part XI — School conditions and teachers bears on credible evidence of system-level learning improvement through this institutional requirement: Percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. The comparison should identify the reference category but avoid presenting it as a natural norm. Policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that delivery counts checked against learner access. Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Materials, facilities and basic services within Part XI — School conditions and teachers must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    Evidence from Materials, facilities and basic services within Part XI — School conditions and teachers bears on credible evidence of system-level learning improvement through this institutional requirement: Where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for learners in temporary or damaged premises. Their circumstances may alter access to enumeration, classification and the service being measured. The review should record non-response, unknown status and excluded locations separately. Combining unknown observations with the majority group biases both estimates and obscures the weakness. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Materials, facilities and basic services within Part XI — School conditions and teachers must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    Part XII

    Finance and distribution

    31

    Public spending by level and function

    A decision concerning Public spending by level and function within Part XII — Finance and distribution should connect credible evidence of system-level learning improvement to the following evidential premise: Its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns resources assigned to education purposes across the system. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Public spending by level and function within Part XII — Finance and distribution must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    A defensible statistic would be based on expenditure classified by level, recurrent or capital use and responsible body. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression. The estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. When several sources exist, consistency is evidence to consider, not proof that common error is absent. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Public spending by level and function within Part XII — Finance and distribution must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    The educational consequence in Public spending by level and function within Part XII — Finance and distribution gives practical meaning to credible evidence of system-level learning improvement: Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: budgets, commitments and actual expenditure distinguished. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Public spending by level and function within Part XII — Finance and distribution must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    Missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. An adequate equity account asks whether basic education services under fiscal pressure are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Public spending by level and function within Part XII — Finance and distribution must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    32

    Incidence of education spending

    Incidence of education spending should be approached as a defined measurement problem. The substantive interest is who benefits from publicly financed places and services, observed for a population and period that are stated before calculation. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Incidence of education spending within Part XII — Finance and distribution must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    The population addressed in Incidence of education spending within Part XII — Finance and distribution tests the adequacy of credible evidence of system-level learning improvement in this respect: These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is unit resources combined with participation across population groups. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Incidence of education spending within Part XII — Finance and distribution must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference. Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that benefit estimates not treated as household income. A responsible commentary distinguishes observation, calculation and interpretation. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Incidence of education spending within Part XII — Finance and distribution must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    For groups excluded before public spending can reach them, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community. Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Incidence of education spending within Part XII — Finance and distribution must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    33

    Protecting equity during fiscal constraint

    A decision concerning Protecting equity during fiscal constraint within Part XII — Finance and distribution should connect credible evidence of system-level learning improvement to the following evidential premise: The report should state the educational consequence before choosing a gap, ratio, threshold or rank. For protecting equity during fiscal constraint, the first requirement is conceptual clarity. The study seeks evidence on whether reductions or delays fall disproportionately on weaker services; it does not infer a learner's circumstances from a national or regional mean. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Protecting equity during fiscal constraint within Part XII — Finance and distribution must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is dated finance and service indicators read together. Numerator, denominator, reference date, unit and exclusions should appear together. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Protecting equity during fiscal constraint within Part XII — Finance and distribution must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: national totals tested against subnational allocation and household costs. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Protecting equity during fiscal constraint within Part XII — Finance and distribution must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    The educational consequence in Protecting equity during fiscal constraint within Part XII — Finance and distribution gives practical meaning to credible evidence of system-level learning improvement: Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include poor households and institutions with little financial reserve. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Protecting equity during fiscal constraint within Part XII — Finance and distribution must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    Part XIII

    Data sources and measurement error

    34

    Administrative records

    A decision concerning Administrative records within Part XIII — Data sources and measurement error should connect credible evidence of system-level learning improvement to the following evidential premise: A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of regular learner, staff, facility and finance information begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement. The measure should preserve both the observed level and the distribution relevant to the claim. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Administrative records within Part XIII — Data sources and measurement error must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use clear definitions, reporting coverage and revision history. This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined. School returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. Reconciliation should record what each source can and cannot represent. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Administrative records within Part XIII — Data sources and measurement error must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. Percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. The comparison should identify the reference category but avoid presenting it as a natural norm. Policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that non-reporting institutions kept visible in aggregates. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Administrative records within Part XIII — Data sources and measurement error must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    The review should record non-response, unknown status and excluded locations separately. Combining unknown observations with the majority group biases both estimates and obscures the weakness. Where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for small, private, non-formal and emergency providers. Their circumstances may alter access to enumeration, classification and the service being measured. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Administrative records within Part XIII — Data sources and measurement error must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    35

    Household surveys

    Comparability in Household surveys within Part XIII — Data sources and measurement error depends on an explicit judgement concerning credible evidence of system-level learning improvement: The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of household surveys lies in making unequal educational experience observable. Here, the relevant phenomenon is population-based evidence beyond enrolled learners, not the administrative convenience of the available categories. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Household surveys within Part XIII — Data sources and measurement error must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    The estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. When several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on probability samples, weights and field dates documented. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Household surveys within Part XIII — Data sources and measurement error must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: sampling and non-response uncertainty carried into group comparisons. Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Household surveys within Part XIII — Data sources and measurement error must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    An adequate equity account asks whether small minorities and mobile households are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. Missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Household surveys within Part XIII — Data sources and measurement error must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    36

    Censuses and population frames

    The population addressed in Censuses and population frames within Part XIII — Data sources and measurement error tests the adequacy of credible evidence of system-level learning improvement in this respect: Its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns broad population coverage and small-area denominators. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Censuses and population frames within Part XIII — Data sources and measurement error must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    Within Censuses and population frames within Part XIII — Data sources and measurement error, the authority should interpret credible evidence of system-level learning improvement against this condition: These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is enumeration date, usual residence and institutional coverage stated. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Censuses and population frames within Part XIII — Data sources and measurement error must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    It does not assign cause from a cross-sectional difference. Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that long intervals and under-enumeration acknowledged. A responsible commentary distinguishes observation, calculation and interpretation. It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Censuses and population frames within Part XIII — Data sources and measurement error must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. For homeless, displaced and geographically isolated people, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Censuses and population frames within Part XIII — Data sources and measurement error must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-20] [REF-21]

    Part XIV

    Disaggregation and intersection

    37

    Single-axis disaggregation

    Single-axis disaggregation should be approached as a defined measurement problem. The substantive interest is separate reporting by sex, wealth, residence or another characteristic, observed for a population and period that are stated before calculation. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Single-axis disaggregation within Part XIV — Disaggregation and intersection must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    A discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is levels, gaps and denominators shown for each category. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Single-axis disaggregation within Part XIV — Disaggregation and intersection must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: one axis not presented as a complete account of marginalisation. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Single-axis disaggregation within Part XIV — Disaggregation and intersection must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-20] [REF-21]

    Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include groups whose disadvantage lies on another unmeasured dimension. These populations may be missing not only from good outcomes but from the denominator itself. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Single-axis disaggregation within Part XIV — Disaggregation and intersection must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    38

    Intersecting categories

    The population addressed in Intersecting categories within Part XIV — Disaggregation and intersection tests the adequacy of credible evidence of system-level learning improvement in this respect: The report should state the educational consequence before choosing a gap, ratio, threshold or rank. For intersecting categories, the first requirement is conceptual clarity. The study seeks evidence on joint distributions such as sex by wealth and residence; it does not infer a learner's circumstances from a national or regional mean. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Intersecting categories within Part XIV — Disaggregation and intersection must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined. School returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. Reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use pre-specified combinations with sufficient observations. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Intersecting categories within Part XIV — Disaggregation and intersection must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-20] [REF-21]

    Policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that empty or unstable cells reported honestly. Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. Percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. The comparison should identify the reference category but avoid presenting it as a natural norm. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Intersecting categories within Part XIV — Disaggregation and intersection must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    Particular scrutiny is required for poor rural girls, disabled learners in remote areas and displaced minorities. Their circumstances may alter access to enumeration, classification and the service being measured. The review should record non-response, unknown status and excluded locations separately. Combining unknown observations with the majority group biases both estimates and obscures the weakness. Where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Intersecting categories within Part XIV — Disaggregation and intersection must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    39

    Small numbers, disclosure and reliability

    The population addressed in Small numbers, disclosure and reliability within Part XIV — Disaggregation and intersection tests the adequacy of credible evidence of system-level learning improvement in this respect: A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of useful detail without unreliable estimates or identification begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement. The measure should preserve both the observed level and the distribution relevant to the claim. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Small numbers, disclosure and reliability within Part XIV — Disaggregation and intersection must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-20] [REF-21]

    Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression. The estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. When several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on suppression, aggregation or qualitative evidence chosen proportionately. The definition should be fixed for the comparison at hand and deviations recorded. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Small numbers, disclosure and reliability within Part XIV — Disaggregation and intersection must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: confidentiality decisions separated from claims that no disparity exists. Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Small numbers, disclosure and reliability within Part XIV — Disaggregation and intersection must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. Missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. An adequate equity account asks whether small communities and learners with rare characteristics are represented at each stage: population frame, collection, valid response, classification, analysis and publication. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Small numbers, disclosure and reliability within Part XIV — Disaggregation and intersection must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    Part XV

    Comparison, uncertainty and change

    40

    Comparing unlike systems

    Learner protection in Comparing unlike systems within Part XV — Comparison, uncertainty and change makes the following aspect of credible evidence of system-level learning improvement material: The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of comparing unlike systems lies in making unequal educational experience observable. Here, the relevant phenomenon is cross-country patterns based on harmonised but bounded concepts, not the administrative convenience of the available categories. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Comparing unlike systems within Part XV — Comparison, uncertainty and change must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    Equity analysis in Comparing unlike systems within Part XV — Comparison, uncertainty and change qualifies any conclusion about credible evidence of system-level learning improvement: These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is metadata tests before numerical comparison. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Comparing unlike systems within Part XV — Comparison, uncertainty and change must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that differences in programme structure and classification remain visible. A responsible commentary distinguishes observation, calculation and interpretation. It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Comparing unlike systems within Part XV — Comparison, uncertainty and change must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community. Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. For countries with incomplete or rapidly changing systems, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Comparing unlike systems within Part XV — Comparison, uncertainty and change must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    41

    Sampling error and other uncertainty

    Within Sampling error and other uncertainty within Part XV — Comparison, uncertainty and change, the authority should interpret credible evidence of system-level learning improvement against this condition: Its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns the range of values reasonably compatible with the observations. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Sampling error and other uncertainty within Part XV — Comparison, uncertainty and change must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-07]

    The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is standard errors, design effects and data-quality qualifications. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Sampling error and other uncertainty within Part XV — Comparison, uncertainty and change must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    Interpretation follows this limitation: rank differences smaller than uncertainty not interpreted. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Sampling error and other uncertainty within Part XV — Comparison, uncertainty and change must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    The distributional review must deliberately include small disaggregated populations. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Sampling error and other uncertainty within Part XV — Comparison, uncertainty and change must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    Part XVI

    Responsible interpretation and action

    43

    Reading disparity without blaming learners

    Within Reading disparity without blaming learners within Part XVI — Responsible interpretation and action, the authority should interpret credible evidence of system-level learning improvement against this condition: The report should state the educational consequence before choosing a gap, ratio, threshold or rank. For reading disparity without blaming learners, the first requirement is conceptual clarity. The study seeks evidence on institutional and social conditions associated with unequal outcomes; it does not infer a learner's circumstances from a national or regional mean. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Reading disparity without blaming learners within Part XVI — Responsible interpretation and action must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-06]

    A defensible statistic would be based on descriptive findings separated from causal claims. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression. The estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. When several sources exist, consistency is evidence to consider, not proof that common error is absent. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Reading disparity without blaming learners within Part XVI — Responsible interpretation and action must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    Public accountability for Reading disparity without blaming learners within Part XVI — Responsible interpretation and action requires a reasoned finding about credible evidence of system-level learning improvement: Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: group identity never treated as a mechanism by itself. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Reading disparity without blaming learners within Part XVI — Responsible interpretation and action must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    Field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. Missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. An adequate equity account asks whether communities subject to stigma are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Reading disparity without blaming learners within Part XVI — Responsible interpretation and action must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    44

    Turning evidence into equitable policy

    Within Turning evidence into equitable policy within Part XVI — Responsible interpretation and action, the authority should interpret credible evidence of system-level learning improvement against this condition: A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of a decision rule linking disparity to service, finance or legal responsibility begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement. The measure should preserve both the observed level and the distribution relevant to the claim. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Turning evidence into equitable policy within Part XVI — Responsible interpretation and action must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    The responsible body for Turning evidence into equitable policy within Part XVI — Responsible interpretation and action must state how credible evidence of system-level learning improvement resolves this question: These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is baseline, intended reach, implementation evidence and review date. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Turning evidence into equitable policy within Part XVI — Responsible interpretation and action must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    Use of the indicator is bounded by the principle that targets accompanied by distributional safeguards. A responsible commentary distinguishes observation, calculation and interpretation. It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference. Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Turning evidence into equitable policy within Part XVI — Responsible interpretation and action must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    For learners farthest below a secured minimum, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community. Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Turning evidence into equitable policy within Part XVI — Responsible interpretation and action must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    45

    A public account beyond the average

    A public account beyond the average within Part XVI — Responsible interpretation and action places this proposition within the evidence required for credible evidence of system-level learning improvement: The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of a public account beyond the average lies in making unequal educational experience observable. Here, the relevant phenomenon is a concise national statement of level, distribution, missingness and remedy, not the administrative convenience of the available categories. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, A public account beyond the average within Part XVI — Responsible interpretation and action must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-12]

    Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is national totals presented beside selected group and place indicators. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, A public account beyond the average within Part XVI — Responsible interpretation and action must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: progress claims bounded by evidence coverage and unresolved gaps. The comparison should show the level for each group as well as any ratio or gap. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, A public account beyond the average within Part XVI — Responsible interpretation and action must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-16]

    Public accountability for A public account beyond the average within Part XVI — Responsible interpretation and action requires a reasoned finding about credible evidence of system-level learning improvement: Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include every learner otherwise hidden by a successful average. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, A public account beyond the average within Part XVI — Responsible interpretation and action must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    Part XVII

    Applied distributional analysis

    46

    Composite measures and the loss of meaning

    The national total supplies context, while the distribution tests whether the total is shared. Where the measure condenses several observations, its construction must remain open to reconstruction from the underlying values. Where it separates groups, classifications must be lawful, meaningful and sufficiently stable for the comparison. The analytical purpose is combining several dimensions into a summary measure. That purpose should be written before the calculation because method follows the intended inference. The evidence base comprises participation, completion, learning and school conditions, each of which describes a different feature of educational opportunity. A result is useful only if a reader can identify the represented population, the reference period and the educational consequence attached to movement. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Composite measures and the loss of meaning within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-03]

    Composite measures and the loss of meaning within Part XVII — Applied distributional analysis places this proposition within the evidence required for credible evidence of system-level learning improvement: Documentation should also state which observations are direct, which are estimated and which are unavailable. Missing values should remain missing unless an explicit estimation method and its effect are shown. The principal methodological questions concern weights, normalisation and substitution. Choices on these matters are not neutral presentation details. They determine how strongly one component, place or population can influence the conclusion. A defensible choice begins with the substantive education question and is then tested against alternative reasonable specifications. If the headline conclusion changes materially, the range of results should be reported. Sensitivity does not make the exercise useless; it prevents a conventional choice from appearing inevitable. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Composite measures and the loss of meaning within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-04]

    No result should be described as equitable solely because one relative measure improved. The main interpretive danger is a low value in one dimension being concealed by a high value in another. Avoiding it requires the underlying counts and distributions to remain visible beside any summary. Analysts should check whether an apparently favourable result arose through changed coverage, population movement, reclassification or concentration on cases nearest a threshold. A disparity can narrow because the better-served group deteriorates, while a composite score can improve although a protected minimum worsens. Direction must therefore be joined to adequacy. Comparison should also examine absolute numbers: a smaller rate in a growing population may still correspond to more learners without the secured opportunity. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Composite measures and the loss of meaning within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-05] [REF-06]

    Source quality should be considered dimension by dimension. Administrative returns may provide frequent local detail but exclude non-participants and non-reporting institutions. Surveys can represent households beyond formal education, subject to sample size, field access and response. Censuses can support small-area denominators at longer intervals, while enumeration rules and population movement affect coverage. A record of agreement should follow reconciliation of concepts rather than simple numerical proximity. A record of disagreement should identify plausible sources and the decision consequence. If the available evidence does not support a proposed level of disaggregation, the limitation should appear in the finding rather than in a remote methodological note. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Composite measures and the loss of meaning within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-12]

    Small cells require protection against disclosure and caution about statistical stability. These duties do not justify silence about a serious disparity. The public account may report a broader group, a range, a qualitative service failure or restricted findings, provided that it states what cannot safely be quantified and which body is responsible for better evidence. Equity review asks who can disappear during construction of the measure. Learners outside school, people in temporary settlements, small language communities, persons with disabilities and households beyond a survey frame can all be absent before analysis begins. Analysts should compare coverage against independent population and service information and preserve an unknown category where classification is incomplete. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Composite measures and the loss of meaning within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-16]

    Composite measures and the loss of meaning within Part XVII — Applied distributional analysis places this proposition within the evidence required for credible evidence of system-level learning improvement: Every response should name the expected population reach and the later observation that will test it. Otherwise a disparity can generate activity without demonstrating that conditions changed for the intended learners. The policy user is policy-makers deciding whether a summary aids or obscures resource allocation. Evidence should lead to a stated decision class: immediate removal of an access barrier, redistribution of staff or finance, further investigation, amendment of a classification, or evaluation of an existing measure. The certainty required depends upon the consequence. Credible indications of severe exclusion can justify protective action before exact prevalence is known, whereas durable allocation formulas require review as evidence improves. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Composite measures and the loss of meaning within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    The two forms answer different questions. Authorities should provide accessible explanations of definitions and invite correction where categories misdescribe experience. Feedback needs a recorded route into classification, service design or further enquiry. People should not be asked repeatedly for sensitive information when no competent body can act upon the answer. Participation by affected communities strengthens both interpretation and legitimacy. Local knowledge can identify seasonal movement, unsafe routes, hidden household costs, language use and service boundaries that national categories miss. Consultation should not be treated as statistical verification, nor should a survey estimate displace credible evidence of a local barrier. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Composite measures and the loss of meaning within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-21] [REF-23]

    A concise account can carry considerable density when every figure retains its population and consequence. The objective is not maximum numerical output. It is a trustworthy connection between unequal educational experience, public responsibility and corrective action. Final reporting should give a clear institutional judgement. It should state what the evidence shows, the strength and limits of that conclusion, the people and places most affected, the action within public authority and the date for review. Changes to definitions, boundaries or population estimates should appear at the point where a series changes. Earlier values should remain available so that revision is not mistaken for real progress. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Composite measures and the loss of meaning within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    47

    Decomposing an observed education gap

    Where the measure condenses several observations, its construction must remain open to reconstruction from the underlying values. Where it separates groups, classifications must be lawful, meaningful and sufficiently stable for the comparison. The analytical purpose is examining how a national disparity is distributed across places and population groups. That purpose should be written before the calculation because method follows the intended inference. The evidence base comprises within-group and between-group differences, each of which describes a different feature of educational opportunity. A result is useful only if a reader can identify the represented population, the reference period and the educational consequence attached to movement. The national total supplies context, while the distribution tests whether the total is shared. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Decomposing an observed education gap within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-04]

    For Decomposing an observed education gap within Part XVII — Applied distributional analysis, a responsible judgement on credible evidence of system-level learning improvement should address the following point: Documentation should also state which observations are direct, which are estimated and which are unavailable. Missing values should remain missing unless an explicit estimation method and its effect are shown. The principal methodological questions concern population shares, outcome levels and overlapping membership. Choices on these matters are not neutral presentation details. They determine how strongly one component, place or population can influence the conclusion. A defensible choice begins with the substantive education question and is then tested against alternative reasonable specifications. If the headline conclusion changes materially, the range of results should be reported. Sensitivity does not make the exercise useless; it prevents a conventional choice from appearing inevitable. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Decomposing an observed education gap within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-05] [REF-06]

    The main interpretive danger is treating a descriptive decomposition as proof of cause. Avoiding it requires the underlying counts and distributions to remain visible beside any summary. Analysts should check whether an apparently favourable result arose through changed coverage, population movement, reclassification or concentration on cases nearest a threshold. A disparity can narrow because the better-served group deteriorates, while a composite score can improve although a protected minimum worsens. Direction must therefore be joined to adequacy. Comparison should also examine absolute numbers: a smaller rate in a growing population may still correspond to more learners without the secured opportunity. No result should be described as equitable solely because one relative measure improved. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Decomposing an observed education gap within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-12]

    For Decomposing an observed education gap within Part XVII — Applied distributional analysis, a responsible judgement on credible evidence of system-level learning improvement should address the following point: Every response should name the expected population reach and the later observation that will test it. Otherwise a disparity can generate activity without demonstrating that conditions changed for the intended learners. The policy user is authorities locating where further enquiry and action are warranted. Evidence should lead to a stated decision class: immediate removal of an access barrier, redistribution of staff or finance, further investigation, amendment of a classification, or evaluation of an existing measure. The certainty required depends upon the consequence. Credible indications of severe exclusion can justify protective action before exact prevalence is known, whereas durable allocation formulas require review as evidence improves. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Decomposing an observed education gap within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-21] [REF-23]

    48

    Setting distribution-sensitive targets

    The evidence base comprises the national level, the least-served group and the lower tail of the distribution, each of which describes a different feature of educational opportunity. A result is useful only if a reader can identify the represented population, the reference period and the educational consequence attached to movement. The national total supplies context, while the distribution tests whether the total is shared. Where the measure condenses several observations, its construction must remain open to reconstruction from the underlying values. Where it separates groups, classifications must be lawful, meaningful and sufficiently stable for the comparison. The analytical purpose is expressing progress as improvement in secured minimums and unjustified gaps. That purpose should be written before the calculation because method follows the intended inference. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Setting distribution-sensitive targets within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-05] [REF-06]

    They determine how strongly one component, place or population can influence the conclusion. A defensible choice begins with the substantive education question and is then tested against alternative reasonable specifications. If the headline conclusion changes materially, the range of results should be reported. Sensitivity does not make the exercise useless; it prevents a conventional choice from appearing inevitable. Documentation should also state which observations are direct, which are estimated and which are unavailable. Missing values should remain missing unless an explicit estimation method and its effect are shown. The principal methodological questions concern baseline stability, ambition and safeguards against exclusion. Choices on these matters are not neutral presentation details. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Setting distribution-sensitive targets within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-12]

    Direction must therefore be joined to adequacy. Comparison should also examine absolute numbers: a smaller rate in a growing population may still correspond to more learners without the secured opportunity. No result should be described as equitable solely because one relative measure improved. The main interpretive danger is meeting a mean target while abandoning those farthest behind. Avoiding it requires the underlying counts and distributions to remain visible beside any summary. Analysts should check whether an apparently favourable result arose through changed coverage, population movement, reclassification or concentration on cases nearest a threshold. A disparity can narrow because the better-served group deteriorates, while a composite score can improve although a protected minimum worsens. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Setting distribution-sensitive targets within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-16]

    The certainty required depends upon the consequence. Credible indications of severe exclusion can justify protective action before exact prevalence is known, whereas durable allocation formulas require review as evidence improves. Every response should name the expected population reach and the later observation that will test it. Otherwise a disparity can generate activity without demonstrating that conditions changed for the intended learners. The policy user is governments linking national commitments to subnational delivery. Evidence should lead to a stated decision class: immediate removal of an access barrier, redistribution of staff or finance, further investigation, amendment of a classification, or evaluation of an existing measure. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Setting distribution-sensitive targets within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    49

    Linking learners to service geography

    Where it separates groups, classifications must be lawful, meaningful and sufficiently stable for the comparison. The analytical purpose is relating participation and learning to the location and capacity of education services. That purpose should be written before the calculation because method follows the intended inference. The evidence base comprises settlement populations, travel conditions, schools, teachers and programme levels, each of which describes a different feature of educational opportunity. A result is useful only if a reader can identify the represented population, the reference period and the educational consequence attached to movement. The national total supplies context, while the distribution tests whether the total is shared. Where the measure condenses several observations, its construction must remain open to reconstruction from the underlying values. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Linking learners to service geography within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-12]

    Evidence from Linking learners to service geography within Part XVII — Applied distributional analysis bears on credible evidence of system-level learning improvement through this institutional requirement: Documentation should also state which observations are direct, which are estimated and which are unavailable. Missing values should remain missing unless an explicit estimation method and its effect are shown. The principal methodological questions concern geographical scale, boundary effects and facility catchments. Choices on these matters are not neutral presentation details. They determine how strongly one component, place or population can influence the conclusion. A defensible choice begins with the substantive education question and is then tested against alternative reasonable specifications. If the headline conclusion changes materially, the range of results should be reported. Sensitivity does not make the exercise useless; it prevents a conventional choice from appearing inevitable. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Linking learners to service geography within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-16]

    Avoiding it requires the underlying counts and distributions to remain visible beside any summary. Analysts should check whether an apparently favourable result arose through changed coverage, population movement, reclassification or concentration on cases nearest a threshold. A disparity can narrow because the better-served group deteriorates, while a composite score can improve although a protected minimum worsens. Direction must therefore be joined to adequacy. Comparison should also examine absolute numbers: a smaller rate in a growing population may still correspond to more learners without the secured opportunity. No result should be described as equitable solely because one relative measure improved. The main interpretive danger is assuming the nearest mapped institution is accessible or appropriate. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Linking learners to service geography within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    The policy user is planners choosing sites, transport support and teacher deployment. Evidence should lead to a stated decision class: immediate removal of an access barrier, redistribution of staff or finance, further investigation, amendment of a classification, or evaluation of an existing measure. The certainty required depends upon the consequence. Credible indications of severe exclusion can justify protective action before exact prevalence is known, whereas durable allocation formulas require review as evidence improves. Every response should name the expected population reach and the later observation that will test it. Otherwise a disparity can generate activity without demonstrating that conditions changed for the intended learners. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Linking learners to service geography within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-03]

    50

    Reconciling conflicting sources

    A result is useful only if a reader can identify the represented population, the reference period and the educational consequence attached to movement. The national total supplies context, while the distribution tests whether the total is shared. Where the measure condenses several observations, its construction must remain open to reconstruction from the underlying values. Where it separates groups, classifications must be lawful, meaningful and sufficiently stable for the comparison. The analytical purpose is interpreting differences between administrative, survey and census estimates. That purpose should be written before the calculation because method follows the intended inference. The evidence base comprises coverage, timing, concepts and reporting incentives, each of which describes a different feature of educational opportunity. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Reconciling conflicting sources within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-16]

    They determine how strongly one component, place or population can influence the conclusion. A defensible choice begins with the substantive education question and is then tested against alternative reasonable specifications. If the headline conclusion changes materially, the range of results should be reported. Sensitivity does not make the exercise useless; it prevents a conventional choice from appearing inevitable. Documentation should also state which observations are direct, which are estimated and which are unavailable. Missing values should remain missing unless an explicit estimation method and its effect are shown. The principal methodological questions concern a documented comparison of population and variable definitions. Choices on these matters are not neutral presentation details. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Reconciling conflicting sources within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    Comparison should also examine absolute numbers: a smaller rate in a growing population may still correspond to more learners without the secured opportunity. No result should be described as equitable solely because one relative measure improved. The main interpretive danger is averaging incompatible estimates into an apparently precise figure. Avoiding it requires the underlying counts and distributions to remain visible beside any summary. Analysts should check whether an apparently favourable result arose through changed coverage, population movement, reclassification or concentration on cases nearest a threshold. A disparity can narrow because the better-served group deteriorates, while a composite score can improve although a protected minimum worsens. Direction must therefore be joined to adequacy. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Reconciling conflicting sources within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-21] [REF-23]

    Evidence from Reconciling conflicting sources within Part XVII — Applied distributional analysis bears on credible evidence of system-level learning improvement through this institutional requirement: Every response should name the expected population reach and the later observation that will test it. Otherwise a disparity can generate activity without demonstrating that conditions changed for the intended learners. The policy user is statistical authorities issuing one bounded account with visible uncertainty. Evidence should lead to a stated decision class: immediate removal of an access barrier, redistribution of staff or finance, further investigation, amendment of a classification, or evaluation of an existing measure. The certainty required depends upon the consequence. Credible indications of severe exclusion can justify protective action before exact prevalence is known, whereas durable allocation formulas require review as evidence improves. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Reconciling conflicting sources within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-04]

    51

    Monitoring marginalisation during severe disruption

    The analytical purpose is maintaining useful distributional evidence when populations and services move rapidly. That purpose should be written before the calculation because method follows the intended inference. The evidence base comprises rapid counts, restored administrative returns and household evidence, each of which describes a different feature of educational opportunity. A result is useful only if a reader can identify the represented population, the reference period and the educational consequence attached to movement. The national total supplies context, while the distribution tests whether the total is shared. Where the measure condenses several observations, its construction must remain open to reconstruction from the underlying values. Where it separates groups, classifications must be lawful, meaningful and sufficiently stable for the comparison. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Monitoring marginalisation during severe disruption within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    The educational consequence in Monitoring marginalisation during severe disruption within Part XVII — Applied distributional analysis gives practical meaning to credible evidence of system-level learning improvement: Documentation should also state which observations are direct, which are estimated and which are unavailable. Missing values should remain missing unless an explicit estimation method and its effect are shown. The principal methodological questions concern dated estimates, revision practice and minimal essential classifications. Choices on these matters are not neutral presentation details. They determine how strongly one component, place or population can influence the conclusion. A defensible choice begins with the substantive education question and is then tested against alternative reasonable specifications. If the headline conclusion changes materially, the range of results should be reported. Sensitivity does not make the exercise useless; it prevents a conventional choice from appearing inevitable. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Monitoring marginalisation during severe disruption within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-21] [REF-23]

    Analysts should check whether an apparently favourable result arose through changed coverage, population movement, reclassification or concentration on cases nearest a threshold. A disparity can narrow because the better-served group deteriorates, while a composite score can improve although a protected minimum worsens. Direction must therefore be joined to adequacy. Comparison should also examine absolute numbers: a smaller rate in a growing population may still correspond to more learners without the secured opportunity. No result should be described as equitable solely because one relative measure improved. The main interpretive danger is using an unstable emergency denominator to assert durable improvement. Avoiding it requires the underlying counts and distributions to remain visible beside any summary. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Monitoring marginalisation during severe disruption within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    The certainty required depends upon the consequence. Credible indications of severe exclusion can justify protective action before exact prevalence is known, whereas durable allocation formulas require review as evidence improves. Every response should name the expected population reach and the later observation that will test it. Otherwise a disparity can generate activity without demonstrating that conditions changed for the intended learners. The policy user is authorities protecting access while rebuilding regular statistics. Evidence should lead to a stated decision class: immediate removal of an access barrier, redistribution of staff or finance, further investigation, amendment of a classification, or evaluation of an existing measure. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Monitoring marginalisation during severe disruption within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-05] [REF-06]

    52

    Communicating uncertainty without losing urgency

    The national total supplies context, while the distribution tests whether the total is shared. Where the measure condenses several observations, its construction must remain open to reconstruction from the underlying values. Where it separates groups, classifications must be lawful, meaningful and sufficiently stable for the comparison. The analytical purpose is explaining what is known strongly enough to justify action and what remains unresolved. That purpose should be written before the calculation because method follows the intended inference. The evidence base comprises point estimates, ranges, quality statements and missing populations, each of which describes a different feature of educational opportunity. A result is useful only if a reader can identify the represented population, the reference period and the educational consequence attached to movement. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Communicating uncertainty without losing urgency within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-21] [REF-23]

    They determine how strongly one component, place or population can influence the conclusion. A defensible choice begins with the substantive education question and is then tested against alternative reasonable specifications. If the headline conclusion changes materially, the range of results should be reported. Sensitivity does not make the exercise useless; it prevents a conventional choice from appearing inevitable. Documentation should also state which observations are direct, which are estimated and which are unavailable. Missing values should remain missing unless an explicit estimation method and its effect are shown. The principal methodological questions concern plain institutional language joined to exact metadata. Choices on these matters are not neutral presentation details. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Communicating uncertainty without losing urgency within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    No result should be described as equitable solely because one relative measure improved. The main interpretive danger is presenting caution as a reason for inaction or urgency as a reason for overstatement. Avoiding it requires the underlying counts and distributions to remain visible beside any summary. Analysts should check whether an apparently favourable result arose through changed coverage, population movement, reclassification or concentration on cases nearest a threshold. A disparity can narrow because the better-served group deteriorates, while a composite score can improve although a protected minimum worsens. Direction must therefore be joined to adequacy. Comparison should also examine absolute numbers: a smaller rate in a growing population may still correspond to more learners without the secured opportunity. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Communicating uncertainty without losing urgency within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-03]

    The policy user is the public, affected communities and responsible decision-makers. Evidence should lead to a stated decision class: immediate removal of an access barrier, redistribution of staff or finance, further investigation, amendment of a classification, or evaluation of an existing measure. The certainty required depends upon the consequence. Credible indications of severe exclusion can justify protective action before exact prevalence is known, whereas durable allocation formulas require review as evidence improves. Every response should name the expected population reach and the later observation that will test it. Otherwise a disparity can generate activity without demonstrating that conditions changed for the intended learners. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Communicating uncertainty without losing urgency within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-12]

    53

    A national marginalisation profile

    That purpose should be written before the calculation because method follows the intended inference. The evidence base comprises population, access, progression, learning, conditions, finance and unresolved evidence gaps, each of which describes a different feature of educational opportunity. A result is useful only if a reader can identify the represented population, the reference period and the educational consequence attached to movement. The national total supplies context, while the distribution tests whether the total is shared. Where the measure condenses several observations, its construction must remain open to reconstruction from the underlying values. Where it separates groups, classifications must be lawful, meaningful and sufficiently stable for the comparison. The analytical purpose is assembling a concise recurring account beyond the national average. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, A national marginalisation profile within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    Public accountability for A national marginalisation profile within Part XVII — Applied distributional analysis requires a reasoned finding about credible evidence of system-level learning improvement: Documentation should also state which observations are direct, which are estimated and which are unavailable. Missing values should remain missing unless an explicit estimation method and its effect are shown. The principal methodological questions concern a stable core with context-specific distributions. Choices on these matters are not neutral presentation details. They determine how strongly one component, place or population can influence the conclusion. A defensible choice begins with the substantive education question and is then tested against alternative reasonable specifications. If the headline conclusion changes materially, the range of results should be reported. Sensitivity does not make the exercise useless; it prevents a conventional choice from appearing inevitable. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, A national marginalisation profile within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-03]

    A disparity can narrow because the better-served group deteriorates, while a composite score can improve although a protected minimum worsens. Direction must therefore be joined to adequacy. Comparison should also examine absolute numbers: a smaller rate in a growing population may still correspond to more learners without the secured opportunity. No result should be described as equitable solely because one relative measure improved. The main interpretive danger is creating an encyclopaedia of indicators without decision priority. Avoiding it requires the underlying counts and distributions to remain visible beside any summary. Analysts should check whether an apparently favourable result arose through changed coverage, population movement, reclassification or concentration on cases nearest a threshold. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, A national marginalisation profile within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-04]

    The educational consequence in A national marginalisation profile within Part XVII — Applied distributional analysis gives practical meaning to credible evidence of system-level learning improvement: Every response should name the expected population reach and the later observation that will test it. Otherwise a disparity can generate activity without demonstrating that conditions changed for the intended learners. The policy user is parliament, ministries, local authorities and communities reviewing educational equity. Evidence should lead to a stated decision class: immediate removal of an access barrier, redistribution of staff or finance, further investigation, amendment of a classification, or evaluation of an existing measure. The certainty required depends upon the consequence. Credible indications of severe exclusion can justify protective action before exact prevalence is known, whereas durable allocation formulas require review as evidence improves. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, A national marginalisation profile within Part XVII — Applied distributional analysis must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-16]

    Part XVIII

    Extended disparity interpretation

    54

    Minimum comparison threshold

    Opportunity to learn, participation and exclusions remain material. A learning-disparity comparison requires a declared construct, represented population, assessment conditions, scale, uncertainty and distribution. National means should be accompanied by lower-tail, threshold and group evidence. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Minimum comparison threshold within Part XVIII — Extended disparity interpretation must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    55

    Within-system interpretation

    Within-system gaps should preserve place, group and institutional context without assigning cause from identity. Counts, levels, absolute gaps and ratios answer different questions. Missing learners and non-participating schools must remain visible. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Within-system interpretation within Part XVIII — Extended disparity interpretation must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-06]

    56

    Between-system interpretation

    Rank differences smaller than uncertainty should not support categorical conclusions. Harmonisation does not remove substantive system difference. Between-system comparison requires metadata tests for curriculum, age, language, sampling and assessment. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Between-system interpretation within Part XVIII — Extended disparity interpretation must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    Part XIX

    Extended analysis of learning disparities

    57

    Assessment participation and the represented learning population

    A learning distribution is defined partly by who participates in the assessment. The target population, eligible population, sampled population and assessed population should be reported separately. Learners can be absent because of illness, displacement, school non-attendance, language barriers, disability, conflict, administrative exclusion or ordinary sampling loss. These routes have different meanings. A result calculated only for participating learners may describe their performance accurately while failing to describe the educational system's full learner population. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Assessment participation and the represented learning population within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-02] [REF-03]

    School exclusion and within-school absence need separate observation where the design permits. If institutions outside the frame differ systematically from included schools, a high learner response rate within sampled schools cannot repair the coverage limitation. Similarly, replacement of inaccessible schools can preserve sample size while changing the population represented. Reports should describe replacement rules and show any material effect on geography or institutional type. Participation should therefore accompany every reported mean, proficiency share or percentile. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Assessment participation and the represented learning population within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    Sensitivity analysis can examine plausible bounds or compare known characteristics of participants and non-participants. Where a numerical adjustment is not defensible, the limitation remains a substantive finding. The absence of learning evidence for a population requiring education attention should not be interpreted as evidence of no disparity. Non-participation should not be assigned a score. Coding an absent learner as below threshold invents performance evidence; removing every absence without comment creates a different bias. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Assessment participation and the represented learning population within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    The report should explain that judgement rather than treat all adaptations as either incomparable or automatically identical. Accommodation and language arrangements affect participation and result validity. An assessment may permit attendance yet fail to elicit the intended construct if the format, communication or response mode is inaccessible. Exemption practices should be reported by reason and learner group. Where an adapted form changes the construct materially, separate interpretation may be necessary. Where it removes an irrelevant barrier, results may belong in the common distribution. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Assessment participation and the represented learning population within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-05] [REF-10] [REF-12]

    A national learning claim should be no broader than that coverage. Public interpretation should connect participation to policy. A low response among remote schools may require field and service improvement; absence among out-of-school children may require a population-based study and re-entry action; assessment exclusion may require accessible design. These are not corrections to be made solely through statistical weighting. The responsible education body should identify which population remains unrepresented and when better evidence will be available. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Assessment participation and the represented learning population within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    58

    Scale, threshold and distribution comparability

    A numerical score has meaning through the tasks, response model, scoring rules and population for which interpretation has been supported. Equal numerical differences should not be assumed to represent equal educational differences unless the scale warrants that inference. A threshold adds a substantive judgement about the knowledge or capability learners should demonstrate; it should not be selected merely because it divides the sample conveniently. Comparisons require a clear statement of what the assessment scale represents. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Scale, threshold and distribution comparability within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    The same mean can accompany a compressed distribution, a wide lower tail or polarisation. Reports should therefore consider percentiles, threshold shares and dispersion where technically sound. The lower part of the distribution is especially important for minimum learning opportunity, while the upper part may reveal whether expansion has altered advanced performance. These summaries should remain tied to uncertainty and assessment coverage. Means provide one description of the centre and can conceal change elsewhere. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Scale, threshold and distribution comparability within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    Descriptions such as basic, adequate or advanced should be treated as definitions under the assessment, not universal attributes of a learner or education system. Threshold comparisons need stable standard-setting and clear labels. A change in the percentage above a threshold may reflect learning, scale revision, task composition or population change. If standards are reset, the break should be visible at the year of change and parallel results reported where possible. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Scale, threshold and distribution comparability within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    Some comparisons may remain defensible at a broad domain while narrower subscales do not. The permissible inference should be stated accordingly. Cross-language and cross-cultural comparability require evidence, not an assumption that translation has preserved difficulty and meaning. Task familiarity, curriculum exposure and response conventions can affect results. Review should examine translation, adaptation, differential item behaviour and opportunity to learn, while avoiding the claim that every detected difference invalidates the whole assessment. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Scale, threshold and distribution comparability within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    League order does not answer those questions. A rank is usually less informative than the estimated difference, uncertainty and distribution. Small rank movement can follow changes in participating systems or sampling variation. Public reporting should avoid categorical language when intervals overlap or scale linkage is weak. The useful question is whether the evidence indicates a material disparity requiring enquiry, what population is affected and which educational condition might be changed. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Scale, threshold and distribution comparability within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    59

    Opportunity to learn and interpretation of achievement gaps

    Opportunity does not determine performance completely, and its measurement is imperfect. It nevertheless prevents a learning gap from being attributed solely to learners or households when institutions supplied different educational conditions. Achievement evidence should be interpreted with opportunity to learn. This includes curriculum entitlement, content actually taught, instructional time, teacher availability, language, materials, attendance and access to support. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Opportunity to learn and interpretation of achievement gaps within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-13] [REF-14]

    Teacher reports, schedules, classroom observation and learner work can add evidence, each with limitations. Teacher self-report may be influenced by recall or expectations; observation covers a short period; work samples are selective. Agreement across sources strengthens interpretation when dates and populations align. Contradictions can identify local variation or weak measurement and should not be resolved by choosing the account most favourable to the system. The official curriculum establishes intended opportunity, not delivered instruction. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Opportunity to learn and interpretation of achievement gaps within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    Total hours also conceal subject allocation and teaching quality. A learner receiving more hours of poorly organised instruction does not necessarily have greater opportunity in the intended domain. Time is therefore one component of the explanatory evidence, not a conversion factor for predicted score. Instructional time should distinguish scheduled, delivered and attended time. Closures, teacher absence, shortened shifts and late entry can reduce delivered or usable time. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Opportunity to learn and interpretation of achievement gaps within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    Facilities may exist but be inaccessible or unsafe. These conditions should be examined at the level where the learning evidence was collected. National resource averages can obscure concentration of weak provision among the same learners whose scores form the lower tail. Resource indicators should be connected to use. Textbooks delivered to a school may not be available in the relevant language, grade or classroom. Teacher qualifications recorded administratively may not match subject assignment. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Opportunity to learn and interpretation of achievement gaps within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    The original disparity and the remedial conditions should remain visible rather than being erased by a new cohort average. Policy conclusions should avoid treating opportunity indicators as excuses for low expectations. Their purpose is to identify conditions within public responsibility and to design support. Where learners received less curriculum exposure, a response may include additional teaching, staff deployment, accessible materials or revised pacing. Later assessment should test whether opportunity and learning changed. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Opportunity to learn and interpretation of achievement gaps within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    60

    Decomposing disparities within education systems

    A large within-school component can coexist with institutional inequality and should not be read as evidence that schools are irrelevant. A national disparity can reflect differences between regions, schools, classrooms and learners. Decomposition can describe where variation is concentrated, provided it is not treated as proof of cause. A large between-school component may indicate segregation, resource distribution or residential pattern; it does not identify which mechanism operates. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Decomposing disparities within education systems within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-09] [REF-14]

    Learners are commonly nested within classes and schools, while policies may operate through districts or providers. Standard errors and models should respect clustering. Small schools and sparsely populated areas may require special treatment, but removal changes the population represented. Reports should state exclusions and avoid presenting a modelled residual as direct observation. The level of analysis should correspond to the sampling and decision structure. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Decomposing disparities within education systems within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    A raw school mean combines prior opportunity, intake, mobility, attendance and current teaching. Adjusted measures can answer bounded questions but depend on variables and assumptions. They should not replace the unadjusted learner outcome or become a definitive quality rank. Adjustment for a condition influenced by the school can also remove part of the very effect under review. The purpose and causal assumptions need explanation. Group composition affects school comparisons. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Decomposing disparities within education systems within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    Policy priority should consider educational severity, population, rights and feasibility rather than statistical contribution alone. Maps and rankings, where used elsewhere, should not expose small communities or imply a boundary creates the disparity. Geographical decomposition should preserve absolute numbers and service context. A small district with a severe gap may need urgent support even though it contributes little to national variance. A populous area with a modest gap may represent many learners. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Decomposing disparities within education systems within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    Follow-up should state the condition examined, action taken and later learning evidence. The appropriate outcome is a decision agenda. Between-region evidence may lead to allocation review; between-school evidence may lead to staffing, admissions or support enquiry; within-school evidence may lead to classroom, language or accessibility review. Each hypothesis requires additional evidence. The decomposition locates questions; it does not authorise blame. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Decomposing disparities within education systems within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    61

    Bounded conclusions between education systems

    Between-system comparison serves public learning when it identifies patterns, plausible questions and alternative institutional arrangements. It becomes misleading when harmonised labels conceal different programme structures, ages, curricula, languages, participation or assessment conditions. Metadata review should precede numerical comparison. Where a material difference cannot be reconciled, the systems may still be described separately without a common rank. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Bounded conclusions between education systems within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03] [REF-16]

    Within-system distributions can overlap substantially even where means differ. Group composition and population coverage matter. An apparent national advantage may not extend to poor, rural, minority-language or disabled learners. Reports should place distributional evidence beside the system result and avoid using nationality as an explanation. Country and system averages should not be interpreted as attributes of every school or learner. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Bounded conclusions between education systems within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    A linked scale can support trend if common items or other methods preserve meaning and security, but linkage error should accompany the estimate. A later higher score should not automatically be described as system improvement where the represented population changed materially. Temporal comparison requires stable linkage. Changes in curriculum, assessment mode, participation, sampling frame or system boundaries can create discontinuity. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Bounded conclusions between education systems within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    Pilots should state the mechanism and review evidence. Adoption based solely on rank proximity or reputation substitutes imitation for analysis. Policy borrowing should attend to authority, capacity, sequence and context. A practice associated with high performance elsewhere may depend upon teacher preparation, finance, curriculum coherence or social conditions absent in the receiving system. The comparison can identify an option, not guarantee its effect. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Bounded conclusions between education systems within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    Interpretation identifies patterns and limitations. Policy consideration proposes further enquiry or action under national authority. Causal judgement requires additional design. This separation permits strong public concern about a learning disparity without false certainty about its source and protects education systems from both complacency and unsupported prescription. The final international statement should distinguish three levels. Recorded fact describes the observed results under stated methods. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Bounded conclusions between education systems within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    62

    Uncertainty, materiality and the duty to respond

    A larger sample may reduce sampling error; it does not correct systematic exclusion or an invalid construct. More decimal places cannot repair weak coverage. The public report should identify which uncertainty could alter the decision and which does not affect the direction of urgent protection. Uncertainty should qualify a disparity claim without neutralising it. Sampling error, non-response, scale linkage, classification and model choice each affect the range of defensible conclusions. They should be described separately because they support different remedies. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Uncertainty, materiality and the duty to respond within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03]

    A small estimated difference may be precise but have limited practical consequence, while an uncertain large gap affecting a protected minimum may require immediate enquiry. Materiality should consider the knowledge or capability involved, the number of learners, distribution, duration and consequences for later progression. The threshold for action also depends upon reversibility. Additional diagnostic support can be introduced and reviewed more readily than a high-stakes classification of schools or learners. Statistical significance should not substitute for educational materiality. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Uncertainty, materiality and the duty to respond within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    Every response should name the competent body, intended population, resources and review date. A remedy should match the evidence. Where participation is selective, improve coverage and examine barriers. Where opportunity to learn differs, address time, teachers, curriculum, language, accessibility or materials. Where scale comparability is weak, improve the assessment before publishing ranks. Where a group disparity persists under several definitions, investigate institutional mechanisms without attributing cause to identity. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Uncertainty, materiality and the duty to respond within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    A national evidence system demonstrates strength when it can acknowledge uncertainty, improve measurement and amend policy. The final test is whether learners receive better educational opportunity and whether remaining disparities continue to be visible rather than whether one annual figure becomes more favourable. Later evidence should be capable of changing the conclusion. Reports should preserve the original estimate, method and limitation, then state why revision occurred. Silent replacement makes apparent improvement impossible to distinguish from correction. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Uncertainty, materiality and the duty to respond within Part XIX — Extended analysis of learning disparities must establish whether observed change is comparable, distributed and plausibly connected to action.

    Part XX

    Targeted improvement planning

    63

    Defining a persistent learning gap

    It should be stated with the learner population, educational domain, geography and evidence period. A national average is insufficient when the plan addresses a local or group disparity. The baseline should preserve the underlying distribution and absolute numbers. If the assessment population excludes learners most at risk, the gap is not adequately defined. The plan should state which observation is direct, which is estimated and which remains unknown. The improvement question concerns a sustained disparity in a declared learning domain, population and period. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Defining a persistent learning gap within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-02]

    Achievement differences may be associated with poverty, language, disability or residence, but those characteristics are not instructional mechanisms. Authorities should examine curriculum, teacher availability, time, attendance, materials, assessment access and learner support. Alternative explanations should remain open until evidence discriminates among them. This discipline prevents a targeted plan from attaching deficit to learners instead of changing institutions. The principal error is a fluctuating score or one cohort difference being treated as proof of persistence. Avoiding it requires evidence on both learning and the opportunity supplied. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Defining a persistent learning gap within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-03] [REF-06]

    Observation and work samples can explain classroom conditions without estimating national prevalence. Learner and teacher accounts can identify barriers. Agreement adds confidence after dates and definitions align; contradiction should guide further enquiry rather than selective reporting. The required evidence includes baseline, assessment coverage, uncertainty, distribution and opportunity to learn. Each source should be used within scope. Administrative records can describe staffing and participation while omitting non-enrolled learners. Assessment provides bounded learning evidence subject to coverage and validity. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Defining a persistent learning gap within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    Local adaptation should remain possible within a common substantive condition. Any departure should be recorded with its reason and expected learner consequence. The governing action is to select a gap that is educationally material and within public influence. Implementation should identify who acts, with what authority, resources and deadline. Dependencies should be sequenced. Teacher guidance without planning time, materials without accessible use, or tutoring without safe attendance cannot deliver the expected mechanism. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Defining a persistent learning gap within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    Targets should therefore include the least-served position and protection against exclusion. Small groups require confidentiality and careful precision, not disappearance from review. Distributional review should ask who is eligible, offered support, participates, receives the intended intensity and demonstrates a later response. Gender, household resources, disability, language, residence and prior opportunity may intersect. A plan can improve its average by reaching learners closest to a threshold while leaving those farthest behind. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Defining a persistent learning gap within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    Professional capability is central. Teachers need subject knowledge, worked examples, diagnostic interpretation and protected time to collaborate. Moderation should examine evidence and reasoning rather than force identical decisions for unlike cases. Staff workload and turnover should be monitored. A plan relying on a few exceptional individuals is not institutionally secure. Leadership should route obstacles to bodies able to change staffing, finance, curriculum or assessment rather than leaving every correction to the classroom. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Defining a persistent learning gap within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    It should report limitations, adverse effects and unresolved learners. A favourable later result may reflect population or assessment change and should be tested against the baseline metadata. Where the intervention is ineffective, adaptation or cessation is responsible improvement. Where benefit depends on temporary support, institutionalisation requires recurrent finance and ordinary ownership. Completion is demonstrated by stronger learning opportunity and a functioning correction route, not by the end of a project. The public account should distinguish authorisation, delivery, use and learning consequence. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Defining a persistent learning gap within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-02]

    64

    From diagnosis to an intervention hypothesis

    A national average is insufficient when the plan addresses a local or group disparity. The baseline should preserve the underlying distribution and absolute numbers. If the assessment population excludes learners most at risk, the gap is not adequately defined. The plan should state which observation is direct, which is estimated and which remains unknown. The improvement question concerns an explicit account of the condition expected to change learning. It should be stated with the learner population, educational domain, geography and evidence period. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, From diagnosis to an intervention hypothesis within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-03] [REF-06]

    Authorities should examine curriculum, teacher availability, time, attendance, materials, assessment access and learner support. Alternative explanations should remain open until evidence discriminates among them. This discipline prevents a targeted plan from attaching deficit to learners instead of changing institutions. The principal error is group identity or low performance being mistaken for a causal explanation. Avoiding it requires evidence on both learning and the opportunity supplied. Achievement differences may be associated with poverty, language, disability or residence, but those characteristics are not instructional mechanisms. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, From diagnosis to an intervention hypothesis within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    Observation and work samples can explain classroom conditions without estimating national prevalence. Learner and teacher accounts can identify barriers. Agreement adds confidence after dates and definitions align; contradiction should guide further enquiry rather than selective reporting. The required evidence includes curriculum exposure, teaching practice, language, time, materials and support. Each source should be used within scope. Administrative records can describe staffing and participation while omitting non-enrolled learners. Assessment provides bounded learning evidence subject to coverage and validity. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, From diagnosis to an intervention hypothesis within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    Any departure should be recorded with its reason and expected learner consequence. The governing action is to state the mechanism and plausible alternatives before choosing activity. Implementation should identify who acts, with what authority, resources and deadline. Dependencies should be sequenced. Teacher guidance without planning time, materials without accessible use, or tutoring without safe attendance cannot deliver the expected mechanism. Local adaptation should remain possible within a common substantive condition. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, From diagnosis to an intervention hypothesis within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    65

    Designing the targeted plan

    The improvement question concerns a bounded sequence connecting resources and actions to learner-facing change. It should be stated with the learner population, educational domain, geography and evidence period. A national average is insufficient when the plan addresses a local or group disparity. The baseline should preserve the underlying distribution and absolute numbers. If the assessment population excludes learners most at risk, the gap is not adequately defined. The plan should state which observation is direct, which is estimated and which remains unknown. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Designing the targeted plan within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    Avoiding it requires evidence on both learning and the opportunity supplied. Achievement differences may be associated with poverty, language, disability or residence, but those characteristics are not instructional mechanisms. Authorities should examine curriculum, teacher availability, time, attendance, materials, assessment access and learner support. Alternative explanations should remain open until evidence discriminates among them. This discipline prevents a targeted plan from attaching deficit to learners instead of changing institutions. The principal error is a list of activities replacing a coherent implementation logic. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Designing the targeted plan within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    Each source should be used within scope. Administrative records can describe staffing and participation while omitting non-enrolled learners. Assessment provides bounded learning evidence subject to coverage and validity. Observation and work samples can explain classroom conditions without estimating national prevalence. Learner and teacher accounts can identify barriers. Agreement adds confidence after dates and definitions align; contradiction should guide further enquiry rather than selective reporting. The required evidence includes responsibility, staff capability, learner support, milestones and correction. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Designing the targeted plan within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    Teacher guidance without planning time, materials without accessible use, or tutoring without safe attendance cannot deliver the expected mechanism. Local adaptation should remain possible within a common substantive condition. Any departure should be recorded with its reason and expected learner consequence. The governing action is to choose a feasible intensity and protect the common educational entitlement. Implementation should identify who acts, with what authority, resources and deadline. Dependencies should be sequenced. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Designing the targeted plan within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    66

    Resourcing equitable implementation

    If the assessment population excludes learners most at risk, the gap is not adequately defined. The plan should state which observation is direct, which is estimated and which remains unknown. The improvement question concerns the staff, time, materials, accessibility and finance required by the selected mechanism. It should be stated with the learner population, educational domain, geography and evidence period. A national average is insufficient when the plan addresses a local or group disparity. The baseline should preserve the underlying distribution and absolute numbers. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Resourcing equitable implementation within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-13] [REF-14]

    This discipline prevents a targeted plan from attaching deficit to learners instead of changing institutions. The principal error is weak institutions being expected to implement with the same nominal allocation. Avoiding it requires evidence on both learning and the opportunity supplied. Achievement differences may be associated with poverty, language, disability or residence, but those characteristics are not instructional mechanisms. Authorities should examine curriculum, teacher availability, time, attendance, materials, assessment access and learner support. Alternative explanations should remain open until evidence discriminates among them. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Resourcing equitable implementation within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    Learner and teacher accounts can identify barriers. Agreement adds confidence after dates and definitions align; contradiction should guide further enquiry rather than selective reporting. The required evidence includes recurrent cost, teacher workload, additional need and household burden. Each source should be used within scope. Administrative records can describe staffing and participation while omitting non-enrolled learners. Assessment provides bounded learning evidence subject to coverage and validity. Observation and work samples can explain classroom conditions without estimating national prevalence. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Resourcing equitable implementation within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    Implementation should identify who acts, with what authority, resources and deadline. Dependencies should be sequenced. Teacher guidance without planning time, materials without accessible use, or tutoring without safe attendance cannot deliver the expected mechanism. Local adaptation should remain possible within a common substantive condition. Any departure should be recorded with its reason and expected learner consequence. The governing action is to direct greater support where barriers and implementation costs are greater. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Resourcing equitable implementation within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-02]

    67

    Monitoring reach, quality and learning response

    A national average is insufficient when the plan addresses a local or group disparity. The baseline should preserve the underlying distribution and absolute numbers. If the assessment population excludes learners most at risk, the gap is not adequately defined. The plan should state which observation is direct, which is estimated and which remains unknown. The improvement question concerns evidence that the intended learners received the intervention as designed and benefited educationally. It should be stated with the learner population, educational domain, geography and evidence period. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Monitoring reach, quality and learning response within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19]

    Authorities should examine curriculum, teacher availability, time, attendance, materials, assessment access and learner support. Alternative explanations should remain open until evidence discriminates among them. This discipline prevents a targeted plan from attaching deficit to learners instead of changing institutions. The principal error is participation counts being treated as learning evidence. Avoiding it requires evidence on both learning and the opportunity supplied. Achievement differences may be associated with poverty, language, disability or residence, but those characteristics are not instructional mechanisms. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Monitoring reach, quality and learning response within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    Administrative records can describe staffing and participation while omitting non-enrolled learners. Assessment provides bounded learning evidence subject to coverage and validity. Observation and work samples can explain classroom conditions without estimating national prevalence. Learner and teacher accounts can identify barriers. Agreement adds confidence after dates and definitions align; contradiction should guide further enquiry rather than selective reporting. The required evidence includes eligibility, offer, take-up, dosage, teaching quality, work and assessment. Each source should be used within scope. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Monitoring reach, quality and learning response within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-02]

    Any departure should be recorded with its reason and expected learner consequence. The governing action is to combine timely implementation evidence with valid learning review. Implementation should identify who acts, with what authority, resources and deadline. Dependencies should be sequenced. Teacher guidance without planning time, materials without accessible use, or tutoring without safe attendance cannot deliver the expected mechanism. Local adaptation should remain possible within a common substantive condition. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Monitoring reach, quality and learning response within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-03] [REF-06]

    68

    Adaptation, institutionalisation and exit

    The improvement question concerns reasoned decisions to continue, change, scale or end the plan. It should be stated with the learner population, educational domain, geography and evidence period. A national average is insufficient when the plan addresses a local or group disparity. The baseline should preserve the underlying distribution and absolute numbers. If the assessment population excludes learners most at risk, the gap is not adequately defined. The plan should state which observation is direct, which is estimated and which remains unknown. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Adaptation, institutionalisation and exit within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-24]

    Avoiding it requires evidence on both learning and the opportunity supplied. Achievement differences may be associated with poverty, language, disability or residence, but those characteristics are not instructional mechanisms. Authorities should examine curriculum, teacher availability, time, attendance, materials, assessment access and learner support. Alternative explanations should remain open until evidence discriminates among them. This discipline prevents a targeted plan from attaching deficit to learners instead of changing institutions. The principal error is temporary measures persisting without benefit or disappearing before durable capability exists. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Adaptation, institutionalisation and exit within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-02]

    Agreement adds confidence after dates and definitions align; contradiction should guide further enquiry rather than selective reporting. The required evidence includes thresholds, adverse effects, unresolved cases, recurrent ownership and later evidence. Each source should be used within scope. Administrative records can describe staffing and participation while omitting non-enrolled learners. Assessment provides bounded learning evidence subject to coverage and validity. Observation and work samples can explain classroom conditions without estimating national prevalence. Learner and teacher accounts can identify barriers. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Adaptation, institutionalisation and exit within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-03] [REF-06]

    Teacher guidance without planning time, materials without accessible use, or tutoring without safe attendance cannot deliver the expected mechanism. Local adaptation should remain possible within a common substantive condition. Any departure should be recorded with its reason and expected learner consequence. The governing action is to retain useful capability while ending ineffective or inequitable arrangements. Implementation should identify who acts, with what authority, resources and deadline. Dependencies should be sequenced. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Adaptation, institutionalisation and exit within Part XX — Targeted improvement planning must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09]

    Part XXI

    National translation after adoption of the 2030 Agenda

    69

    Fixing the contemporaneous institutional baseline

    National planning should begin from the institutional position that existed on 10 March 2018. The Incheon Declaration expressed the education community's commitment to inclusive and equitable quality education and lifelong learning; Addis supplied a financing framework; the 2030 Agenda established Goal 4 and its targets; the Education 2030 Framework for Action supplied implementation guidance; and resolution 71/313 established the global indicator framework. National assessment programmes should preserve their stated constructs and populations while documenting any relation to global monitoring.[REF-25] [REF-26] [REF-27] [REF-28] [REF-31]

    It does not prove that national law, plans, budgets, information or delivery arrangements already satisfy the commitment. Each country should identify which obligations and targets can be acted upon under existing authority, which require legislative or administrative change, and which depend on clarification through later competent decisions. Unsettled detail should be recorded as such rather than filled by anticipation. The distinction between adoption and implementation is essential. Adoption establishes an agreed direction and permits governments to begin alignment. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Fixing the contemporaneous institutional baseline within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-07] [REF-12] [REF-27]

    National authorities should map the newly adopted targets against those instruments and against the education stages, populations and institutions already in law and plans. This avoids two risks: abandoning useful evidence because terminology changed, and claiming continuity where a new target has broader scope or different educational substance. The baseline should preserve existing education commitments and national evidence. The new agenda does not erase the right to education, the unfinished Education for All undertaking, programme structures or established statistical series. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Fixing the contemporaneous institutional baseline within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-09] [REF-25]

    The register should be revised openly as competent bodies act. A dated commitment register can support institutional accuracy. For each relevant proposition, it should record the adopting body, date, legal or policy status, national authority, present implementing instrument and unresolved question. This is not an administrative inventory for its own sake. It prevents a proposed measure from being represented as an obligation, a declaration from being treated as proof of delivery, and a later decision from being projected backwards. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Fixing the contemporaneous institutional baseline within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-19] [REF-27]

    Public communication should use the same discipline. Governments can state that the 2030 Agenda has been adopted and that national alignment is commencing. They should not state that every indicator, national milestone or implementation mechanism has already been internationally settled. A precise account strengthens credibility because it makes clear which choices belong to national democratic and administrative processes and which follow directly from adopted commitments. It also establishes a reliable point from which later implementation can be judged. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Fixing the contemporaneous institutional baseline within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-25] [REF-26] [REF-27]

    70

    Selecting priorities without narrowing the commitment

    The breadth of Goal 4 requires sequencing, not selective abandonment. A national plan cannot improve every condition simultaneously, yet it should retain a complete map of early childhood, primary and secondary education, technical and vocational learning, tertiary participation, adult learning, relevant skills, equality, literacy, learning environments, scholarships and teachers as they appear in the adopted targets. A first improvement priority should be chosen because evidence shows a serious and remediable break, not because other elements have ceased to matter. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Selecting priorities without narrowing the commitment within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-25] [REF-27]

    Priority selection should apply four tests. The entitlement test asks which population and educational condition are at stake. The consequence test asks the scale and severity of the denial. The actionability test asks whether a competent authority has a plausible means of change. The equity test asks whether the measure will reach learners farthest from the secured opportunity. A priority that scores highly on visibility but weakly on consequence or equity should be reconsidered. The reasons for selection should be published beside the evidence and known limitations. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Selecting priorities without narrowing the commitment within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-03] [REF-12]

    It should also distinguish poor data from satisfactory conditions. Where the least-served population is weakly observed, strengthening coverage can itself become an immediate priority while urgent service evidence supports proportionate protection. National averages should not determine the sequence alone. A moderate national gap may conceal acute failure in one district or population; a large aggregate shortfall may require broad system expansion alongside targeted support. Analysis should retain the national level, absolute number affected, subnational and group distributions and the minimum educational floor. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Selecting priorities without narrowing the commitment within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-06] [REF-24]

    Dependencies should influence sequence. An assessment reform cannot improve learning without curriculum alignment, teacher capability and participation. Expansion of secondary places may depend on primary completion, trained staff and facilities. Adult learning may require flexible provision, recognition and learner support. The plan should make these dependencies visible and decide which must precede or accompany the selected measure. Otherwise a high-level commitment can be converted into an isolated activity unable to change the learner-facing condition. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Selecting priorities without narrowing the commitment within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-03] [REF-16] [REF-18]

    Selection should remain revisable. New evidence may show that the diagnosed mechanism was wrong, that another population is more severely affected or that implementation capacity is insufficient. Revision is not a retreat from ambition when reasons and consequences are published. It is a condition of responsible improvement. The complete commitment map should remain in view so that repeated concentration on one readily measured target does not produce silent neglect of lifelong learning, equality, educational quality or populations outside formal schooling. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Selecting priorities without narrowing the commitment within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-25] [REF-27]

    71

    From global target to national improvement proposition

    For example, a commitment to equitable quality education is too broad to guide one implementation decision; a proposition to improve regular attendance for a defined remote population through transport, staffing and calendar changes can be examined and corrected. The narrower proposition remains connected to the universal commitment and should not be mistaken for its completion. A national improvement proposition should translate a broad target into a bounded statement of change. It should name the population, present educational condition, institutional mechanism, responsible body, resources, time and evidence of success. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, From global target to national improvement proposition within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-07] [REF-12] [REF-27]

    These mechanisms require different responses. A government should compare administrative, household, assessment and local service evidence and state where inference remains uncertain. Consultation with teachers, learners and communities can identify mechanisms, but it should not replace representative population evidence when prevalence is claimed. Diagnosis should precede instrument choice. A low completion rate may reflect late entry, repetition, household cost, distance, school safety, language, disability exclusion, teacher shortage or unreliable records. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, From global target to national improvement proposition within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-03] [REF-24]

    The intervention hypothesis should explain how the proposed measure changes the barrier. Additional materials will not improve learning if teachers lack time or knowledge to use them; professional guidance will not improve attendance where transport is decisive; a new indicator will not correct exclusion without authority and resources. The plan should identify necessary dependencies and foreseeable adverse effects. It should specify which part of the hypothesis is established by evidence and which remains to be tested during implementation. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, From global target to national improvement proposition within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-06] [REF-12] [REF-16]

    The public record should show the relationship between the global target, national definition and selected measure, including any material difference from an existing series. National adaptation should preserve educational substance. Targets may need national definitions for programme levels, age groups, language and institutional responsibility. Adaptation is legitimate where it makes the commitment operational and comparable over time. It is not legitimate where it narrows the entitled population, lowers expectations for disadvantaged groups or converts learning into attendance alone. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, From global target to national improvement proposition within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-10] [REF-25]

    The proposition should end with a decision rule. Authorities should state what evidence would justify continuation, expansion, adaptation or cessation and when that decision will occur. A pilot that continues because it attracts support rather than because it changes the intended condition is not an improvement method. Conversely, an intervention should not be abandoned merely because early outcomes are uncertain where delivery has not reached intended intensity. Review must distinguish theory failure, implementation failure, measurement weakness and insufficient time. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, From global target to national improvement proposition within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-11] [REF-23]

    72

    Aligning authority, finance and professional capability

    National policy may set the priority, but regional administrations, municipalities, schools, training institutions or other bodies may control staffing, facilities and learner support. The plan should allocate each function to the body able to perform it and identify escalation where local authority is insufficient. Coordination should not allow responsibility to become diffuse. One public owner should remain answerable for whether the learner-facing condition changes. Implementation requires a chain of competent authority. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Aligning authority, finance and professional capability within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-12] [REF-17] [REF-19]

    Finance should cover the complete intervention rather than visible start-up items. Staff preparation, salaries, accessible materials, transport, maintenance, guidance, assessment, evidence and review may all be necessary. The Addis Ababa Action Agenda places national action within a broader financing context, but an international commitment does not supply a national cost estimate. Authorities should identify recurrent and capital requirements, the source and timing of funds, distribution rules and the conditions for continuity after temporary support. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Aligning authority, finance and professional capability within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-26]

    Announced expenditure is not proof of delivery: review should trace authorization, transfer, institutional use and learner consequence. Where households continue to bear material cost, formal fee policy should not be represented as full accessibility. Allocation should respond to unequal cost and starting capacity. Equal per-learner funding can reproduce inequality where remoteness, disability, language, insecurity or weak infrastructure makes adequate provision more expensive. Formulae should state which need factors are recognised and should be checked against actual receipt. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Aligning authority, finance and professional capability within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-06] [REF-10] [REF-13]

    A plan dependent on exceptional individuals or uncompensated workload is not institutionally secure and can deepen disparity between strong and weak institutions. Teachers and institutional leaders require capability proportionate to the change. A new curriculum, assessment or inclusion expectation needs more than notification. Professional learning should provide subject substance, practical examples, time for collaboration and a route to support. Staffing and turnover should be examined in the locations expected to implement first. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Aligning authority, finance and professional capability within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-18]

    International cooperation should strengthen ordinary national capacity. External support may finance initial expansion, evidence, technical work or regional learning, but roles, conditions and exit should be clear. Parallel activities and reporting arrangements can fragment public authority. Assistance should align with a national improvement proposition, use compatible records and establish how essential functions will enter recurrent provision. Its success should be judged by national capability and equitable learner opportunity, not the duration or visibility of the supporting activity. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Aligning authority, finance and professional capability within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-22] [REF-26] [REF-27]

    73

    Monitoring delivery, reach and educational consequence

    Monitoring should follow the causal sequence of the improvement proposition. Inputs show whether resources and staff were available; delivery evidence shows whether the measure operated; reach shows which eligible learners participated and with what intensity; educational evidence shows whether the intended condition changed. These stages should not be collapsed. A budget can be executed without materials arriving, a programme can operate without reaching disadvantaged learners, and participation can rise without improvement in learning or progression. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Monitoring delivery, reach and educational consequence within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-03] [REF-06] [REF-12]

    Where a new target requires a new measure, authorities should preserve the preceding series and identify any break. Apparent improvement caused by revised population estimates, wider institutional reporting or changed assessment participation should be separated from educational change. Comparable trends are valuable, but continuity should not be asserted where concepts differ materially. The baseline should retain numerator, denominator, population, date, geography, definition and exclusions. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Monitoring delivery, reach and educational consequence within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-23] [REF-24]

    Disaggregation must remain lawful, meaningful and safe. Small numbers may require controlled access or combined reporting, but the affected population and public responsibility should not disappear. Equity monitoring should show eligibility, offer, take-up, attendance, completion and outcome for relevant groups and places. A plan can improve its average by reaching learners already closest to the desired condition. It should therefore report the least-served position and protect against exclusion or deterioration. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Monitoring delivery, reach and educational consequence within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-10] [REF-13]

    Agreement across sources strengthens a conclusion only after definitions and dates align; disagreement should guide investigation rather than selective reporting. Learning evidence requires population coverage and opportunity to learn. Assessment results should identify the domain, eligible population, participation and exclusions. A favourable mean among tested learners cannot represent those outside school or absent from assessment. Classroom observation, work samples and learner accounts can illuminate mechanism without establishing national prevalence. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Monitoring delivery, reach and educational consequence within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-03] [REF-24]

    Public reporting should connect the result to a decision. A concise account should state what was implemented, who received it, what changed, what remains uncertain, which adverse effects occurred and whether the measure will continue, adapt, expand or cease. The next review date and responsible authority should be visible. This makes monitoring a means of correction rather than an obligation to produce favourable figures. The adopted agenda gains national credibility when evidence can alter action and expose populations who remain underserved. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Monitoring delivery, reach and educational consequence within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-25] [REF-27]

    74

    First national decisions after adoption

    Governments can designate a competent coordinating authority, preserve existing sector responsibilities, assemble a dated commitment register and commission a baseline review. The review should map every adopted education target against national law, plans, budgets and evidence. It should identify urgent gaps and unsettled definitions separately. This creates a disciplined bridge between global adoption and national action. The immediate national decision is to establish governance for alignment without pretending that implementation detail is complete. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, First national decisions after adoption within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-25] [REF-27]

    Where credible evidence reveals severe exclusion or harm, protective action need not await a perfect estimate. Longer-term allocation, however, should be reviewed as coverage improves. The plan should guard against choosing only learners and institutions most likely to produce rapid favourable results. A first priority should be small enough for accountable action and important enough to change educational opportunity. It should name the population and condition, state why it takes precedence, and preserve the wider commitment map. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, First national decisions after adoption within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-06] [REF-12]

    A financing gap should be described rather than hidden through reduced educational substance or transfer of cost to poor households. The first budget decision should identify recurrent implications. Temporary finance may permit testing, but teachers, learner support, accessible facilities, maintenance and evidence cannot be sustained by an announcement. The national authority should state the source, timing and distribution of finance and how external cooperation relates to ordinary provision. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, First national decisions after adoption within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-26]

    Accurate status is a condition of accountable planning, not a reason for delay. The first public report should be candid about chronology. It can state that three major texts relevant to education and financing had been adopted by the cut-off, including the 2030 Agenda on the preceding day. It should also state that later implementation instruments are outside the record. This protects the difference between an adopted commitment, national policy choice and future institutional clarification. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, First national decisions after adoption within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-25] [REF-26] [REF-27]

    It should examine authority, delivery, reach, professional capability, finance, learner experience and early educational consequence. It should record adverse findings and adapt the measure where the hypothesis or delivery proves weak. The transition from global commitment to national improvement is complete only when ordinary institutions can sustain the changed condition and correct foreseeable departures; the end of a project or reporting period is not evidence of that result. The first review should test whether institutions learned, not merely whether a plan was issued. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, First national decisions after adoption within Part XXI — National translation after adoption of the 2030 Agenda must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-12] [REF-23]

    Part XXII

    Quality and equity within the contemporaneous indicator framework

    75

    The hierarchy from goal to observation

    Authority and inference narrow at each stage. A measure can support judgement on part of a target without becoming the target's complete meaning. The indicator framework should be read through a hierarchy of public meaning. Goal 4 states the overarching commitment; each target identifies an educational change or condition; an indicator selects one observable aspect; a national measure applies definitions, sources and calculations; and a released value describes a population and period. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, The hierarchy from goal to observation within Part XXII — Quality and equity within the contemporaneous indicator framework must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-27] [REF-28] [REF-29]

    Interpretation should therefore state the specific proposition supported and name the remaining parts of the target that require other evidence. This distinction is essential for quality and equity because both are multidimensional. A learning indicator can provide important evidence on one domain and population, but cannot establish free access, regular participation, safety, curricular breadth or recognised progression. A parity measure can reveal a relative relationship while concealing low levels for both groups or exclusion from the denominator. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, The hierarchy from goal to observation within Part XXII — Quality and equity within the contemporaneous indicator framework must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-03] [REF-12]

    The General Assembly's 2017 adoption of the global framework supports common monitoring while recognising that indicators undergo methodological development. It does not justify presenting unlike assessments as one scale or every operational definition as fixed for all purposes. A responsible national record should retain the indicator version, metadata, source and calculation used at each release. Where refinement changes the population or construct, the effect should be assessed and the series broken when continuity is not defensible. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, The hierarchy from goal to observation within Part XXII — Quality and equity within the contemporaneous indicator framework must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-16] [REF-23] [REF-31] [REF-33]

    The common measure supports international orientation; the national measure supports domestic decision. Neither should be preferred merely because it produces a more favourable result. A correspondence statement should identify population, education level, event, reference period and material divergence so that readers can see whether two values answer the same question. National measures may add detail relevant to law, programme structure and known barriers, provided their relation to the global concept is documented. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, The hierarchy from goal to observation within Part XXII — Quality and equity within the contemporaneous indicator framework must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-19] [REF-28]

    An unavailable measure identifies a statistical gap; an unfavourable value identifies an observed condition; an incomplete target judgement requires a broader evidentiary account. These are different findings with different responsible bodies and remedies. Public reporting should preserve this hierarchy in its language. It should not say that a goal was achieved because one indicator improved, or that an indicator failed because a source is temporarily unavailable. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, The hierarchy from goal to observation within Part XXII — Quality and equity within the contemporaneous indicator framework must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-17] [REF-29]

    76

    Interpreting quality beyond one outcome

    Its evidence can include curriculum, teachers, instructional time, facilities, safety, accessibility, assessment, learner work and recognised qualifications. No single item establishes the whole. Inputs can be present but unused; learning can be observed among a selected population; completion can carry weak or uncertain educational value. The framework should therefore be interpreted as a related set rather than a search for one quality proxy. Quality concerns the educational substance and conditions through which learners participate, learn, complete and progress. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Interpreting quality beyond one outcome within Part XXII — Quality and equity within the contemporaneous indicator framework must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-09] [REF-12] [REF-28]

    Age-based results may represent a broader population while using different educational correspondence. School-based assessment cannot represent children outside provision without an explicit population model. Participation and exclusions should accompany the result so that improved coverage is not mistaken for deteriorating learning or vice versa. Learning measures require a specified domain, target population, instrument, language, administration and proficiency threshold. Grade-based results describe learners who reached the grade and participated under stated conditions. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Interpreting quality beyond one outcome within Part XXII — Quality and equity within the contemporaneous indicator framework must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-03] [REF-24]

    A report should state which part of curriculum is observed and should not infer general school quality from a limited instrument. Where a target concerns citizenship or sustainable development, policy presence and learner capability should remain separate stages. Relevant and effective learning outcomes should not be narrowed to whatever domain is easiest to compare. National curricula and public purposes remain material. Comparative measures can establish a common bounded domain, while national evidence addresses additional knowledge, capabilities and progression. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Interpreting quality beyond one outcome within Part XXII — Quality and equity within the contemporaneous indicator framework must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-15] [REF-24] [REF-27]

    Reporting should connect teacher evidence to the learners and institutions exposed to the condition rather than assume that a national workforce total describes classroom opportunity uniformly. Teachers are both a means-of-implementation concern and a condition of quality across targets. Counts should be interpreted with preparation, assignment, attendance, workload and support. A national ratio can conceal shortage by subject, level or locality. Qualification definitions vary and need national metadata. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Interpreting quality beyond one outcome within Part XXII — Quality and equity within the contemporaneous indicator framework must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-04] [REF-18] [REF-28]

    Quality interpretation should lead to action. If low learning coincides with weak participation, response should not concentrate solely on assessed pupils. If teacher shortage is local, a national recruitment measure may be too broad or too slow. If facilities exclude learners with disabilities, an aggregate infrastructure count is insufficient. The public account should connect the observed condition to the competent authority, resource and later review without claiming attribution beyond the evidence. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Interpreting quality beyond one outcome within Part XXII — Quality and equity within the contemporaneous indicator framework must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-13] [REF-17]

    77

    Defining the assessment claim

    It does not observe learning in its entirety. Responsible interpretation begins by naming the construct, target population, curriculum or framework relation, language, mode, reference period and reporting scale. The result should be no broader than those design choices permit. A reading score cannot stand for complete educational quality, and a national mean cannot describe every learner's opportunity. A large-scale learning assessment produces evidence about performance on a defined set of tasks under specified administration and scoring conditions. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Defining the assessment claim within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03] [REF-33]

    A sample-based monitoring assessment may estimate population patterns without producing defensible individual scores. A census assessment may still be unsuitable for high-stakes individual decisions if the construct, reliability or administration was designed for aggregate use. An available result should not acquire a new purpose without renewed validity and fairness examination. The assessment purpose should be fixed before results are examined. System monitoring, curriculum review, international comparison, certification and classroom diagnosis demand different coverage, precision and consequences. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Defining the assessment claim within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-12] [REF-23]

    School-based assessments omit children outside school and may omit learners absent on the day, excluded from testing or placed in unrecognised provision. The participation record should show eligibility, exclusions, absence and completed instruments; the commentary should explain how these states affect the estimate. The target population is an essential part of the claim. An age-based assessment and a grade-based assessment answer different questions where enrolment, repetition, acceleration or exclusion varies. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Defining the assessment claim within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-06] [REF-32]

    Performance on a narrow item set should not be represented as general capability without evidence that the set adequately samples the intended domain. Construct definition should identify the knowledge and cognitive activity elicited, not merely the subject label. Reading may concern retrieval, interpretation, evaluation and use across text forms; mathematics may concern concepts, procedures, application and reasoning. Item formats, language and time constraints influence which of these are observed. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Defining the assessment claim within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-32] [REF-33]

    Opportunity to learn belongs in interpretation. Learners cannot reasonably be judged against content, language or formats to which they had no meaningful access, although absence of opportunity does not make low learning unimportant. The appropriate conclusion may be that the system failed to secure both opportunity and outcome. Curriculum mapping, teacher evidence and learner work can help distinguish an assessment mismatch from a substantive gap, while neither should be used automatically to dismiss the result. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Defining the assessment claim within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-12] [REF-32]

    78

    Scales, thresholds and uncertainty

    Scale points are not natural units of learning. A ten-point difference has meaning only through the scale construction, score distribution, uncertainty and substantive descriptions attached to performance. Reports should avoid expressions that turn an arbitrary unit into a simple quantity of curriculum learned or time gained unless an explicit and defensible linking study supports that interpretation. Assessment scales order or locate performance according to a statistical model. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Scales, thresholds and uncertainty within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-23]

    Learners immediately above and below a cut may have nearly indistinguishable performance, while learners within one category may differ substantially. Reports should present level shares alongside distributions, uncertainty and, where relevant, sensitivity to reasonable threshold choices. Proficiency levels can make results intelligible when cut scores and descriptors reflect a coherent construct. They also compress variation and create apparent discontinuity around a threshold. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Scales, thresholds and uncertainty within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-03] [REF-33]

    Means, percentiles, level distributions and item-domain evidence can complement the threshold, provided the set remains concise and linked to decisions. A minimum proficiency threshold has normative importance but should not become the sole account of learning. It shows the share reaching a specified standard; it does not show how far learners below the threshold are from it, whether learners above it continue to progress, or which domain components are weak. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Scales, thresholds and uncertainty within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-27] [REF-31]

    Sampling error should accompany estimates from samples, but it is not the only uncertainty. Non-response, exclusions, translation, administration departures, missing items, scoring disagreement and model assumptions may influence results. A narrow confidence interval does not correct systematic undercoverage or construct weakness. Technical quality assurance should examine each source of error and state which is quantified, bounded through procedure or unresolved. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Scales, thresholds and uncertainty within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-33]

    Statistical significance concerns compatibility with a sampling model; educational significance concerns the magnitude and consequence of the difference. Multiple comparisons increase the likelihood of chance findings. Reports should pre-specify principal comparisons, present intervals, examine robustness and resist converting every ordinal position into a judgement of system quality. Small differences should not be ranked as certain educational differences. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Scales, thresholds and uncertainty within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-33]

    79

    Comparability across populations and time

    Translation should retain meaning and difficulty as far as possible, but formal equivalence of wording does not assure that tasks function alike across languages and cultures. Differential item evidence and expert review can identify threats, while remaining limitations should narrow the claim. Comparability is an empirical property, not a consequence of using the same assessment title. Versions must preserve sufficient construct, item, administration and scaling continuity. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Comparability across populations and time within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-33]

    Trend interpretation requires a stable relation among construct, target population, administration and scale. Curriculum reform, changes in enrolment, altered exclusion rules, new test modes or disruptions to schooling can change the meaning of the population result. A trend may still be estimable through common items, bridge studies and documented adjustments, but the report should identify what remained comparable and which change cannot be separated from educational performance. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Comparability across populations and time within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-16] [REF-23]

    International comparison requires compatible ages or grades, populations and constructs. A common instrument strengthens comparison, yet countries may differ in school coverage, language, repetition, migration and participation. National context should explain the result without redefining the common measure or excusing a failed entitlement. League positions should not replace examination of distributions, minimum proficiency and conditions that may be changed by policy. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Comparability across populations and time within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-32]

    Low participation among learners with disabilities, displaced learners, remote communities or language minorities may create a favourable measured average while concealing unmet need. Accommodations should preserve the intended construct and be recorded. When some learners cannot access the ordinary instrument, alternative evidence should keep them within the public account rather than mark them as irrelevant to system performance. Assessment participation can itself be an equity result. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Comparability across populations and time within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-30]

    Where linking is not defensible, reports should present parallel evidence and explain convergence or divergence qualitatively. Refusal of a false common scale is not a failure of monitoring; it protects the meaning of each result. Linking distinct assessments to a common scale requires evidence of construct alignment, common items or populations, model fit and stability. A numerical conversion created from aggregate correlations may conceal important domain differences. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Comparability across populations and time within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-33]

    80

    Distribution, equity and missing learners

    Group comparisons should retain the levels on both sides: parity at a low level is not success, and a narrowing gap caused by deterioration in the initially advantaged group is not equitable progress. Where lawful and sufficiently precise, analysis should consider sex, wealth, residence, disability, language, migration and other nationally relevant dimensions. National means should be accompanied by the distribution and the absolute number of learners below any defined minimum. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Distribution, equity and missing learners within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-06]

    Differences can reflect opportunity, resources, language of instruction, safety, discrimination, health, prior schooling and assessment access. The report should move from observed distribution to plausible mechanisms and then to evidence and authority for correction. Labelling a population as low-performing without examining these conditions can reproduce stigma. Disaggregation should serve a public question and not invite deterministic claims. Group identity is not a causal explanation. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Distribution, equity and missing learners within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-05] [REF-10]

    Intersectional analysis may reveal disadvantage hidden in one-dimensional averages, but cell sizes and disclosure risk constrain detail. Analysis can pool years where the construct is stable, use broader categories, employ protected access or combine quantitative and qualitative evidence. Suppression should protect people without erasing the existence of the educational issue or the obligation to investigate it. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Distribution, equity and missing learners within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-13]

    Household evidence, administrative reconciliation and targeted studies can estimate or describe the missing population. The national report should not call the tested distribution universal when coverage does not support that claim. Missing learners require a separate account. Those not enrolled, absent, displaced, institutionalised or outside recognised schools may face the greatest deprivation and yet have no assessment result. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Distribution, equity and missing learners within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-30] [REF-32]

    The purpose is to identify conditions capable of public action, not to construct an unbounded list of correlations. Equity-oriented reporting should connect findings to resources and teaching conditions without assuming that inputs automatically produce learning. Teacher availability, preparation, language capability, materials, instructional time, accessibility and learner support can be examined as possible mechanisms. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Distribution, equity and missing learners within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-32]

    Part XXII

    Responsible use of large-scale assessment evidence

    77

    Defining the assessment claim

    It does not observe learning in its entirety. Responsible interpretation begins by naming the construct, target population, curriculum or framework relation, language, mode, reference period and reporting scale. The result should be no broader than those design choices permit. A reading score cannot stand for complete educational quality, and a national mean cannot describe every learner's opportunity. A large-scale learning assessment produces evidence about performance on a defined set of tasks under specified administration and scoring conditions. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Defining the assessment claim within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-03] [REF-33]

    A sample-based monitoring assessment may estimate population patterns without producing defensible individual scores. A census assessment may still be unsuitable for high-stakes individual decisions if the construct, reliability or administration was designed for aggregate use. An available result should not acquire a new purpose without renewed validity and fairness examination. The assessment purpose should be fixed before results are examined. System monitoring, curriculum review, international comparison, certification and classroom diagnosis demand different coverage, precision and consequences. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Defining the assessment claim within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-12] [REF-23]

    School-based assessments omit children outside school and may omit learners absent on the day, excluded from testing or placed in unrecognised provision. The participation record should show eligibility, exclusions, absence and completed instruments; the commentary should explain how these states affect the estimate. The target population is an essential part of the claim. An age-based assessment and a grade-based assessment answer different questions where enrolment, repetition, acceleration or exclusion varies. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Defining the assessment claim within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-06] [REF-32]

    Performance on a narrow item set should not be represented as general capability without evidence that the set adequately samples the intended domain. Construct definition should identify the knowledge and cognitive activity elicited, not merely the subject label. Reading may concern retrieval, interpretation, evaluation and use across text forms; mathematics may concern concepts, procedures, application and reasoning. Item formats, language and time constraints influence which of these are observed. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Defining the assessment claim within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-32] [REF-33]

    Opportunity to learn belongs in interpretation. Learners cannot reasonably be judged against content, language or formats to which they had no meaningful access, although absence of opportunity does not make low learning unimportant. The appropriate conclusion may be that the system failed to secure both opportunity and outcome. Curriculum mapping, teacher evidence and learner work can help distinguish an assessment mismatch from a substantive gap, while neither should be used automatically to dismiss the result. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Defining the assessment claim within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-12] [REF-32]

    78

    Scales, thresholds and uncertainty

    Scale points are not natural units of learning. A ten-point difference has meaning only through the scale construction, score distribution, uncertainty and substantive descriptions attached to performance. Reports should avoid expressions that turn an arbitrary unit into a simple quantity of curriculum learned or time gained unless an explicit and defensible linking study supports that interpretation. Assessment scales order or locate performance according to a statistical model. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Scales, thresholds and uncertainty within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-23]

    Learners immediately above and below a cut may have nearly indistinguishable performance, while learners within one category may differ substantially. Reports should present level shares alongside distributions, uncertainty and, where relevant, sensitivity to reasonable threshold choices. Proficiency levels can make results intelligible when cut scores and descriptors reflect a coherent construct. They also compress variation and create apparent discontinuity around a threshold. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Scales, thresholds and uncertainty within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-03] [REF-33]

    Means, percentiles, level distributions and item-domain evidence can complement the threshold, provided the set remains concise and linked to decisions. A minimum proficiency threshold has normative importance but should not become the sole account of learning. It shows the share reaching a specified standard; it does not show how far learners below the threshold are from it, whether learners above it continue to progress, or which domain components are weak. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Scales, thresholds and uncertainty within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-27] [REF-31]

    Sampling error should accompany estimates from samples, but it is not the only uncertainty. Non-response, exclusions, translation, administration departures, missing items, scoring disagreement and model assumptions may influence results. A narrow confidence interval does not correct systematic undercoverage or construct weakness. Technical quality assurance should examine each source of error and state which is quantified, bounded through procedure or unresolved. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Scales, thresholds and uncertainty within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-33]

    Statistical significance concerns compatibility with a sampling model; educational significance concerns the magnitude and consequence of the difference. Multiple comparisons increase the likelihood of chance findings. Reports should pre-specify principal comparisons, present intervals, examine robustness and resist converting every ordinal position into a judgement of system quality. Small differences should not be ranked as certain educational differences. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Scales, thresholds and uncertainty within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-23] [REF-33]

    79

    Comparability across populations and time

    Translation should retain meaning and difficulty as far as possible, but formal equivalence of wording does not assure that tasks function alike across languages and cultures. Differential item evidence and expert review can identify threats, while remaining limitations should narrow the claim. Comparability is an empirical property, not a consequence of using the same assessment title. Versions must preserve sufficient construct, item, administration and scaling continuity. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Comparability across populations and time within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-33]

    Trend interpretation requires a stable relation among construct, target population, administration and scale. Curriculum reform, changes in enrolment, altered exclusion rules, new test modes or disruptions to schooling can change the meaning of the population result. A trend may still be estimable through common items, bridge studies and documented adjustments, but the report should identify what remained comparable and which change cannot be separated from educational performance. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Comparability across populations and time within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-16] [REF-23]

    International comparison requires compatible ages or grades, populations and constructs. A common instrument strengthens comparison, yet countries may differ in school coverage, language, repetition, migration and participation. National context should explain the result without redefining the common measure or excusing a failed entitlement. League positions should not replace examination of distributions, minimum proficiency and conditions that may be changed by policy. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, Comparability across populations and time within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-32]

    Low participation among learners with disabilities, displaced learners, remote communities or language minorities may create a favourable measured average while concealing unmet need. Accommodations should preserve the intended construct and be recorded. When some learners cannot access the ordinary instrument, alternative evidence should keep them within the public account rather than mark them as irrelevant to system performance. Assessment participation can itself be an equity result. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, Comparability across populations and time within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-30]

    Where linking is not defensible, reports should present parallel evidence and explain convergence or divergence qualitatively. Refusal of a false common scale is not a failure of monitoring; it protects the meaning of each result. Linking distinct assessments to a common scale requires evidence of construct alignment, common items or populations, model fit and stability. A numerical conversion created from aggregate correlations may conceal important domain differences. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, Comparability across populations and time within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-02] [REF-33]

    80

    Distribution, equity and missing learners

    Group comparisons should retain the levels on both sides: parity at a low level is not success, and a narrowing gap caused by deterioration in the initially advantaged group is not equitable progress. Where lawful and sufficiently precise, analysis should consider sex, wealth, residence, disability, language, migration and other nationally relevant dimensions. National means should be accompanied by the distribution and the absolute number of learners below any defined minimum. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, Distribution, equity and missing learners within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-01] [REF-06]

    Differences can reflect opportunity, resources, language of instruction, safety, discrimination, health, prior schooling and assessment access. The report should move from observed distribution to plausible mechanisms and then to evidence and authority for correction. Labelling a population as low-performing without examining these conditions can reproduce stigma. Disaggregation should serve a public question and not invite deterministic claims. Group identity is not a causal explanation. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, Distribution, equity and missing learners within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-05] [REF-10]

    Intersectional analysis may reveal disadvantage hidden in one-dimensional averages, but cell sizes and disclosure risk constrain detail. Analysis can pool years where the construct is stable, use broader categories, employ protected access or combine quantitative and qualitative evidence. Suppression should protect people without erasing the existence of the educational issue or the obligation to investigate it. The comparison remains valid only where classification, timing and participation are made explicit. For evidence standards for system-level improvement claims, Distribution, equity and missing learners within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-10] [REF-13]

    Household evidence, administrative reconciliation and targeted studies can estimate or describe the missing population. The national report should not call the tested distribution universal when coverage does not support that claim. Missing learners require a separate account. Those not enrolled, absent, displaced, institutionalised or outside recognised schools may face the greatest deprivation and yet have no assessment result. The finding should be read alongside the educational entitlement and the conditions needed to realise it. For evidence standards for system-level improvement claims, Distribution, equity and missing learners within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-30] [REF-32]

    The purpose is to identify conditions capable of public action, not to construct an unbounded list of correlations. Equity-oriented reporting should connect findings to resources and teaching conditions without assuming that inputs automatically produce learning. Teacher availability, preparation, language capability, materials, instructional time, accessibility and learner support can be examined as possible mechanisms. The authority should distinguish observed association from explanation and attributed effect. For evidence standards for system-level improvement claims, Distribution, equity and missing learners within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-17] [REF-32]

    81

    From results to proportionate action

    Assessment results should initiate diagnosis before prescription. A weak domain may arise from curriculum coverage, teaching knowledge, language, materials, attendance, assessment mismatch or accumulated prior gaps. Different explanations imply different responses. Authorities should combine assessment evidence with curriculum review, observation, learner work and institutional evidence and should record which explanation is sufficiently supported. The resulting judgement should identify both the evidential limitation and the available remedy. For evidence standards for system-level improvement claims, From results to proportionate action within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-12] [REF-32]

    Population estimates can guide system priorities and resource distribution; they do not establish the competence or misconduct of an individual teacher. Institutional comparisons can identify a need for review; they do not establish sole institutional causation. High-stakes use requires stronger precision, stable participation, fair process and evidence of control. Consequences must respect the level of inference. The institutional consequence is a documented decision with a responsible body and review point. For evidence standards for system-level improvement claims, From results to proportionate action within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-08] [REF-29]

    A balanced target may combine improvement in the mean or proficiency share with a floor for the lowest-performing group and a requirement not to reduce participation. Baseline, period, population and assessment version should be fixed, with rules for interpreting breaks in series. Targets should be ambitious, feasible and distribution-sensitive. A national average target can be met while the least-served population remains unchanged. The public account should therefore preserve the population base and the limit of inference. For evidence standards for system-level improvement claims, From results to proportionate action within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-27] [REF-31]

    Technical documentation should remain accessible, while a public summary explains in plain institutional language what the assessment can and cannot show. Public communication should separate the observed estimate, uncertainty, interpretation and proposed response. Visual simplicity should not erase the population or confidence interval. Rankings should be avoided when differences are unstable or constructs unlike. The equity consequence is to identify learners absent from the observed distribution. For evidence standards for system-level improvement claims, From results to proportionate action within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-03] [REF-33]

    Follow-up should test whether action reached the intended learners and altered the diagnosed condition. Repeating the assessment may show later population change, but interim evidence should examine implementation, opportunity and learner work. Improvement should not be attributed automatically to the intervention; the review should consider population, policy and measurement change. A responsible assessment cycle ends in a reasoned decision, not in publication alone. The operational consequence is to connect the finding to resources, professional action and follow-up. For evidence standards for system-level improvement claims, From results to proportionate action within Part XXII — Responsible use of large-scale assessment evidence must establish whether observed change is comparable, distributed and plausibly connected to action.[REF-16] [REF-24] [REF-32]

    Part XXIII

    Standards for a system-level improvement claim

    82

    The form of the claim

    A claim of system-level learning improvement should state the population, learning domain, period, assessment, magnitude and distribution to which it refers. “Learning improved” is incomplete if the evidence concerns only enrolled learners in selected grades, one assessed domain or a sample that excludes important territories. The public statement should identify the baseline and later estimate, their uncertainty, the rules governing participation and any material change in coverage.[REF-02] [REF-23] [REF-33]

    System level does not mean national mean alone. It concerns the reach and operation of public arrangements across the relevant system and should therefore examine variation among places, institutions and learner groups. Improvement in the mean accompanied by deterioration among learners below a substantive minimum is a qualified result, not unambiguous system progress. The standard should require absolute levels and distributions alongside change.[REF-01] [REF-06] [REF-32]

    The claim should distinguish achievement, improvement and target attainment. A country may improve while remaining far below a minimum; another may maintain a high level without a measurable short-term increase; a third may meet an aggregate target while leaving an excluded population unmeasured. Each statement has a different evidentiary basis and public meaning. Combining them into one success label impedes fair comparison and useful action.[REF-27] [REF-31]

    Precision of language should follow precision of evidence. “The estimate increased” may be justified where uncertainty overlaps substantially; “evidence is consistent with improvement” may be justified where several sources converge but comparability is imperfect; “the system caused the increase” requires a causal design and evidence of mechanism. Responsible reporting uses the strongest formulation supported, not the strongest formulation politically desired.[REF-16] [REF-24]

    83

    Establishing a defensible baseline

    The baseline should precede or coincide with the period of action and use a population and construct relevant to the objective. A retrospective baseline chosen after seeing several years of data invites selection of an unusually low point. If circumstances require retrospective selection, the rule and all available candidate periods should be disclosed. A multi-year baseline may reduce volatility when definitions and constructs are stable.[REF-02] [REF-33]

    Baseline coverage should be reconciled with the education population. Enrolment expansion can lower the assessed mean by including learners who previously lacked access even while the system improves in equity and total learning. Conversely, exclusion of struggling learners can raise the mean without improving anyone's capability. Reports should present participation, absolute numbers and composition so that changes in who is assessed are not misread as changes in how well the system teaches.[REF-01] [REF-32]

    The baseline instrument requires a retained record of tasks or specifications, administration, scoring, scaling and quality assurance sufficient to support later linkage. Secure retention does not require public release of protected items, but it does require institutional memory. If the evidence necessary for linkage was not retained, a later series should acknowledge the break rather than reconstruct certainty from incomplete documentation.[REF-12] [REF-33]

    External conditions at baseline should be recorded selectively where they bear on interpretation: conflict, displacement, major population movement, disaster, prolonged interruption, curriculum reform or exceptional assessment disruption. Context does not redefine the learning standard. It helps explain the starting point, identify relevant comparison and determine which authority must act.[REF-30] [REF-32]

    84

    Demonstrating comparable change

    Change requires comparable constructs and populations. Common scale labels or proficiency categories do not themselves establish equivalence. The assessment authority should show how content coverage, cognitive demand, language, mode, administration and scaling were maintained or linked. Where a change was intentional, bridge evidence should estimate its effect. Where that effect cannot be separated, the result should be reported as a new baseline.[REF-23] [REF-33]

    Trend uncertainty includes more than sampling error. Shifts in non-response, exclusions, accommodations, school participation and data cleaning can alter the estimate. The technical review should quantify effects where possible and provide sensitivity when alternative reasonable treatments affect the conclusion. A result that changes sign under plausible assumptions cannot sustain an unqualified claim.[REF-02] [REF-33]

    Several assessment cycles strengthen evidence of sustained change, but frequency must fit the educational mechanism. A one-cycle increase may reflect cohort composition or transient conditions. Repetition can show persistence, while interim evidence should verify changes in teaching, materials, attendance and learner opportunity. Waiting for repeated outcome measures should not delay action where current evidence reveals serious deprivation.[REF-12] [REF-32]

    Convergence among assessments supports a claim only where each source contributes relevant and sufficiently independent information. Two scores derived from the same learner sample and overlapping items do not constitute two independent confirmations. Administrative, classroom and qualitative evidence can clarify reach and mechanism, but should not be represented as an equivalent measure of the same outcome.[REF-16] [REF-24]

    85

    From temporal change to causal contribution

    A system-level outcome can change after reform without having changed because of reform. A causal contribution claim should specify the intervention, expected sequence, population reached, intensity, implementation period and mechanism linking action to learning. It should examine other policies, demographic change, economic conditions and assessment changes capable of producing the observed pattern.[REF-24] [REF-32]

    The evidentiary design should be proportionate to the claim. Phased implementation, comparison groups, discontinuities, exposure gradients or other structured comparisons may strengthen causal inference when ethically and operationally appropriate. No design is self-interpreting. Selection, spillovers, concurrent change and implementation variation should be examined, and the result should remain bounded to the studied conditions.[REF-16] [REF-23]

    Mechanism evidence is particularly important for national reform. A policy announcement or expenditure record does not establish altered instruction. Evidence should follow the chain from authority and finance to materials, professional support, classroom use, learner participation and domain-specific work. A broken link can explain why a plausible policy did not produce the expected result and can direct correction more effectively than a general verdict.[REF-17] [REF-32]

    Attribution among levels should reflect control. National authorities may be accountable for curriculum, finance and staffing rules; local bodies for allocation and support; institutions for organisation and teaching conditions; professionals for decisions within their competence. A national learning change should not be assigned wholly to one level without evidence. Shared contribution requires specified duties rather than diffuse collective praise or blame.[REF-08] [REF-29]

    86

    Equity, adverse effects and public proof

    An improvement claim should test whether benefits reached learners with the weakest baseline opportunity. Pre-specified group and place analysis reduces selective reporting, while intersectional analysis may reveal hidden disadvantage. Small numbers require confidentiality and careful uncertainty, but should not permit the disappearance of displaced learners, persons with disabilities or remote communities from the public account.[REF-10] [REF-30]

    Adverse effects should be sought actively. Pressure to raise scores can narrow curriculum, divert teaching from unassessed domains, increase exclusion or concentrate resources on learners near a threshold. Evidence of assessment participation, subject time, admissions, transfers and learner experience can reveal such effects. An improved target measure does not establish net educational improvement if another essential entitlement was impaired.[REF-08] [REF-32]

    Public proof should be reproducible at the level allowed by confidentiality and assessment security. The authority should publish definitions, population rules, estimates, uncertainty, trend methods, exclusions and revision policy. Independent reviewers should be able to test the reasoning and calculations without gaining access to personally identifiable learner records or live secure items.[REF-03] [REF-33]

    Corrections should be visible. If a coding, weighting, linking or population error changes the estimate, the issuing body should identify the affected release, explain the correction and state whether conclusions or decisions change. Silent replacement weakens institutional memory and may leave earlier public claims in circulation. A revision is evidence of quality control when handled openly.[REF-02] [REF-33]

    The minimum acceptable claim therefore joins a bounded learning statement, defensible baseline, comparable later observation, uncertainty, distribution, coverage, contextual explanation and proportionate wording. A causal claim additionally requires verified implementation, a credible comparison or contribution design, mechanism evidence and consideration of alternatives. These standards do not obstruct recognition of progress. They ensure that recognition belongs to real learning among a known population and can guide the next public decision.[REF-27] [REF-32]

    References

    1. REF-01

      Education for All Global Monitoring Report Team. Reaching the Marginalized — EFA Global Monitoring Report 2010. 2010.

      Principal contemporaneous analysis of intersecting disadvantage and education marginalisation.

      https://unesdoc.unesco.org/ark:/48223/pf0000186606
    2. REF-02

      UNESCO Institute for Statistics. Global Education Digest 2010: Comparing Education Statistics Across the World. 2010.

      Comparative education statistics, definitions and limitations.

      https://uis.unesco.org/sites/default/files/documents/global-education-digest-2010-comparing-education-statistics-across-the-world-en.pdf
    3. REF-03

      UNESCO Institute for Statistics. Education Indicators: Technical Guidelines. 2009.

      Definitions and interpretation of participation, progression, completion and resource indicators.

      https://uis.unesco.org/sites/default/files/documents/education-indicators-technical-guidelines-en_0.pdf
    4. REF-04

      United Nations. The Millennium Development Goals Report 2010. 2010.

      Global and regional monitoring of primary education, gender, poverty and related development conditions.

      https://www.un.org/millenniumgoals/pdf/MDG%20Report%202010%20En%20r15%20-low%20res%2020100615%20-.pdf
    5. REF-05

      United Nations Development Programme. Human Development Report 2010: The Real Wealth of Nations — Pathways to Human Development. 2010.

      Distribution-sensitive human development concepts and evidence available before the cut-off.

      https://hdr.undp.org/content/human-development-report-2010
    6. REF-06

      UNICEF. Progress for Children: Achieving the MDGs with Equity, Number 9. 2010.

      Equity-focused child indicators and comparison between population groups.

      https://www.unicef.org/reports/progress-children-no-9
    7. REF-07

      World Education Forum. The Dakar Framework for Action: Education for All — Meeting Our Collective Commitments. 2000.

      Commitments to equitable access, quality, measurable outcomes and accountable national planning.

      https://unesdoc.unesco.org/ark:/48223/pf0000121147
    8. REF-08

      United Nations General Assembly. Convention on the Rights of the Child. 1989.

      Rights concerning non-discrimination, identity, education and development.

      https://www.ohchr.org/en/instruments-mechanisms/instruments/convention-rights-child
    9. REF-09

      United Nations Committee on Economic, Social and Cultural Rights. General Comment No. 13: The Right to Education. 1999.

      Interpretation of availability, accessibility, acceptability and adaptability.

      https://www.refworld.org/legal/general/cescr/1999/en/37937
    10. REF-10

      United Nations General Assembly. Convention on the Rights of Persons with Disabilities. 2006.

      Non-discrimination, accessibility, inclusive education and disability data safeguards.

      https://www.ohchr.org/en/instruments-mechanisms/instruments/convention-rights-persons-disabilities
    11. REF-11

      United Nations. Guiding Principles on Internal Displacement. 1998.

      Principles relevant to protection, documentation, education and non-discrimination of displaced persons.

      https://www.ohchr.org/en/special-procedures/sr-internally-displaced-persons/international-standards
    12. REF-12

      UNESCO and UNICEF. A Human Rights-Based Approach to Education for All. 2007.

      Rights-based planning, equality, participation, accountability and education quality.

      https://unesdoc.unesco.org/ark:/48223/pf0000154861
    13. REF-13

      Education for All Global Monitoring Report Team. Overcoming Inequality: Why Governance Matters — EFA Global Monitoring Report 2009. 2008.

      Governance, finance and unequal educational opportunity.

      https://unesdoc.unesco.org/ark:/48223/pf0000177683
    14. REF-14

      World Bank. World Development Report 2006: Equity and Development. 2005.

      Concepts of unequal opportunity, institutions and equitable public action.

      https://documents.worldbank.org/curated/en/435331468127174418/pdf/322040World0Development0Report02006.pdf
    15. REF-15

      World Bank. Safeguarding Education During Economic Crisis. 2009.

      Risks to budgets, households, participation and long-term human development during economic crisis.

      https://documents1.worldbank.org/curated/en/489131468340200911/pdf/485120WP0Avert10Box338912B01PUBLIC1.pdf
    16. REF-16

      Organisation for Economic Co-operation and Development. Education at a Glance 2010: OECD Indicators. 2010.

      Comparative participation, progression, expenditure and outcomes evidence with system-level metadata.

      https://doi.org/10.1787/eag-2010-en
    17. REF-17

      European Commission. Europe 2020: A Strategy for Smart, Sustainable and Inclusive Growth. 2010.

      Contemporaneous European policy context for education, inclusion, employment and headline indicators.

      https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:52010DC2020
    18. REF-18

      European Commission. Youth on the Move: An Initiative to Unleash the Potential of Young People to Achieve Smart, Sustainable and Inclusive Growth in the European Union. 2010.

      European education, mobility, attainment and youth inclusion policy context.

      https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:52010DC0477
    19. REF-19

      European Commission. A Renewed Commitment to Social Europe: Reinforcing the Open Method of Coordination for Social Protection and Social Inclusion. 2008.

      Social inclusion monitoring, common objectives and context-sensitive indicators.

      https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:52008DC0418
    20. REF-20

      United Nations General Assembly. Resolution 64/250: Assistance to Haiti in the Aftermath of the Recent Earthquake. 2010.

      Contemporaneous recognition of humanitarian and reconstruction needs and national leadership.

      https://undocs.org/A/RES/64/250
    21. REF-21

      United Nations Office for the Coordination of Humanitarian Affairs. Haiti Revised Humanitarian Appeal. 2010.

      Displacement, service disruption and humanitarian education context.

      https://reliefweb.int/report/haiti/haiti-revised-humanitarian-appeal-2010
    22. REF-22

      European Commission. European Union Response to the Earthquake in Haiti. 2010.

      European humanitarian and recovery support, coordination and Haitian ownership.

      https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:52010DC0056
    23. REF-23

      United Nations Economic and Social Council. Principles and Recommendations for Population and Housing Censuses, Revision 2. 2008.

      Official principles for population coverage, definitions, classifications and data quality.

      https://unstats.un.org/unsd/demographic-social/Standards-and-Methods/files/Principles_and_Recommendations/Population-and-Housing-Censuses/Series_M67Rev2-E.pdf
    24. REF-24

      United Nations Statistics Division. Designing Household Survey Samples: Practical Guidelines. 2005.

      Sample design, estimation, precision and non-response guidance.

      https://unstats.un.org/unsd/demographic/sources/surveys/Handbook23June05.pdf
    25. REF-25

      World Education Forum 2015. Incheon Declaration: Education 2030 — Towards Inclusive and Equitable Quality Education and Lifelong Learning for All. 2015.

      Education commitments adopted at Incheon and their contemporaneous institutional status.

      https://unesdoc.unesco.org/ark:/48223/pf0000233137
    26. REF-26

      United Nations General Assembly. Addis Ababa Action Agenda of the Third International Conference on Financing for Development. 2015.

      Adopted global financing framework and relevant principles for domestic public finance and international cooperation.

      https://undocs.org/A/RES/69/313
    27. REF-27

      United Nations General Assembly. Transforming Our World: The 2030 Agenda for Sustainable Development. 2015.

      Agenda adopted on 25 September 2015, including Goal 4 and its education targets, as available at the evidence cut-off.

      https://undocs.org/A/RES/70/1
    28. REF-28

      UNESCO and the World Education Forum 2015 co-convening agencies. Education 2030 Framework for Action. 2015.

      Implementation guidance adopted at the 4 November 2015 high-level meeting; used without later indicator arrangements, statistics or results.

      https://www.unesco.org/en/articles/education-2030-framework-action-be-formally-adopted-and-launched
    29. REF-29

      United Nations Statistical Commission. Report on the Forty-Seventh Session (8–11 March 2016). 2016.

      Contemporaneous Statistical Commission decision treating the proposed global indicator framework as a practical starting point, subject to future refinement.

      https://unstats.un.org/UNSDWebsite/statcom/session_47/documents/2016-34-FinalReport-E.pdf
    30. REF-30

      United Nations General Assembly. New York Declaration for Refugees and Migrants. 2016.

      Adopted commitment relevant to access to education for refugee and migrant children within the cut-off.

      https://undocs.org/A/RES/71/1
    31. REF-31

      United Nations General Assembly. Work of the Statistical Commission Pertaining to the 2030 Agenda for Sustainable Development. 2017.

      Adopted global indicator framework and its status by the evidence cut-off.

      https://undocs.org/A/RES/71/313
    32. REF-32

      World Bank. World Development Report 2018: Learning to Realize Education's Promise. 2018.

      Contemporaneous analysis of the global learning crisis, learning measurement and the alignment of education systems around learning.

      https://www.worldbank.org/en/publication/wdr2018
    33. REF-33

      UNESCO Institute for Statistics. SDG 4 Data Digest 2017: The Quality Factor — Strengthening National Data to Monitor Sustainable Development Goal 4. 2017.

      Contemporaneous guidance on learning assessment data, national capacity, comparability and quality assurance.

      https://uis.unesco.org/sites/default/files/documents/quality-factor-strengthening-national-data-2017-en.pdf