ICEQC-R-2017-11 — Accountability in Education: Purpose, Proportionality and the Limits of Attribution cover

تقرير بحث موضوعي

ICEQC-R-2017-11 — Accountability in Education: Purpose, Proportionality and the Limits of Attribution

A global standards interpretation of educational responsibility, proportionate consequence, causal restraint and remedy

تاريخ النشر
فئة البحث
تفسير المعايير
التقرير النموذج الأصلي
دراسة تفسيرية للمعايير
النطاق الجغرافي
Global
تاريخ انتهاء صلاحية الأدلة
الجهة المسؤولة
مديرية البحوث والسياسات في ICEQC
ICEQC-R-2017-11 — Accountability in Education: Purpose, Proportionality and the Limits of Attribution cover

Publication record

This is the controlled English edition. Evidence and institutional status are stated as at the evidence cut-off date.

Executive summary

Accountability evidence supports improvement when it identifies a bounded educational condition, preserves context and uncertainty, assigns response to an authority capable of remedy, and returns to verify whether learners benefited. It distorts improvement when an available measure becomes the objective, when a rank substitutes for diagnosis, when sanctions attach to unstable differences or when institutions are rewarded for excluding difficult cases. The design should therefore be judged by the educational behaviour it encourages as well as by the accuracy of each released value.

Balanced use requires several forms of evidence. Participation, completion, learning, teaching conditions, resources, learner experience and distribution answer different questions. One should not silently compensate for a failed essential condition in another. External assessment and inspection can provide comparability and independent challenge; institutional evidence can explain mechanism and support rapid correction; learner, family and teacher evidence can show whether provision is usable. Each source should remain within its population and inferential scope.

Consequences should be proportionate to evidence strength and institutional control. A confirmed safety or rights failure requires urgent action. A learning gap may require support, diagnostic inquiry and follow-up. An unstable difference or a new indicator may justify stronger evidence rather than sanction. Institutions should be able to correct factual error and explain context without vetoing a valid finding. Accountability should protect candour: a system that punishes every adverse finding encourages concealment rather than learning.

The practical standard for executive summary concerns conversely, immediate action on a weak finding can waste resources or create new inequity. The institutional consequence follows from whether constructive use requires disciplined movement from claim to evidence, from evidence to judgement, and from judgement to correction. Evaluation contributes to institutional improvement only when a finding is translated into an authorised decision, a feasible change and later verification. A report can be methodologically careful yet have little educational value if its conclusions remain detached from responsibility, resources and learner-facing conditions.

For executive summary, the material distinction is between it is not a general reputation judgement. A contrary reading would overlook that its strength depends on the evidence observed, coverage, selection, definitions and uncertainty. Institutional leaders should separate confirmed finding, plausible explanation and proposed response. The same pattern may have several causes, and the authority to address each cause may lie at institutional, local or national level. An evaluation finding is a bounded statement about a defined population, provision, period and criterion.

A defensible account of executive summary distinguishes the commissioning record should identify the public or educational question, intended users, decision that may follow, applicable standards, required independence and protection of participants. The public account remains incomplete unless it explains how it should not predetermine a favourable conclusion or a preferred intervention. Evaluators need access to contrary and missing evidence; institutions need an opportunity to correct factual error and explain context without acquiring a veto over judgement. Constructive use begins before fieldwork.

Review of executive summary is credible only where it explains a recurring educational weakness may require diagnostic enquiry and a bounded improvement trial. For the learners concerned, the decisive consideration is whether a promising practice may merit continuation or wider examination, but a favourable local association does not establish that extension will work elsewhere. Findings should be classified by consequence and evidentiary confidence. A confirmed failure of safety, legality or learner protection requires immediate authorised action even where causal explanation remains incomplete.

Evidence concerning executive summary should establish additional meetings, policies, training events or purchases are activities; they do not establish changed teaching, access, support or learning. This matters because a response record should name the finding, responsible body, affected population, authorised resource, intended operating condition, delivery date, evidence and review decision. Dependencies outside institutional authority should be escalated rather than restated as local obligations. Institutional response should address the condition identified.

In assessing executive summary, authorities must determine participation does not make every account equally probative. The resulting interpretation should show why the evaluation should record who was reached, who was absent, how candour was protected and how disagreement affected the final claim. Participation improves both interpretation and feasibility. Learners can show whether intended provision was usable; teachers can identify workload, materials and instructional constraints; families and communities can reveal access barriers; leaders can explain finance and authority.

Evidence concerning executive summary should establish response design should ask who can use the ordinary arrangement, who bears additional cost and whether correction in one area transfers disadvantage elsewhere. A proportionate conclusion must also recognise that equity should be examined within every major finding. An aggregate improvement can coexist with deterioration for a smaller or less visible population. Data may omit remote learners, displaced populations, persons with disabilities or those outside recognised institutions.

A defensible account of executive summary distinguishes a facility recommendation should examine capital, maintenance and accessibility. The resulting interpretation should show why a data recommendation should examine definitions, collection burden, coverage and authorised use. Recommendations that ignore the conditions of delivery can turn an institutional limitation into an individual performance judgement. Capacity is part of causal reasoning. A teacher-facing recommendation should examine competence, workload, materials, leadership and time.

The practical standard for executive summary concerns where definitions, assessment or population changed, the institution should not manufacture a trend. This matters because it may still report current status and implementation evidence, with limitations stated. Follow-up should return to the original population and criterion. It should distinguish whether the response was initiated, completed, operating, reaching the intended population and associated with the intended condition.

Evidence concerning executive summary should establish it should not reward polished documentation over candid identification of weakness. The evidence must therefore clarify how a strong institution can explain what it did not know, which conditions remained outside its authority and why a decision was revised after evidence. External scrutiny has a constructive role when it tests the reasoning chain, treatment of adverse evidence, proportionality of response and credibility of follow-up.

For executive summary, the material distinction is between this record preserves institutional learning without confusing self-report with independent assurance. The resulting interpretation should show why the central recommendation is a finding-to-improvement record. It joins the evaluative claim to its criterion, population, evidence, limitation, responsible authority, action, resource, milestone, equity test, operating verification and review decision.

Key findings

    Scope and method

    Part I

    From finding to authorised response

    1

    What national averages conceal

    The practical standard for what national averages conceal concerns the measure should preserve both the observed level and the distribution relevant to the claim. The institutional consequence follows from whether a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns national participation or attainment. Its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal.[REF-01] [REF-07]

    Evidence concerning what national averages conceal should establish it should require examination of definitions, timing, migration, duplication and non-response. The resulting interpretation should show why the resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is the weighted experience of the population included in the denominator. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source.[REF-02] [REF-03]

    Evidence concerning what national averages conceal should establish reference points therefore need substantive meaning. The resulting interpretation should show why where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: distribution within countries, not a league table between them. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups.[REF-04] [REF-06]

    what national averages conceal cannot be judged without identifying coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. The institutional consequence follows from whether if direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include remote rural learners, urban informal settlements and displaced populations. These populations may be missing not only from good outcomes but from the denominator itself.[REF-08] [REF-09]

    For what national averages conceal, the material distinction is between it should not be used to rank schools or communities where differences in population and opportunity to learn are uncontrolled. For the learners concerned, the decisive consideration is whether monitoring after action must preserve the original baseline and follow both reach and outcome. Improvement in an average does not demonstrate that the intended group benefited; participation, intensity and learner experience should be checked. The final public statement should identify remaining uncertainty and the next evidentiary step rather than convert a partial result into assurance. The decision record should connect the finding to a responsible authority, available intervention and review date. An indicator is useful when it can alter service location, staffing, language support, accessibility, household assistance or another defined condition.[REF-10] [REF-12]

    2

    Marginalisation as accumulated disadvantage

    A defensible account of marginalisation as accumulated disadvantage distinguishes a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. For the learners concerned, the decisive consideration is whether the report should state the educational consequence before choosing a gap, ratio, threshold or rank. Marginalisation as accumulated disadvantage should be approached as a defined measurement problem. The substantive interest is distance from a socially secured educational minimum, observed for a population and period that are stated before calculation. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-02] [REF-03]

    marginalisation as accumulated disadvantage cannot be judged without identifying school returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. A proportionate conclusion must also recognise that reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use the joint effect of exclusion, weak provision and adverse social conditions. This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined.[REF-04] [REF-06]

    Review of marginalisation as accumulated disadvantage is credible only where it explains policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. A proportionate conclusion must also recognise that the governing caution is that multiple indicators read together over the learner's course. Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. Percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. The comparison should identify the reference category but avoid presenting it as a natural norm.[REF-08] [REF-09]

    marginalisation as accumulated disadvantage requires a decision about where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. The resulting interpretation should show why where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for children facing poverty, gender disadvantage, disability or minority status. Their circumstances may alter access to enumeration, classification and the service being measured. The review should record non-response, unknown status and excluded locations separately. Combining unknown observations with the majority group biases both estimates and obscures the weakness.[REF-10] [REF-12]

    For marginalisation as accumulated disadvantage, the material distinction is between targets should specify the population expected to benefit and guard against gains achieved by concentrating on learners nearest a threshold. The institutional consequence follows from whether accountability lies in changed opportunity, not in the favourable movement of an indicator alone. Policy use should begin with a question that the competent body can answer. The evidence may justify further enquiry, immediate removal of a barrier, redistribution of resources or evaluation of an existing measure. These are different decisions and require different certainty. Urgent protection need not await a perfect estimate where credible evidence shows serious exclusion, but long-term allocation should be reviewed as coverage improves.[REF-13] [REF-14]

    3

    From monitoring commitment to decision

    from monitoring commitment to decision requires a decision about a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. A proportionate conclusion must also recognise that the report should state the educational consequence before choosing a gap, ratio, threshold or rank. For from monitoring commitment to decision, the first requirement is conceptual clarity. The study seeks evidence on evidence capable of changing resource or service decisions; it does not infer a learner's circumstances from a national or regional mean. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-04] [REF-06]

    Institutional action on from monitoring commitment to decision should be tested against the estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. For the learners concerned, the decisive consideration is whether when several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on a link between observed disparity, responsible body and remedy. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression.[REF-08] [REF-09]

    Comparative interpretation of from monitoring commitment to decision depends upon results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. The resulting interpretation should show why if a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: indicators selected for action rather than visibility.[REF-10] [REF-12]

    In assessing from monitoring commitment to decision, authorities must determine analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. The institutional consequence follows from whether missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. An adequate equity account asks whether groups absent from routine plans and budget classifications are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation.[REF-13] [REF-14]

    A defensible account of from monitoring commitment to decision distinguishes communities should be able to question both the category and the conclusion drawn from it. The institutional consequence follows from whether if a measure creates incentives to exclude difficult cases, narrow the denominator or reclassify non-completion, an independent check is required. The report should recognise such behaviour as a measurement risk without assuming misconduct in every discrepancy. A proportionate response records who acts, the condition to be changed and the evidence that will show whether reach was equitable. Local interpretation is valuable because national classifications cannot capture every barrier; it should operate within common definitions sufficient for aggregation.[REF-15] [REF-16]

    4

    Defining the population entitled to education

    Institutional action on defining the population entitled to education should be tested against without that purpose, disaggregation can multiply figures without improving public judgement. The evidence must therefore clarify how the measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of resident and temporarily absent learners within the relevant age or programme group begins by naming the decision the evidence may inform.[REF-08] [REF-09]

    defining the population entitled to education requires a decision about these are substantive attributes because they determine who can appear in the evidence. A contrary reading would overlook that a figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is a denominator consistent with the right, level and reference period under review. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules.[REF-10] [REF-12]

    In assessing defining the population entitled to education, authorities must determine a responsible commentary distinguishes observation, calculation and interpretation. The public account remains incomplete unless it explains how it states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference. Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that census, survey and administrative estimates reconciled openly.[REF-13] [REF-14]

    The central question in defining the population entitled to education is it should involve statistical judgement, legal safeguards and knowledge of the affected community. The public account remains incomplete unless it explains how suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. For unregistered residents, migrants and displaced children, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical.[REF-15] [REF-16]

    Public responsibility for defining the population entitled to education begins with this makes the indicator a means of scrutiny rather than a decorative measure of concern. The resulting interpretation should show why public accountability requires a concise explanation of the result and its boundary. Readers should be able to see who was counted, who was not, what period the figure covers, how large the underlying population is and which comparisons are justified. A policy conclusion should be no broader than that evidence. Subsequent reports should distinguish real change from late reporting, revised population estimates and altered definitions. Where a disparity remains, the responsible body should state whether the obstacle is knowledge, authority, resources or implementation.[REF-17] [REF-19]

    Part II

    Designing improvement action

    5

    Age, grade and programme populations

    age, grade and programme populations cannot be judged without identifying the measure should preserve both the observed level and the distribution relevant to the claim. The public account remains incomplete unless it explains how a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of age, grade and programme populations lies in making unequal educational experience observable. Here, the relevant phenomenon is age-specific, grade-specific and programme-specific participation, not the administrative convenience of the available categories.[REF-10] [REF-12]

    Public responsibility for age, grade and programme populations begins with a discrepancy is not resolved by selecting the more favourable source. The public account remains incomplete unless it explains how it should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is separate denominators for different educational questions. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs.[REF-13] [REF-14]

    For age, grade and programme populations, the material distinction is between explanation requires evidence on institutions, resources, households and prior conditions. A proportionate conclusion must also recognise that interpretation follows this limitation: exact age, school age and enrolled population never substituted silently. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose.[REF-15] [REF-16]

    age, grade and programme populations cannot be judged without identifying if direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. The institutional consequence follows from whether confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include over-age entrants and learners repeating grades. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete.[REF-17] [REF-19]

    6

    Population movement and disrupted residence

    Evidence concerning population movement and disrupted residence should establish its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The public account remains incomplete unless it explains how the measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns education status amid migration, displacement and return.[REF-13] [REF-14]

    Review of population movement and disrupted residence is credible only where it explains school returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. A contrary reading would overlook that reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use dated location and residence rules with sensitivity to mobility. This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined.[REF-15] [REF-16]

    population movement and disrupted residence cannot be judged without identifying the comparison should identify the reference category but avoid presenting it as a natural norm. For the learners concerned, the decisive consideration is whether policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that origin, current location and service responsibility distinguished. Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. Percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation.[REF-17] [REF-19]

    The practical standard for population movement and disrupted residence concerns their circumstances may alter access to enumeration, classification and the service being measured. A proportionate conclusion must also recognise that the review should record non-response, unknown status and excluded locations separately. Combining unknown observations with the majority group biases both estimates and obscures the weakness. Where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for families affected by conflict, disaster and seasonal movement.[REF-20] [REF-21]

    7

    Entry at the official starting age

    For entry at the official starting age, the material distinction is between the measure should preserve both the observed level and the distribution relevant to the claim. For the learners concerned, the decisive consideration is whether a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. Entry at the official starting age should be approached as a defined measurement problem. The substantive interest is timely admission to the first grade of primary education, observed for a population and period that are stated before calculation.[REF-15] [REF-16]

    The practical standard for entry at the official starting age concerns the estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. A proportionate conclusion must also recognise that when several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on new entrants of official age relative to the corresponding population. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression.[REF-17] [REF-19]

    The practical standard for entry at the official starting age concerns if a conclusion changes under a reasonable specification, that instability is part of the finding. The evidence must therefore clarify how national averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: late entry examined alongside non-entry. Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved.[REF-20] [REF-21]

    Evidence concerning entry at the official starting age should establish missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. A contrary reading would overlook that an adequate equity account asks whether children facing fees, distance, disability or documentation barriers are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination.[REF-23] [REF-24]

    8

    Attendance beyond enrolment

    Public responsibility for attendance beyond enrolment begins with the measure should preserve both the observed level and the distribution relevant to the claim. The resulting interpretation should show why a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. For attendance beyond enrolment, the first requirement is conceptual clarity. The study seeks evidence on actual participation during a stated recent period; it does not infer a learner's circumstances from a national or regional mean.[REF-17] [REF-19]

    Evidence concerning attendance beyond enrolment should establish a survey estimate should disclose weights and uncertainty. For the learners concerned, the decisive consideration is whether a census figure should disclose enumeration rules. These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is presence measured independently of registration status. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions.[REF-20] [REF-21]

    A defensible account of attendance beyond enrolment distinguishes a responsible commentary distinguishes observation, calculation and interpretation. The resulting interpretation should show why it states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference. Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that frequency, season and reasons for absence retained.[REF-23] [REF-24]

    Public responsibility for attendance beyond enrolment begins with the public report can still state that a disparity was examined, whether action is required and which body will monitor it. The evidence must therefore clarify how for working children, carers and learners affected by illness, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community. Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells.[REF-01] [REF-07]

    Part III

    Verification and institutional learning

    9

    Out-of-school status

    In assessing out-of-school status, authorities must determine a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The evidence must therefore clarify how the report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of children in the relevant age group not participating at the defined level begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-20] [REF-21]

    For out-of-school status, the material distinction is between numerator, denominator, reference date, unit and exclusions should appear together. The institutional consequence follows from whether if a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is a transparent residual from compatible population and participation concepts.[REF-23] [REF-24]

    Evidence concerning out-of-school status should establish the comparison should show the level for each group as well as any ratio or gap. The institutional consequence follows from whether a ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: never-enrolled and formerly enrolled children separated.[REF-01] [REF-07]

    A defensible account of out-of-school status distinguishes these populations may be missing not only from good outcomes but from the denominator itself. The institutional consequence follows from whether coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include children invisible to school registers.[REF-02] [REF-03]

    10

    Repetition and grade survival

    repetition and grade survival cannot be judged without identifying here, the relevant phenomenon is movement through grades without avoidable delay or exit, not the administrative convenience of the available categories. For the learners concerned, the decisive consideration is whether the measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of repetition and grade survival lies in making unequal educational experience observable.[REF-23] [REF-24]

    repetition and grade survival cannot be judged without identifying this formulation requires the reporting body to preserve the population base and the observation period beside the result. This matters because counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined. School returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. Reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use cohort or reconstructed-cohort evidence with explicit assumptions.[REF-01] [REF-07]

    A defensible account of repetition and grade survival distinguishes percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. For the learners concerned, the decisive consideration is whether the comparison should identify the reference category but avoid presenting it as a natural norm. Policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that repeaters distinguished from re-entrants and transfers. Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses.[REF-02] [REF-03]

    The practical standard for repetition and grade survival concerns the review should record non-response, unknown status and excluded locations separately. The public account remains incomplete unless it explains how combining unknown observations with the majority group biases both estimates and obscures the weakness. Where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for learners in overcrowded or intermittently operating schools. Their circumstances may alter access to enumeration, classification and the service being measured.[REF-04] [REF-06]

    11

    Transition between education levels

    A defensible account of transition between education levels distinguishes its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The institutional consequence follows from whether the measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns entry to the next level after completion of the preceding one.[REF-01] [REF-07]

    The practical standard for transition between education levels concerns the estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. A proportionate conclusion must also recognise that when several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on matched completion and new-entry populations over coherent periods. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression.[REF-02] [REF-03]

    For transition between education levels, the material distinction is between nor should a group estimate be read as a description of every member. A contrary reading would overlook that within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: capacity constraints distinguished from learner attainment. Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution.[REF-04] [REF-06]

    For transition between education levels, the material distinction is between missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. The evidence must therefore clarify how an adequate equity account asks whether rural learners and those unable to relocate are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination.[REF-08] [REF-09]

    12

    Completion and educational entitlement

    In assessing completion and educational entitlement, authorities must determine a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The evidence must therefore clarify how the report should state the educational consequence before choosing a gap, ratio, threshold or rank. Completion and educational entitlement should be approached as a defined measurement problem. The substantive interest is finishing the final grade or meeting recognised programme requirements, observed for a population and period that are stated before calculation. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-02] [REF-03]

    Comparative interpretation of completion and educational entitlement depends upon a national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. The resulting interpretation should show why a survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is completion defined separately from sitting or passing an examination. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions.[REF-04] [REF-06]

    A defensible account of completion and educational entitlement distinguishes apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. For the learners concerned, the decisive consideration is whether use of the indicator is bounded by the principle that late completion and alternative pathways reported. A responsible commentary distinguishes observation, calculation and interpretation. It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference.[REF-08] [REF-09]

    Evidence concerning completion and educational entitlement should establish suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. A contrary reading would overlook that the public report can still state that a disparity was examined, whether action is required and which body will monitor it. For over-age learners and people returning after interruption, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community.[REF-10] [REF-12]

    Part I

    Purpose and interpretive frame

    1

    What national averages conceal

    The practical standard for what national averages conceal concerns the measure should preserve both the observed level and the distribution relevant to the claim. The institutional consequence follows from whether a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns national participation or attainment. Its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal.[REF-01] [REF-07]

    Evidence concerning what national averages conceal should establish it should require examination of definitions, timing, migration, duplication and non-response. The resulting interpretation should show why the resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is the weighted experience of the population included in the denominator. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source.[REF-02] [REF-03]

    Evidence concerning what national averages conceal should establish reference points therefore need substantive meaning. The resulting interpretation should show why where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: distribution within countries, not a league table between them. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups.[REF-04] [REF-06]

    what national averages conceal cannot be judged without identifying coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. The institutional consequence follows from whether if direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include remote rural learners, urban informal settlements and displaced populations. These populations may be missing not only from good outcomes but from the denominator itself.[REF-08] [REF-09]

    For what national averages conceal, the material distinction is between it should not be used to rank schools or communities where differences in population and opportunity to learn are uncontrolled. For the learners concerned, the decisive consideration is whether monitoring after action must preserve the original baseline and follow both reach and outcome. Improvement in an average does not demonstrate that the intended group benefited; participation, intensity and learner experience should be checked. The final public statement should identify remaining uncertainty and the next evidentiary step rather than convert a partial result into assurance. The decision record should connect the finding to a responsible authority, available intervention and review date. An indicator is useful when it can alter service location, staffing, language support, accessibility, household assistance or another defined condition.[REF-10] [REF-12]

    2

    Marginalisation as accumulated disadvantage

    A defensible account of marginalisation as accumulated disadvantage distinguishes a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. For the learners concerned, the decisive consideration is whether the report should state the educational consequence before choosing a gap, ratio, threshold or rank. Marginalisation as accumulated disadvantage should be approached as a defined measurement problem. The substantive interest is distance from a socially secured educational minimum, observed for a population and period that are stated before calculation. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-02] [REF-03]

    marginalisation as accumulated disadvantage cannot be judged without identifying school returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. A proportionate conclusion must also recognise that reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use the joint effect of exclusion, weak provision and adverse social conditions. This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined.[REF-04] [REF-06]

    Review of marginalisation as accumulated disadvantage is credible only where it explains policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. A proportionate conclusion must also recognise that the governing caution is that multiple indicators read together over the learner's course. Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. Percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. The comparison should identify the reference category but avoid presenting it as a natural norm.[REF-08] [REF-09]

    marginalisation as accumulated disadvantage requires a decision about where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. The resulting interpretation should show why where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for children facing poverty, gender disadvantage, disability or minority status. Their circumstances may alter access to enumeration, classification and the service being measured. The review should record non-response, unknown status and excluded locations separately. Combining unknown observations with the majority group biases both estimates and obscures the weakness.[REF-10] [REF-12]

    For marginalisation as accumulated disadvantage, the material distinction is between targets should specify the population expected to benefit and guard against gains achieved by concentrating on learners nearest a threshold. The institutional consequence follows from whether accountability lies in changed opportunity, not in the favourable movement of an indicator alone. Policy use should begin with a question that the competent body can answer. The evidence may justify further enquiry, immediate removal of a barrier, redistribution of resources or evaluation of an existing measure. These are different decisions and require different certainty. Urgent protection need not await a perfect estimate where credible evidence shows serious exclusion, but long-term allocation should be reviewed as coverage improves.[REF-13] [REF-14]

    3

    From monitoring commitment to decision

    from monitoring commitment to decision requires a decision about a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. A proportionate conclusion must also recognise that the report should state the educational consequence before choosing a gap, ratio, threshold or rank. For from monitoring commitment to decision, the first requirement is conceptual clarity. The study seeks evidence on evidence capable of changing resource or service decisions; it does not infer a learner's circumstances from a national or regional mean. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-04] [REF-06]

    Institutional action on from monitoring commitment to decision should be tested against the estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. For the learners concerned, the decisive consideration is whether when several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on a link between observed disparity, responsible body and remedy. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression.[REF-08] [REF-09]

    Comparative interpretation of from monitoring commitment to decision depends upon results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. The resulting interpretation should show why if a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: indicators selected for action rather than visibility.[REF-10] [REF-12]

    In assessing from monitoring commitment to decision, authorities must determine analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. The institutional consequence follows from whether missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. An adequate equity account asks whether groups absent from routine plans and budget classifications are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation.[REF-13] [REF-14]

    A defensible account of from monitoring commitment to decision distinguishes communities should be able to question both the category and the conclusion drawn from it. The institutional consequence follows from whether if a measure creates incentives to exclude difficult cases, narrow the denominator or reclassify non-completion, an independent check is required. The report should recognise such behaviour as a measurement risk without assuming misconduct in every discrepancy. A proportionate response records who acts, the condition to be changed and the evidence that will show whether reach was equitable. Local interpretation is valuable because national classifications cannot capture every barrier; it should operate within common definitions sufficient for aggregation.[REF-15] [REF-16]

    Part II

    Population and denominator

    4

    Defining the population entitled to education

    Institutional action on defining the population entitled to education should be tested against without that purpose, disaggregation can multiply figures without improving public judgement. The evidence must therefore clarify how the measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of resident and temporarily absent learners within the relevant age or programme group begins by naming the decision the evidence may inform.[REF-08] [REF-09]

    defining the population entitled to education requires a decision about these are substantive attributes because they determine who can appear in the evidence. A contrary reading would overlook that a figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is a denominator consistent with the right, level and reference period under review. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules.[REF-10] [REF-12]

    In assessing defining the population entitled to education, authorities must determine a responsible commentary distinguishes observation, calculation and interpretation. The public account remains incomplete unless it explains how it states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference. Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that census, survey and administrative estimates reconciled openly.[REF-13] [REF-14]

    The central question in defining the population entitled to education is it should involve statistical judgement, legal safeguards and knowledge of the affected community. The public account remains incomplete unless it explains how suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. For unregistered residents, migrants and displaced children, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical.[REF-15] [REF-16]

    Public responsibility for defining the population entitled to education begins with this makes the indicator a means of scrutiny rather than a decorative measure of concern. The resulting interpretation should show why public accountability requires a concise explanation of the result and its boundary. Readers should be able to see who was counted, who was not, what period the figure covers, how large the underlying population is and which comparisons are justified. A policy conclusion should be no broader than that evidence. Subsequent reports should distinguish real change from late reporting, revised population estimates and altered definitions. Where a disparity remains, the responsible body should state whether the obstacle is knowledge, authority, resources or implementation.[REF-17] [REF-19]

    5

    Age, grade and programme populations

    age, grade and programme populations cannot be judged without identifying the measure should preserve both the observed level and the distribution relevant to the claim. The public account remains incomplete unless it explains how a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of age, grade and programme populations lies in making unequal educational experience observable. Here, the relevant phenomenon is age-specific, grade-specific and programme-specific participation, not the administrative convenience of the available categories.[REF-10] [REF-12]

    Public responsibility for age, grade and programme populations begins with a discrepancy is not resolved by selecting the more favourable source. The public account remains incomplete unless it explains how it should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is separate denominators for different educational questions. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs.[REF-13] [REF-14]

    For age, grade and programme populations, the material distinction is between explanation requires evidence on institutions, resources, households and prior conditions. A proportionate conclusion must also recognise that interpretation follows this limitation: exact age, school age and enrolled population never substituted silently. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose.[REF-15] [REF-16]

    age, grade and programme populations cannot be judged without identifying if direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. The institutional consequence follows from whether confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include over-age entrants and learners repeating grades. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete.[REF-17] [REF-19]

    6

    Population movement and disrupted residence

    Evidence concerning population movement and disrupted residence should establish its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The public account remains incomplete unless it explains how the measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns education status amid migration, displacement and return.[REF-13] [REF-14]

    Review of population movement and disrupted residence is credible only where it explains school returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. A contrary reading would overlook that reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use dated location and residence rules with sensitivity to mobility. This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined.[REF-15] [REF-16]

    population movement and disrupted residence cannot be judged without identifying the comparison should identify the reference category but avoid presenting it as a natural norm. For the learners concerned, the decisive consideration is whether policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that origin, current location and service responsibility distinguished. Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. Percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation.[REF-17] [REF-19]

    The practical standard for population movement and disrupted residence concerns their circumstances may alter access to enumeration, classification and the service being measured. A proportionate conclusion must also recognise that the review should record non-response, unknown status and excluded locations separately. Combining unknown observations with the majority group biases both estimates and obscures the weakness. Where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for families affected by conflict, disaster and seasonal movement.[REF-20] [REF-21]

    Part III

    Access and participation

    7

    Entry at the official starting age

    For entry at the official starting age, the material distinction is between the measure should preserve both the observed level and the distribution relevant to the claim. For the learners concerned, the decisive consideration is whether a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. Entry at the official starting age should be approached as a defined measurement problem. The substantive interest is timely admission to the first grade of primary education, observed for a population and period that are stated before calculation.[REF-15] [REF-16]

    The practical standard for entry at the official starting age concerns the estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. A proportionate conclusion must also recognise that when several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on new entrants of official age relative to the corresponding population. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression.[REF-17] [REF-19]

    The practical standard for entry at the official starting age concerns if a conclusion changes under a reasonable specification, that instability is part of the finding. The evidence must therefore clarify how national averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: late entry examined alongside non-entry. Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved.[REF-20] [REF-21]

    Evidence concerning entry at the official starting age should establish missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. A contrary reading would overlook that an adequate equity account asks whether children facing fees, distance, disability or documentation barriers are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination.[REF-23] [REF-24]

    8

    Attendance beyond enrolment

    Public responsibility for attendance beyond enrolment begins with the measure should preserve both the observed level and the distribution relevant to the claim. The resulting interpretation should show why a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. For attendance beyond enrolment, the first requirement is conceptual clarity. The study seeks evidence on actual participation during a stated recent period; it does not infer a learner's circumstances from a national or regional mean.[REF-17] [REF-19]

    Evidence concerning attendance beyond enrolment should establish a survey estimate should disclose weights and uncertainty. For the learners concerned, the decisive consideration is whether a census figure should disclose enumeration rules. These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is presence measured independently of registration status. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions.[REF-20] [REF-21]

    A defensible account of attendance beyond enrolment distinguishes a responsible commentary distinguishes observation, calculation and interpretation. The resulting interpretation should show why it states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference. Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that frequency, season and reasons for absence retained.[REF-23] [REF-24]

    Public responsibility for attendance beyond enrolment begins with the public report can still state that a disparity was examined, whether action is required and which body will monitor it. The evidence must therefore clarify how for working children, carers and learners affected by illness, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community. Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells.[REF-01] [REF-07]

    9

    Out-of-school status

    In assessing out-of-school status, authorities must determine a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The evidence must therefore clarify how the report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of children in the relevant age group not participating at the defined level begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-20] [REF-21]

    For out-of-school status, the material distinction is between numerator, denominator, reference date, unit and exclusions should appear together. The institutional consequence follows from whether if a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is a transparent residual from compatible population and participation concepts.[REF-23] [REF-24]

    Evidence concerning out-of-school status should establish the comparison should show the level for each group as well as any ratio or gap. The institutional consequence follows from whether a ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: never-enrolled and formerly enrolled children separated.[REF-01] [REF-07]

    A defensible account of out-of-school status distinguishes these populations may be missing not only from good outcomes but from the denominator itself. The institutional consequence follows from whether coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include children invisible to school registers.[REF-02] [REF-03]

    Part IV

    Progression and completion

    10

    Repetition and grade survival

    repetition and grade survival cannot be judged without identifying here, the relevant phenomenon is movement through grades without avoidable delay or exit, not the administrative convenience of the available categories. For the learners concerned, the decisive consideration is whether the measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of repetition and grade survival lies in making unequal educational experience observable.[REF-23] [REF-24]

    repetition and grade survival cannot be judged without identifying this formulation requires the reporting body to preserve the population base and the observation period beside the result. This matters because counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined. School returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. Reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use cohort or reconstructed-cohort evidence with explicit assumptions.[REF-01] [REF-07]

    A defensible account of repetition and grade survival distinguishes percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. For the learners concerned, the decisive consideration is whether the comparison should identify the reference category but avoid presenting it as a natural norm. Policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that repeaters distinguished from re-entrants and transfers. Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses.[REF-02] [REF-03]

    The practical standard for repetition and grade survival concerns the review should record non-response, unknown status and excluded locations separately. The public account remains incomplete unless it explains how combining unknown observations with the majority group biases both estimates and obscures the weakness. Where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for learners in overcrowded or intermittently operating schools. Their circumstances may alter access to enumeration, classification and the service being measured.[REF-04] [REF-06]

    11

    Transition between education levels

    A defensible account of transition between education levels distinguishes its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The institutional consequence follows from whether the measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns entry to the next level after completion of the preceding one.[REF-01] [REF-07]

    The practical standard for transition between education levels concerns the estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. A proportionate conclusion must also recognise that when several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on matched completion and new-entry populations over coherent periods. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression.[REF-02] [REF-03]

    For transition between education levels, the material distinction is between nor should a group estimate be read as a description of every member. A contrary reading would overlook that within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: capacity constraints distinguished from learner attainment. Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution.[REF-04] [REF-06]

    For transition between education levels, the material distinction is between missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. The evidence must therefore clarify how an adequate equity account asks whether rural learners and those unable to relocate are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination.[REF-08] [REF-09]

    12

    Completion and educational entitlement

    In assessing completion and educational entitlement, authorities must determine a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The evidence must therefore clarify how the report should state the educational consequence before choosing a gap, ratio, threshold or rank. Completion and educational entitlement should be approached as a defined measurement problem. The substantive interest is finishing the final grade or meeting recognised programme requirements, observed for a population and period that are stated before calculation. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-02] [REF-03]

    Comparative interpretation of completion and educational entitlement depends upon a national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. The resulting interpretation should show why a survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is completion defined separately from sitting or passing an examination. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions.[REF-04] [REF-06]

    A defensible account of completion and educational entitlement distinguishes apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. For the learners concerned, the decisive consideration is whether use of the indicator is bounded by the principle that late completion and alternative pathways reported. A responsible commentary distinguishes observation, calculation and interpretation. It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference.[REF-08] [REF-09]

    Evidence concerning completion and educational entitlement should establish suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. A contrary reading would overlook that the public report can still state that a disparity was examined, whether action is required and which body will monitor it. For over-age learners and people returning after interruption, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community.[REF-10] [REF-12]

    Part V

    Learning and assessment

    13

    Minimum learning outcomes

    Public responsibility for minimum learning outcomes begins with the measure should preserve both the observed level and the distribution relevant to the claim. A proportionate conclusion must also recognise that a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. For minimum learning outcomes, the first requirement is conceptual clarity. The study seeks evidence on demonstrated knowledge or skill against a declared domain; it does not infer a learner's circumstances from a national or regional mean.[REF-04] [REF-06]

    minimum learning outcomes requires a decision about it should require examination of definitions, timing, migration, duplication and non-response. The institutional consequence follows from whether the resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is assessment evidence whose population and conditions are known. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source.[REF-08] [REF-09]

    minimum learning outcomes cannot be judged without identifying reference points therefore need substantive meaning. The evidence must therefore clarify how where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: results interpreted with opportunity to learn and participation. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups.[REF-10] [REF-12]

    minimum learning outcomes cannot be judged without identifying protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The public account remains incomplete unless it explains how the distributional review must deliberately include learners taught in an unfamiliar language. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. Confidentiality is essential, especially where identity or status creates risk.[REF-13] [REF-14]

    14

    Assessment participation

    assessment participation cannot be judged without identifying the report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public account remains incomplete unless it explains how a distribution-sensitive account of who was eligible, present, absent and excluded from testing begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage.[REF-08] [REF-09]

    assessment participation requires a decision about counts reveal scale; rates permit comparison; neither is sufficient alone. A proportionate conclusion must also recognise that source coverage must be tested before sources are combined. School returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. Reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use a participation profile accompanying every result distribution. This formulation requires the reporting body to preserve the population base and the observation period beside the result.[REF-10] [REF-12]

    Institutional action on assessment participation should be tested against the comparison should identify the reference category but avoid presenting it as a natural norm. The resulting interpretation should show why policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that non-participation never treated as low attainment or ignored. Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. Percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation.[REF-13] [REF-14]

    The practical standard for assessment participation concerns the review should record non-response, unknown status and excluded locations separately. The public account remains incomplete unless it explains how combining unknown observations with the majority group biases both estimates and obscures the weakness. Where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for learners with disabilities and remote candidates. Their circumstances may alter access to enumeration, classification and the service being measured.[REF-15] [REF-16]

    15

    Distribution of achievement

    distribution of achievement requires a decision about a proportionate conclusion must also recognise that the report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public account remains incomplete unless it explains how the public value of distribution of achievement lies in making unequal educational experience observable. Here, the relevant phenomenon is variation across the full score or proficiency distribution, not the administrative convenience of the available categories. The measure should preserve both the observed level and the distribution relevant to the claim. Comparative interpretation of distribution of achievement depends upon a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage.[REF-10] [REF-12]

    Evidence concerning distribution of achievement should establish when several sources exist, consistency is evidence to consider, not proof that common error is absent. A proportionate conclusion must also recognise that a defensible statistic would be based on percentiles and threshold shares beside a mean. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression. The estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality.[REF-13] [REF-14]

    For distribution of achievement, the material distinction is between results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. The public account remains incomplete unless it explains how if a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: uncertainty and scale properties stated.[REF-15] [REF-16]

    Review of distribution of achievement is credible only where it explains field arrangements need relevant languages, accessible formats and safe participation. This matters because analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. Missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. An adequate equity account asks whether learners concentrated below a minimum proficiency threshold are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population.[REF-17] [REF-19]

    Part VI

    Gender and household resources

    16

    Gender parity and its limits

    gender parity and its limits requires a decision about a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The public account remains incomplete unless it explains how the report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns differences between girls and boys in access, progression and learning. Its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-13] [REF-14]

    In assessing gender parity and its limits, authorities must determine its metadata should travel with every published value. A contrary reading would overlook that at minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is female-to-male ratios read with levels and absolute gaps.[REF-15] [REF-16]

    Evidence concerning gender parity and its limits should establish a responsible commentary distinguishes observation, calculation and interpretation. The public account remains incomplete unless it explains how it states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference. Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that parity not confused with adequacy for either group.[REF-17] [REF-19]

    For gender parity and its limits, the material distinction is between the public report can still state that a disparity was examined, whether action is required and which body will monitor it. The evidence must therefore clarify how for girls in poor rural households and boys exposed to hazardous work, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community. Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells.[REF-20] [REF-21]

    17

    Household wealth gradients

    A defensible account of household wealth gradients distinguishes household wealth gradients should be approached as a defined measurement problem. A contrary reading would overlook that the substantive interest is education outcomes across relative household resource groups, observed for a population and period that are stated before calculation. Public responsibility for household wealth gradients begins with the measure should preserve both the observed level and the distribution relevant to the claim. A contrary reading would overlook that a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank.[REF-15] [REF-16]

    household wealth gradients cannot be judged without identifying numerator, denominator, reference date, unit and exclusions should appear together. For the learners concerned, the decisive consideration is whether if a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is a documented asset or consumption classification within each setting.[REF-17] [REF-19]

    Comparative interpretation of household wealth gradients depends upon the comparison should show the level for each group as well as any ratio or gap. The institutional consequence follows from whether a ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: wealth ranks not assumed equivalent across countries or time.[REF-20] [REF-21]

    household wealth gradients cannot be judged without identifying coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. A contrary reading would overlook that if direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include children in the poorest quintile and those near classification boundaries. These populations may be missing not only from good outcomes but from the denominator itself.[REF-23] [REF-24]

    18

    Costs borne by households

    costs borne by households cannot be judged without identifying the report should state the educational consequence before choosing a gap, ratio, threshold or rank. For the learners concerned, the decisive consideration is whether for costs borne by households, the first requirement is conceptual clarity. The study seeks evidence on fees, materials, transport, clothing and foregone labour; it does not infer a learner's circumstances from a national or regional mean. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage.[REF-17] [REF-19]

    costs borne by households cannot be judged without identifying school returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. The resulting interpretation should show why reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use participation read against direct and indirect education costs. This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined.[REF-20] [REF-21]

    Institutional action on costs borne by households should be tested against percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. The public account remains incomplete unless it explains how the comparison should identify the reference category but avoid presenting it as a natural norm. Policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that nominal fee abolition checked against remaining expenditure. Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses.[REF-23] [REF-24]

    Review of costs borne by households is credible only where it explains where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. The resulting interpretation should show why where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for large families and households hit by economic crisis. Their circumstances may alter access to enumeration, classification and the service being measured. The review should record non-response, unknown status and excluded locations separately. Combining unknown observations with the majority group biases both estimates and obscures the weakness.[REF-01] [REF-07]

    Part VII

    Place and service geography

    19

    Rural and urban residence

    rural and urban residence requires a decision about the measure should preserve both the observed level and the distribution relevant to the claim. The evidence must therefore clarify how a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of education outcomes by a declared settlement classification begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement.[REF-20] [REF-21]

    Comparative interpretation of rural and urban residence depends upon this distinction is essential for participation and progression. The institutional consequence follows from whether the estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. When several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on residence linked to service availability and travel conditions. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states.[REF-23] [REF-24]

    Review of rural and urban residence is credible only where it explains if a conclusion changes under a reasonable specification, that instability is part of the finding. The evidence must therefore clarify how national averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: national rural definitions preserved and comparison limitations stated. Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved.[REF-01] [REF-07]

    rural and urban residence cannot be judged without identifying analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. The resulting interpretation should show why missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. An adequate equity account asks whether remote villages, pastoral populations and peri-urban settlements are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation.[REF-02] [REF-03]

    20

    Subnational administrative disparity

    subnational administrative disparity cannot be judged without identifying a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. For the learners concerned, the decisive consideration is whether the report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of subnational administrative disparity lies in making unequal educational experience observable. Here, the relevant phenomenon is variation between provinces, districts or comparable areas, not the administrative convenience of the available categories. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-23] [REF-24]

    A defensible account of subnational administrative disparity distinguishes at minimum this includes population, geography, date, collection method, classification and known exclusions. For the learners concerned, the decisive consideration is whether a national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is area estimates with population size and precision. Its metadata should travel with every published value.[REF-01] [REF-07]

    The practical standard for subnational administrative disparity concerns it states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. The evidence must therefore clarify how it does not assign cause from a cross-sectional difference. Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that administrative rankings not mistaken for causal explanations. A responsible commentary distinguishes observation, calculation and interpretation.[REF-02] [REF-03]

    Institutional action on subnational administrative disparity should be tested against it should involve statistical judgement, legal safeguards and knowledge of the affected community. The resulting interpretation should show why suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. For small districts and areas with incomplete reporting, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical.[REF-04] [REF-06]

    21

    Distance, isolation and transport

    The practical standard for distance, isolation and transport concerns the report should state the educational consequence before choosing a gap, ratio, threshold or rank. A contrary reading would overlook that the indicator question concerns physical accessibility of the nearest appropriate service. Its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage.[REF-01] [REF-07]

    distance, isolation and transport requires a decision about if a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. This matters because administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is travel time, route safety and seasonal interruption rather than straight-line distance alone. Numerator, denominator, reference date, unit and exclusions should appear together.[REF-02] [REF-03]

    For distance, isolation and transport, the material distinction is between statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. For the learners concerned, the decisive consideration is whether explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: household reports and facility mapping reconciled. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible.[REF-04] [REF-06]

    Evidence concerning distance, isolation and transport should establish coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. The public account remains incomplete unless it explains how if direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include learners with limited mobility and communities cut off seasonally. These populations may be missing not only from good outcomes but from the denominator itself.[REF-08] [REF-09]

    Part VIII

    Disability, language and identity

    22

    Disability-sensitive education data

    disability-sensitive education data cannot be judged without identifying the substantive interest is participation and learning by functional difficulty and support requirement, observed for a population and period that are stated before calculation. The resulting interpretation should show why the measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. Disability-sensitive education data should be approached as a defined measurement problem.[REF-02] [REF-03]

    A defensible account of disability-sensitive education data distinguishes where estimates are revised, both the reason and effect of revision should remain accessible. The institutional consequence follows from whether measurement should use questions designed for comparable reporting without defining the child by diagnosis alone. This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined. School returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. Reconciliation should record what each source can and cannot represent.[REF-04] [REF-06]

    disability-sensitive education data cannot be judged without identifying percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. A proportionate conclusion must also recognise that the comparison should identify the reference category but avoid presenting it as a natural norm. Policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that identification, environment and accommodation kept analytically distinct. Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses.[REF-08] [REF-09]

    disability-sensitive education data requires a decision about combining unknown observations with the majority group biases both estimates and obscures the weakness. The institutional consequence follows from whether where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for learners whose impairments are not recorded by schools. Their circumstances may alter access to enumeration, classification and the service being measured. The review should record non-response, unknown status and excluded locations separately.[REF-10] [REF-12]

    23

    Language of home and instruction

    Comparative interpretation of language of home and instruction depends upon the study seeks evidence on alignment between learner language, teaching and assessment; it does not infer a learner's circumstances from a national or regional mean. A proportionate conclusion must also recognise that the measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. For language of home and instruction, the first requirement is conceptual clarity.[REF-04] [REF-06]

    A defensible account of language of home and instruction distinguishes the estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. For the learners concerned, the decisive consideration is whether when several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on language categories reflecting local use and instructional practice. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression.[REF-08] [REF-09]

    Public responsibility for language of home and instruction begins with results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. The evidence must therefore clarify how if a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution. A defensible account of language of home and instruction distinguishes nor should a group estimate be read as a description of every member. The public account remains incomplete unless it explains how within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: small groups not erased through broad national labels.[REF-10] [REF-12]

    Public responsibility for language of home and instruction begins with missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. The institutional consequence follows from whether an adequate equity account asks whether minority-language and multilingual learners are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination.[REF-13] [REF-14]

    24

    Ethnicity, indigeneity and protected identity

    The practical standard for ethnicity, indigeneity and protected identity concerns a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The evidence must therefore clarify how the report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of disparity associated with historically excluded identity groups begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-08] [REF-09]

    A defensible account of ethnicity, indigeneity and protected identity distinguishes a figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. The resulting interpretation should show why operationally, the measure is lawful, voluntary and contextually meaningful classification. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. These are substantive attributes because they determine who can appear in the evidence.[REF-10] [REF-12]

    A defensible account of ethnicity, indigeneity and protected identity distinguishes it does not assign cause from a cross-sectional difference. The public account remains incomplete unless it explains how apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that self-identification protected and non-response reported. A responsible commentary distinguishes observation, calculation and interpretation. It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods.[REF-13] [REF-14]

    ethnicity, indigeneity and protected identity requires a decision about this balance is contextual rather than mechanical. The institutional consequence follows from whether it should involve statistical judgement, legal safeguards and knowledge of the affected community. Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. For communities exposed to discrimination or forced assimilation, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable.[REF-15] [REF-16]

    Part IX

    Conflict, disaster and mobility

    25

    Education under conflict and insecurity

    Public responsibility for education under conflict and insecurity begins with here, the relevant phenomenon is access, attendance and learning where violence alters service and movement, not the administrative convenience of the available categories. This matters because the measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of education under conflict and insecurity lies in making unequal educational experience observable.[REF-10] [REF-12]

    For education under conflict and insecurity, the material distinction is between numerator, denominator, reference date, unit and exclusions should appear together. The resulting interpretation should show why if a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is location- and time-specific observation with explicit coverage gaps.[REF-13] [REF-14]

    The central question in education under conflict and insecurity is statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. A proportionate conclusion must also recognise that explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: absence caused by insecurity distinguished from ordinary dropout. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible.[REF-15] [REF-16]

    Comparative interpretation of education under conflict and insecurity depends upon confidentiality is essential, especially where identity or status creates risk. The evidence must therefore clarify how protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include learners in insecure areas and host communities. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero.[REF-17] [REF-19]

    27

    Refugees, displaced persons and migrants

    Public responsibility for refugees, displaced persons and migrants begins with the report should state the educational consequence before choosing a gap, ratio, threshold or rank. A proportionate conclusion must also recognise that refugees, displaced persons and migrants should be approached as a defined measurement problem. The substantive interest is educational participation across changing legal and residential situations, observed for a population and period that are stated before calculation. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage.[REF-15] [REF-16]

    A defensible account of refugees, displaced persons and migrants distinguishes when several sources exist, consistency is evidence to consider, not proof that common error is absent. The institutional consequence follows from whether a defensible statistic would be based on status, origin, current location and service access recorded separately. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression. The estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality.[REF-17] [REF-19]

    In assessing refugees, displaced persons and migrants, authorities must determine within-group variation and unmeasured intersecting conditions remain material. The resulting interpretation should show why the analytical rule is clear: mobility never converted into duplicate enrolment or unexplained disappearance. Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member.[REF-20] [REF-21]

    Public responsibility for refugees, displaced persons and migrants begins with attrition at any stage can produce an apparently complete indicator from a selective population. The institutional consequence follows from whether field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. Missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. An adequate equity account asks whether undocumented migrants, refugees and internally displaced learners are represented at each stage: population frame, collection, valid response, classification, analysis and publication.[REF-23] [REF-24]

    Part X

    School conditions and teachers

    28

    Teacher availability and distribution

    teacher availability and distribution requires a decision about the measure should preserve both the observed level and the distribution relevant to the claim. A contrary reading would overlook that a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. For teacher availability and distribution, the first requirement is conceptual clarity. The study seeks evidence on access to competent teaching across schools and subjects; it does not infer a learner's circumstances from a national or regional mean.[REF-17] [REF-19]

    The practical standard for teacher availability and distribution concerns a figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. The public account remains incomplete unless it explains how operationally, the measure is teachers present and assigned relative to learner and curriculum need. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. These are substantive attributes because they determine who can appear in the evidence.[REF-20] [REF-21]

    Institutional action on teacher availability and distribution should be tested against it states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. For the learners concerned, the decisive consideration is whether it does not assign cause from a cross-sectional difference. Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that payroll totals not substituted for classroom availability. A responsible commentary distinguishes observation, calculation and interpretation.[REF-23] [REF-24]

    Public responsibility for teacher availability and distribution begins with it should involve statistical judgement, legal safeguards and knowledge of the affected community. The resulting interpretation should show why suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. For schools serving poor, remote or displaced communities, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical.[REF-01] [REF-07]

    29

    Class size, multi-grade teaching and time

    Public responsibility for class size, multi-grade teaching and time begins with a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. This matters because the report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of the instructional conditions experienced by learners begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-20] [REF-21]

    In assessing class size, multi-grade teaching and time, authorities must determine administrative records should be reconciled with population-based evidence where their coverage differs. The public account remains incomplete unless it explains how a discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is class organisation, scheduled time and delivered time considered together. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners.[REF-23] [REF-24]

    class size, multi-grade teaching and time requires a decision about a ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. A proportionate conclusion must also recognise that reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: simple pupil-teacher ratios not treated as a complete quality measure. The comparison should show the level for each group as well as any ratio or gap.[REF-01] [REF-07]

    Review of class size, multi-grade teaching and time is credible only where it explains protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. A proportionate conclusion must also recognise that the distributional review must deliberately include early grades and mixed-age classes. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. Confidentiality is essential, especially where identity or status creates risk.[REF-02] [REF-03]

    30

    Materials, facilities and basic services

    The practical standard for materials, facilities and basic services concerns the report should state the educational consequence before choosing a gap, ratio, threshold or rank. For the learners concerned, the decisive consideration is whether the public value of materials, facilities and basic services lies in making unequal educational experience observable. Here, the relevant phenomenon is usable learning resources and safe, accessible school conditions, not the administrative convenience of the available categories. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage.[REF-23] [REF-24]

    A defensible account of materials, facilities and basic services distinguishes this matters because school returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. A contrary reading would overlook that reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use availability joined to condition, accessibility and regular use. This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. Institutional action on materials, facilities and basic services should be tested against source coverage must be tested before sources are combined.[REF-01] [REF-07]

    A defensible account of materials, facilities and basic services distinguishes the comparison should identify the reference category but avoid presenting it as a natural norm. The institutional consequence follows from whether policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that delivery counts checked against learner access. Disparity measures should not replace the underlying distributions. A difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. Percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation.[REF-02] [REF-03]

    Public responsibility for materials, facilities and basic services begins with the review should record non-response, unknown status and excluded locations separately. This matters because combining unknown observations with the majority group biases both estimates and obscures the weakness. Where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for learners in temporary or damaged premises. Their circumstances may alter access to enumeration, classification and the service being measured.[REF-04] [REF-06]

    Part XI

    Finance and distribution

    31

    Public spending by level and function

    Public responsibility for public spending by level and function begins with a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. This matters because the report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns resources assigned to education purposes across the system. Its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-01] [REF-07]

    public spending by level and function requires a decision about analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. The institutional consequence follows from whether this distinction is essential for participation and progression. The estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. When several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on expenditure classified by level, recurrent or capital use and responsible body. The definition should be fixed for the comparison at hand and deviations recorded.[REF-02] [REF-03]

    Public responsibility for public spending by level and function begins with within-group variation and unmeasured intersecting conditions remain material. The evidence must therefore clarify how the analytical rule is clear: budgets, commitments and actual expenditure distinguished. Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member.[REF-04] [REF-06]

    Institutional action on public spending by level and function should be tested against field arrangements need relevant languages, accessible formats and safe participation. The evidence must therefore clarify how analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. Missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. An adequate equity account asks whether basic education services under fiscal pressure are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population.[REF-08] [REF-09]

    32

    Incidence of education spending

    A defensible account of incidence of education spending distinguishes a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. For the learners concerned, the decisive consideration is whether the report should state the educational consequence before choosing a gap, ratio, threshold or rank. Incidence of education spending should be approached as a defined measurement problem. The substantive interest is who benefits from publicly financed places and services, observed for a population and period that are stated before calculation. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-02] [REF-03]

    Comparative interpretation of incidence of education spending depends upon these are substantive attributes because they determine who can appear in the evidence. A proportionate conclusion must also recognise that a figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is unit resources combined with participation across population groups. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions. A national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules.[REF-04] [REF-06]

    Public responsibility for incidence of education spending begins with it does not assign cause from a cross-sectional difference. For the learners concerned, the decisive consideration is whether apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that benefit estimates not treated as household income. A responsible commentary distinguishes observation, calculation and interpretation. It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods.[REF-08] [REF-09]

    incidence of education spending requires a decision about the public report can still state that a disparity was examined, whether action is required and which body will monitor it. The resulting interpretation should show why for groups excluded before public spending can reach them, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community. Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells.[REF-10] [REF-12]

    33

    Protecting equity during fiscal constraint

    In assessing protecting equity during fiscal constraint, authorities must determine a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The resulting interpretation should show why the report should state the educational consequence before choosing a gap, ratio, threshold or rank. For protecting equity during fiscal constraint, the first requirement is conceptual clarity. The study seeks evidence on whether reductions or delays fall disproportionately on weaker services; it does not infer a learner's circumstances from a national or regional mean. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-04] [REF-06]

    Evidence concerning protecting equity during fiscal constraint should establish it should require examination of definitions, timing, migration, duplication and non-response. The resulting interpretation should show why the resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is dated finance and service indicators read together. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs. A discrepancy is not resolved by selecting the more favourable source.[REF-08] [REF-09]

    In assessing protecting equity during fiscal constraint, authorities must determine the comparison should show the level for each group as well as any ratio or gap. For the learners concerned, the decisive consideration is whether a ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: national totals tested against subnational allocation and household costs.[REF-10] [REF-12]

    The practical standard for protecting equity during fiscal constraint concerns if direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. A contrary reading would overlook that confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include poor households and institutions with little financial reserve. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete.[REF-13] [REF-14]

    Part XII

    Data sources and measurement error

    34

    Administrative records

    A defensible account of administrative records distinguishes without that purpose, disaggregation can multiply figures without improving public judgement. The public account remains incomplete unless it explains how the measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of regular learner, staff, facility and finance information begins by naming the decision the evidence may inform.[REF-08] [REF-09]

    administrative records requires a decision about reconciliation should record what each source can and cannot represent. The public account remains incomplete unless it explains how where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use clear definitions, reporting coverage and revision history. This formulation requires the reporting body to preserve the population base and the observation period beside the result. Counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined. School returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail.[REF-10] [REF-12]

    Evidence concerning administrative records should establish a difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. The resulting interpretation should show why percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. The comparison should identify the reference category but avoid presenting it as a natural norm. Policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that non-reporting institutions kept visible in aggregates. Disparity measures should not replace the underlying distributions.[REF-13] [REF-14]

    For administrative records, the material distinction is between where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. The resulting interpretation should show why particular scrutiny is required for small, private, non-formal and emergency providers. Their circumstances may alter access to enumeration, classification and the service being measured. The review should record non-response, unknown status and excluded locations separately. Combining unknown observations with the majority group biases both estimates and obscures the weakness. Where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated.[REF-15] [REF-16]

    35

    Household surveys

    The practical standard for household surveys concerns here, the relevant phenomenon is population-based evidence beyond enrolled learners, not the administrative convenience of the available categories. For the learners concerned, the decisive consideration is whether the measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of household surveys lies in making unequal educational experience observable.[REF-10] [REF-12]

    Evidence concerning household surveys should establish analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. The evidence must therefore clarify how this distinction is essential for participation and progression. The estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. When several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on probability samples, weights and field dates documented. The definition should be fixed for the comparison at hand and deviations recorded.[REF-13] [REF-14]

    household surveys cannot be judged without identifying if a conclusion changes under a reasonable specification, that instability is part of the finding. This matters because national averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member. Within-group variation and unmeasured intersecting conditions remain material. The analytical rule is clear: sampling and non-response uncertainty carried into group comparisons. Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved.[REF-15] [REF-16]

    household surveys requires a decision about missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. A proportionate conclusion must also recognise that an adequate equity account asks whether small minorities and mobile households are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination.[REF-17] [REF-19]

    36

    Censuses and population frames

    The practical standard for censuses and population frames concerns a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The resulting interpretation should show why the report should state the educational consequence before choosing a gap, ratio, threshold or rank. The indicator question concerns broad population coverage and small-area denominators. Its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-13] [REF-14]

    The central question in censuses and population frames is at minimum this includes population, geography, date, collection method, classification and known exclusions. The resulting interpretation should show why a national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is enumeration date, usual residence and institutional coverage stated. Its metadata should travel with every published value.[REF-15] [REF-16]

    censuses and population frames cannot be judged without identifying apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. The resulting interpretation should show why use of the indicator is bounded by the principle that long intervals and under-enumeration acknowledged. A responsible commentary distinguishes observation, calculation and interpretation. It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference.[REF-17] [REF-19]

    For censuses and population frames, the material distinction is between suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. A proportionate conclusion must also recognise that the public report can still state that a disparity was examined, whether action is required and which body will monitor it. For homeless, displaced and geographically isolated people, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community.[REF-20] [REF-21]

    Part XIII

    Disaggregation and intersection

    37

    Single-axis disaggregation

    Review of single-axis disaggregation is credible only where it explains a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The public account remains incomplete unless it explains how the report should state the educational consequence before choosing a gap, ratio, threshold or rank. Single-axis disaggregation should be approached as a defined measurement problem. The substantive interest is separate reporting by sex, wealth, residence or another characteristic, observed for a population and period that are stated before calculation. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-15] [REF-16]

    A defensible account of single-axis disaggregation distinguishes administrative records should be reconciled with population-based evidence where their coverage differs. A proportionate conclusion must also recognise that a discrepancy is not resolved by selecting the more favourable source. It should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is levels, gaps and denominators shown for each category. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners.[REF-17] [REF-19]

    single-axis disaggregation requires a decision about the comparison should show the level for each group as well as any ratio or gap. For the learners concerned, the decisive consideration is whether a ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. A defensible account of single-axis disaggregation distinguishes explanation requires evidence on institutions, resources, households and prior conditions. A proportionate conclusion must also recognise that interpretation follows this limitation: one axis not presented as a complete account of marginalisation.[REF-20] [REF-21]

    single-axis disaggregation requires a decision about confidentiality is essential, especially where identity or status creates risk. A contrary reading would overlook that protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include groups whose disadvantage lies on another unmeasured dimension. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero.[REF-23] [REF-24]

    38

    Intersecting categories

    Institutional action on intersecting categories should be tested against the study seeks evidence on joint distributions such as sex by wealth and residence; it does not infer a learner's circumstances from a national or regional mean. For the learners concerned, the decisive consideration is whether the measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. For intersecting categories, the first requirement is conceptual clarity.[REF-17] [REF-19]

    Comparative interpretation of intersecting categories depends upon this formulation requires the reporting body to preserve the population base and the observation period beside the result. A proportionate conclusion must also recognise that counts reveal scale; rates permit comparison; neither is sufficient alone. Source coverage must be tested before sources are combined. School returns may describe enrolled learners well while saying little about children outside institutions, whereas household enquiries may reach non-enrolled children but provide limited school detail. Reconciliation should record what each source can and cannot represent. Where estimates are revised, both the reason and effect of revision should remain accessible. Measurement should use pre-specified combinations with sufficient observations.[REF-20] [REF-21]

    The central question in intersecting categories is a difference in means may reflect the lower tail, the upper tail or change across the whole range; these possibilities call for different responses. The institutional consequence follows from whether percentage-point gaps, ratios and relative risks answer different questions and should not be exchanged without explanation. The comparison should identify the reference category but avoid presenting it as a natural norm. Policy significance depends upon the educational consequence and the number of learners affected, not solely upon statistical separation. The governing caution is that empty or unstable cells reported honestly. Disparity measures should not replace the underlying distributions.[REF-23] [REF-24]

    Institutional action on intersecting categories should be tested against combining unknown observations with the majority group biases both estimates and obscures the weakness. The evidence must therefore clarify how where sample size is limited, several years or compatible areas may sometimes be combined, provided the loss of time or place specificity is stated. Where combination would be misleading, a descriptive case record can establish a service problem without pretending to estimate prevalence. Particular scrutiny is required for poor rural girls, disabled learners in remote areas and displaced minorities. Their circumstances may alter access to enumeration, classification and the service being measured. The review should record non-response, unknown status and excluded locations separately.[REF-01] [REF-07]

    39

    Small numbers, disclosure and reliability

    Comparative interpretation of small numbers, disclosure and reliability depends upon a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The public account remains incomplete unless it explains how the report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of useful detail without unreliable estimates or identification begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-20] [REF-21]

    In assessing small numbers, disclosure and reliability, authorities must determine this distinction is essential for participation and progression. A contrary reading would overlook that the estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality. When several sources exist, consistency is evidence to consider, not proof that common error is absent. A defensible statistic would be based on suppression, aggregation or qualitative evidence chosen proportionately. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states.[REF-23] [REF-24]

    The practical standard for small numbers, disclosure and reliability concerns within-group variation and unmeasured intersecting conditions remain material. The public account remains incomplete unless it explains how the analytical rule is clear: confidentiality decisions separated from claims that no disparity exists. Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member.[REF-01] [REF-07]

    Review of small numbers, disclosure and reliability is credible only where it explains attrition at any stage can produce an apparently complete indicator from a selective population. A proportionate conclusion must also recognise that field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination. Missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. An adequate equity account asks whether small communities and learners with rare characteristics are represented at each stage: population frame, collection, valid response, classification, analysis and publication.[REF-02] [REF-03]

    Part XIV

    Comparison, uncertainty and change

    40

    Comparing unlike systems

    For comparing unlike systems, the material distinction is between here, the relevant phenomenon is cross-country patterns based on harmonised but bounded concepts, not the administrative convenience of the available categories. For the learners concerned, the decisive consideration is whether the measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. The public value of comparing unlike systems lies in making unequal educational experience observable.[REF-23] [REF-24]

    Review of comparing unlike systems is credible only where it explains at minimum this includes population, geography, date, collection method, classification and known exclusions. The resulting interpretation should show why a national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. A survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is metadata tests before numerical comparison. Its metadata should travel with every published value.[REF-01] [REF-07]

    Public responsibility for comparing unlike systems begins with apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. A contrary reading would overlook that use of the indicator is bounded by the principle that differences in programme structure and classification remain visible. A responsible commentary distinguishes observation, calculation and interpretation. It states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference.[REF-02] [REF-03]

    A defensible account of comparing unlike systems distinguishes it should involve statistical judgement, legal safeguards and knowledge of the affected community. A contrary reading would overlook that suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. For countries with incomplete or rapidly changing systems, a single group label may conceal important internal differences. Disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This balance is contextual rather than mechanical.[REF-04] [REF-06]

    41

    Sampling error and other uncertainty

    Review of sampling error and other uncertainty is credible only where it explains the report should state the educational consequence before choosing a gap, ratio, threshold or rank. The resulting interpretation should show why the indicator question concerns the range of values reasonably compatible with the observations. Its object is not to divide a population into convenient labels, but to show whether educational opportunity is distributed in a manner that a national total cannot reveal. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage.[REF-01] [REF-07]

    A defensible account of sampling error and other uncertainty distinguishes a discrepancy is not resolved by selecting the more favourable source. This matters because it should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is standard errors, design effects and data-quality qualifications. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs.[REF-02] [REF-03]

    sampling error and other uncertainty cannot be judged without identifying the comparison should show the level for each group as well as any ratio or gap. This matters because a ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups. Reference points therefore need substantive meaning. Where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: rank differences smaller than uncertainty not interpreted.[REF-04] [REF-06]

    Comparative interpretation of sampling error and other uncertainty depends upon coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. The institutional consequence follows from whether if direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero. Confidentiality is essential, especially where identity or status creates risk. Protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include small disaggregated populations. These populations may be missing not only from good outcomes but from the denominator itself.[REF-08] [REF-09]

    Part XV

    Responsible interpretation and action

    43

    Reading disparity without blaming learners

    The central question in reading disparity without blaming learners is a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. A contrary reading would overlook that the report should state the educational consequence before choosing a gap, ratio, threshold or rank. For reading disparity without blaming learners, the first requirement is conceptual clarity. The study seeks evidence on institutional and social conditions associated with unequal outcomes; it does not infer a learner's circumstances from a national or regional mean. The measure should preserve both the observed level and the distribution relevant to the claim.[REF-04] [REF-06]

    The practical standard for reading disparity without blaming learners concerns when several sources exist, consistency is evidence to consider, not proof that common error is absent. The resulting interpretation should show why a defensible statistic would be based on descriptive findings separated from causal claims. The definition should be fixed for the comparison at hand and deviations recorded. Analysts need to show whether the observation refers to a stock on one date, activity over a period or a flow between states. This distinction is essential for participation and progression. The estimate should retain its unrounded numerator and denominator for checking, although published precision should not exceed data quality.[REF-08] [REF-09]

    reading disparity without blaming learners cannot be judged without identifying within-group variation and unmeasured intersecting conditions remain material. The institutional consequence follows from whether the analytical rule is clear: group identity never treated as a mechanism by itself. Results should be tested for sensitivity to plausible alternative definitions, particularly where age bands, residence, wealth grouping or programme equivalence are involved. If a conclusion changes under a reasonable specification, that instability is part of the finding. National averages should remain available as context, yet never as a substitute for the distribution. Nor should a group estimate be read as a description of every member.[REF-10] [REF-12]

    A defensible account of reading disparity without blaming learners distinguishes missingness is itself patterned evidence when it clusters by location or social condition, although its magnitude should not be guessed. The public account remains incomplete unless it explains how an adequate equity account asks whether communities subject to stigma are represented at each stage: population frame, collection, valid response, classification, analysis and publication. Attrition at any stage can produce an apparently complete indicator from a selective population. Field arrangements need relevant languages, accessible formats and safe participation. Analysts should also examine who answers on behalf of whom, since proxy response may be necessary yet less reliable for attendance, impairment or discrimination.[REF-13] [REF-14]

    44

    Turning evidence into equitable policy

    Institutional action on turning evidence into equitable policy should be tested against the measure should preserve both the observed level and the distribution relevant to the claim. The public account remains incomplete unless it explains how a national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage. The report should state the educational consequence before choosing a gap, ratio, threshold or rank. A distribution-sensitive account of a decision rule linking disparity to service, finance or legal responsibility begins by naming the decision the evidence may inform. Without that purpose, disaggregation can multiply figures without improving public judgement.[REF-08] [REF-09]

    turning evidence into equitable policy cannot be judged without identifying a national estimate assembled from local reports should disclose reporting completeness and treatment of missing institutions. The resulting interpretation should show why a survey estimate should disclose weights and uncertainty. A census figure should disclose enumeration rules. These are substantive attributes because they determine who can appear in the evidence. A figure detached from them may be arithmetically correct yet unsuitable for an equity judgement. Operationally, the measure is baseline, intended reach, implementation evidence and review date. Its metadata should travel with every published value. At minimum this includes population, geography, date, collection method, classification and known exclusions.[REF-10] [REF-12]

    The practical standard for turning evidence into equitable policy concerns a responsible commentary distinguishes observation, calculation and interpretation. A proportionate conclusion must also recognise that it states whether a disparity is large in educational terms, whether it is estimated precisely enough for the proposed comparison and whether it persists across sources or periods. It does not assign cause from a cross-sectional difference. Apparent exceptions should be examined rather than removed, because they may reveal classification error, a local policy difference or a population not adequately represented elsewhere. Use of the indicator is bounded by the principle that targets accompanied by distributional safeguards.[REF-13] [REF-14]

    In assessing turning evidence into equitable policy, authorities must determine disaggregation should proceed far enough to reveal a plausible service disparity but stop before estimates become unsafe or persons identifiable. This matters because this balance is contextual rather than mechanical. It should involve statistical judgement, legal safeguards and knowledge of the affected community. Suppression rules need explanation, and restricted analysis may be preferable to public release of small cells. The public report can still state that a disparity was examined, whether action is required and which body will monitor it. For learners farthest below a secured minimum, a single group label may conceal important internal differences.[REF-15] [REF-16]

    45

    A public account beyond the average

    Review of a public account beyond the average is credible only where it explains the report should state the educational consequence before choosing a gap, ratio, threshold or rank. The resulting interpretation should show why the public value of a public account beyond the average lies in making unequal educational experience observable. Here, the relevant phenomenon is a concise national statement of level, distribution, missingness and remedy, not the administrative convenience of the available categories. The measure should preserve both the observed level and the distribution relevant to the claim. A national figure remains necessary for scale and direction, yet it cannot answer which learners receive the service, how far groups lie from an acceptable minimum or whether apparent progress reflects changed coverage.[REF-10] [REF-12]

    A defensible account of a public account beyond the average distinguishes a discrepancy is not resolved by selecting the more favourable source. This matters because it should require examination of definitions, timing, migration, duplication and non-response. The resulting indicator should be reproducible from stated components, while any necessary estimation remains distinguishable from direct observation. The preferred construction is national totals presented beside selected group and place indicators. Numerator, denominator, reference date, unit and exclusions should appear together. If a proportion is reported, its underlying population count remains material: identical percentages can describe very different evidentiary strength and numbers of affected learners. Administrative records should be reconciled with population-based evidence where their coverage differs.[REF-13] [REF-14]

    Institutional action on a public account beyond the average should be tested against reference points therefore need substantive meaning. The public account remains incomplete unless it explains how where a minimum entitlement or policy threshold is relevant, the distance of every group from that threshold should be visible. Statistical association can identify where disadvantage is concentrated, but it does not establish why the disparity arose. Explanation requires evidence on institutions, resources, households and prior conditions. Interpretation follows this limitation: progress claims bounded by evidence coverage and unresolved gaps. The comparison should show the level for each group as well as any ratio or gap. A ratio can approach one because the more advantaged group deteriorates; a small absolute gap can coexist with severe deprivation for all groups.[REF-15] [REF-16]

    The central question in a public account beyond the average is confidentiality is essential, especially where identity or status creates risk. A proportionate conclusion must also recognise that protection, however, should lead to careful access and publication rules; it should not make an affected population analytically disappear. The distributional review must deliberately include every learner otherwise hidden by a successful average. These populations may be missing not only from good outcomes but from the denominator itself. Coverage assessment should compare survey frames, census listings, administrative registers and local knowledge without assuming that any one is complete. If direct estimation is impossible, the report should state the evidence gap and use appropriate qualitative or service information rather than assign zero.[REF-17] [REF-19]

    Part XVI

    Applied distributional analysis

    46

    Composite measures and the loss of meaning

    Review of composite measures and the loss of meaning is credible only where it explains a result is useful only if a reader can identify the represented population, the reference period and the educational consequence attached to movement. The public account remains incomplete unless it explains how the national total supplies context, while the distribution tests whether the total is shared. Where the measure condenses several observations, its construction must remain open to reconstruction from the underlying values. Where it separates groups, classifications must be lawful, meaningful and sufficiently stable for the comparison. The analytical purpose is combining several dimensions into a summary measure. That purpose should be written before the calculation because method follows the intended inference. The evidence base comprises participation, completion, learning and school conditions, each of which describes a different feature of educational opportunity.[REF-01] [REF-03]

    composite measures and the loss of meaning requires a decision about they determine how strongly one component, place or population can influence the conclusion. The institutional consequence follows from whether a defensible choice begins with the substantive education question and is then tested against alternative reasonable specifications. If the headline conclusion changes materially, the range of results should be reported. Sensitivity does not make the exercise useless; it prevents a conventional choice from appearing inevitable. Documentation should also state which observations are direct, which are estimated and which are unavailable. Missing values should remain missing unless an explicit estimation method and its effect are shown. The principal methodological questions concern weights, normalisation and substitution. Choices on these matters are not neutral presentation details.[REF-02] [REF-04]

    The practical standard for composite measures and the loss of meaning concerns avoiding it requires the underlying counts and distributions to remain visible beside any summary. For the learners concerned, the decisive consideration is whether analysts should check whether an apparently favourable result arose through changed coverage, population movement, reclassification or concentration on cases nearest a threshold. A disparity can narrow because the better-served group deteriorates, while a composite score can improve although a protected minimum worsens. Direction must therefore be joined to adequacy. Comparison should also examine absolute numbers: a smaller rate in a growing population may still correspond to more learners without the secured opportunity. No result should be described as equitable solely because one relative measure improved. The main interpretive danger is a low value in one dimension being concealed by a high value in another.[REF-05] [REF-06]

    Institutional action on composite measures and the loss of meaning should be tested against a record of disagreement should identify plausible sources and the decision consequence. The resulting interpretation should show why if the available evidence does not support a proposed level of disaggregation, the limitation should appear in the finding rather than in a remote methodological note. Source quality should be considered dimension by dimension. Administrative returns may provide frequent local detail but exclude non-participants and non-reporting institutions. Surveys can represent households beyond formal education, subject to sample size, field access and response. Censuses can support small-area denominators at longer intervals, while enumeration rules and population movement affect coverage. A record of agreement should follow reconciliation of concepts rather than simple numerical proximity.[REF-08] [REF-12]

    The central question in composite measures and the loss of meaning is these duties do not justify silence about a serious disparity. A contrary reading would overlook that the public account may report a broader group, a range, a qualitative service failure or restricted findings, provided that it states what cannot safely be quantified and which body is responsible for better evidence. Equity review asks who can disappear during construction of the measure. Learners outside school, people in temporary settlements, small language communities, persons with disabilities and households beyond a survey frame can all be absent before analysis begins. Analysts should compare coverage against independent population and service information and preserve an unknown category where classification is incomplete. Small cells require protection against disclosure and caution about statistical stability.[REF-13] [REF-16]

    composite measures and the loss of meaning requires a decision about every response should name the expected population reach and the later observation that will test it. A contrary reading would overlook that otherwise a disparity can generate activity without demonstrating that conditions changed for the intended learners. The policy user is policy-makers deciding whether a summary aids or obscures resource allocation. Evidence should lead to a stated decision class: immediate removal of an access barrier, redistribution of staff or finance, further investigation, amendment of a classification, or evaluation of an existing measure. The certainty required depends upon the consequence. Credible indications of severe exclusion can justify protective action before exact prevalence is known, whereas durable allocation formulas require review as evidence improves.[REF-17] [REF-19]

    Evidence concerning composite measures and the loss of meaning should establish feedback needs a recorded route into classification, service design or further enquiry. For the learners concerned, the decisive consideration is whether people should not be asked repeatedly for sensitive information when no competent body can act upon the answer. Participation by affected communities strengthens both interpretation and legitimacy. Local knowledge can identify seasonal movement, unsafe routes, hidden household costs, language use and service boundaries that national categories miss. Consultation should not be treated as statistical verification, nor should a survey estimate displace credible evidence of a local barrier. The two forms answer different questions. Authorities should provide accessible explanations of definitions and invite correction where categories misdescribe experience.[REF-21] [REF-23]

    The central question in composite measures and the loss of meaning is it should state what the evidence shows, the strength and limits of that conclusion, the people and places most affected, the action within public authority and the date for review. A contrary reading would overlook that changes to definitions, boundaries or population estimates should appear at the point where a series changes. Earlier values should remain available so that revision is not mistaken for real progress. A concise account can carry considerable density when every figure retains its population and consequence. The objective is not maximum numerical output. It is a trustworthy connection between unequal educational experience, public responsibility and corrective action. Final reporting should give a clear institutional judgement.[REF-23] [REF-24]

    47

    Decomposing an observed education gap

    decomposing an observed education gap requires a decision about the evidence base comprises within-group and between-group differences, each of which describes a different feature of educational opportunity. For the learners concerned, the decisive consideration is whether a result is useful only if a reader can identify the represented population, the reference period and the educational consequence attached to movement. The national total supplies context, while the distribution tests whether the total is shared. Where the measure condenses several observations, its construction must remain open to reconstruction from the underlying values. Where it separates groups, classifications must be lawful, meaningful and sufficiently stable for the comparison. The analytical purpose is examining how a national disparity is distributed across places and population groups. That purpose should be written before the calculation because method follows the intended inference.[REF-02] [REF-04]

    Institutional action on decomposing an observed education gap should be tested against if the headline conclusion changes materially, the range of results should be reported. A contrary reading would overlook that sensitivity does not make the exercise useless; it prevents a conventional choice from appearing inevitable. Documentation should also state which observations are direct, which are estimated and which are unavailable. Missing values should remain missing unless an explicit estimation method and its effect are shown. The principal methodological questions concern population shares, outcome levels and overlapping membership. Choices on these matters are not neutral presentation details. They determine how strongly one component, place or population can influence the conclusion. A defensible choice begins with the substantive education question and is then tested against alternative reasonable specifications.[REF-05] [REF-06]

    Review of decomposing an observed education gap is credible only where it explains avoiding it requires the underlying counts and distributions to remain visible beside any summary. The public account remains incomplete unless it explains how analysts should check whether an apparently favourable result arose through changed coverage, population movement, reclassification or concentration on cases nearest a threshold. A disparity can narrow because the better-served group deteriorates, while a composite score can improve although a protected minimum worsens. Direction must therefore be joined to adequacy. Comparison should also examine absolute numbers: a smaller rate in a growing population may still correspond to more learners without the secured opportunity. No result should be described as equitable solely because one relative measure improved. The main interpretive danger is treating a descriptive decomposition as proof of cause.[REF-08] [REF-12]

    In assessing decomposing an observed education gap, authorities must determine credible indications of severe exclusion can justify protective action before exact prevalence is known, whereas durable allocation formulas require review as evidence improves. The evidence must therefore clarify how every response should name the expected population reach and the later observation that will test it. Otherwise a disparity can generate activity without demonstrating that conditions changed for the intended learners. The policy user is authorities locating where further enquiry and action are warranted. Evidence should lead to a stated decision class: immediate removal of an access barrier, redistribution of staff or finance, further investigation, amendment of a classification, or evaluation of an existing measure. The certainty required depends upon the consequence.[REF-21] [REF-23]

    48

    Setting distribution-sensitive targets

    The central question in setting distribution-sensitive targets is where the measure condenses several observations, its construction must remain open to reconstruction from the underlying values. This matters because where it separates groups, classifications must be lawful, meaningful and sufficiently stable for the comparison. The analytical purpose is expressing progress as improvement in secured minimums and unjustified gaps. That purpose should be written before the calculation because method follows the intended inference. The evidence base comprises the national level, the least-served group and the lower tail of the distribution, each of which describes a different feature of educational opportunity. A result is useful only if a reader can identify the represented population, the reference period and the educational consequence attached to movement. The national total supplies context, while the distribution tests whether the total is shared.[REF-05] [REF-06]

    The practical standard for setting distribution-sensitive targets concerns documentation should also state which observations are direct, which are estimated and which are unavailable. The evidence must therefore clarify how missing values should remain missing unless an explicit estimation method and its effect are shown. The principal methodological questions concern baseline stability, ambition and safeguards against exclusion. Choices on these matters are not neutral presentation details. They determine how strongly one component, place or population can influence the conclusion. A defensible choice begins with the substantive education question and is then tested against alternative reasonable specifications. If the headline conclusion changes materially, the range of results should be reported. Sensitivity does not make the exercise useless; it prevents a conventional choice from appearing inevitable.[REF-08] [REF-12]

    The central question in setting distribution-sensitive targets is a disparity can narrow because the better-served group deteriorates, while a composite score can improve although a protected minimum worsens. The resulting interpretation should show why direction must therefore be joined to adequacy. Comparison should also examine absolute numbers: a smaller rate in a growing population may still correspond to more learners without the secured opportunity. No result should be described as equitable solely because one relative measure improved. The main interpretive danger is meeting a mean target while abandoning those farthest behind. Avoiding it requires the underlying counts and distributions to remain visible beside any summary. Analysts should check whether an apparently favourable result arose through changed coverage, population movement, reclassification or concentration on cases nearest a threshold.[REF-13] [REF-16]

    The practical standard for setting distribution-sensitive targets concerns credible indications of severe exclusion can justify protective action before exact prevalence is known, whereas durable allocation formulas require review as evidence improves. The resulting interpretation should show why every response should name the expected population reach and the later observation that will test it. Otherwise a disparity can generate activity without demonstrating that conditions changed for the intended learners. The policy user is governments linking national commitments to subnational delivery. Evidence should lead to a stated decision class: immediate removal of an access barrier, redistribution of staff or finance, further investigation, amendment of a classification, or evaluation of an existing measure. The certainty required depends upon the consequence.[REF-23] [REF-24]

    49

    Linking learners to service geography

    Evidence concerning linking learners to service geography should establish that purpose should be written before the calculation because method follows the intended inference. The evidence must therefore clarify how the evidence base comprises settlement populations, travel conditions, schools, teachers and programme levels, each of which describes a different feature of educational opportunity. A result is useful only if a reader can identify the represented population, the reference period and the educational consequence attached to movement. The national total supplies context, while the distribution tests whether the total is shared. Where the measure condenses several observations, its construction must remain open to reconstruction from the underlying values. Where it separates groups, classifications must be lawful, meaningful and sufficiently stable for the comparison. The analytical purpose is relating participation and learning to the location and capacity of education services.[REF-08] [REF-12]

    A defensible account of linking learners to service geography distinguishes they determine how strongly one component, place or population can influence the conclusion. The public account remains incomplete unless it explains how a defensible choice begins with the substantive education question and is then tested against alternative reasonable specifications. If the headline conclusion changes materially, the range of results should be reported. Sensitivity does not make the exercise useless; it prevents a conventional choice from appearing inevitable. Documentation should also state which observations are direct, which are estimated and which are unavailable. Missing values should remain missing unless an explicit estimation method and its effect are shown. The principal methodological questions concern geographical scale, boundary effects and facility catchments. Choices on these matters are not neutral presentation details.[REF-13] [REF-16]

    The central question in linking learners to service geography is no result should be described as equitable solely because one relative measure improved. For the learners concerned, the decisive consideration is whether the main interpretive danger is assuming the nearest mapped institution is accessible or appropriate. Avoiding it requires the underlying counts and distributions to remain visible beside any summary. Analysts should check whether an apparently favourable result arose through changed coverage, population movement, reclassification or concentration on cases nearest a threshold. A disparity can narrow because the better-served group deteriorates, while a composite score can improve although a protected minimum worsens. Direction must therefore be joined to adequacy. Comparison should also examine absolute numbers: a smaller rate in a growing population may still correspond to more learners without the secured opportunity.[REF-17] [REF-19]

    linking learners to service geography cannot be judged without identifying the certainty required depends upon the consequence. The institutional consequence follows from whether credible indications of severe exclusion can justify protective action before exact prevalence is known, whereas durable allocation formulas require review as evidence improves. Every response should name the expected population reach and the later observation that will test it. Otherwise a disparity can generate activity without demonstrating that conditions changed for the intended learners. The policy user is planners choosing sites, transport support and teacher deployment. Evidence should lead to a stated decision class: immediate removal of an access barrier, redistribution of staff or finance, further investigation, amendment of a classification, or evaluation of an existing measure.[REF-01] [REF-03]

    50

    Reconciling conflicting sources

    reconciling conflicting sources cannot be judged without identifying the national total supplies context, while the distribution tests whether the total is shared. The public account remains incomplete unless it explains how where the measure condenses several observations, its construction must remain open to reconstruction from the underlying values. Where it separates groups, classifications must be lawful, meaningful and sufficiently stable for the comparison. The analytical purpose is interpreting differences between administrative, survey and census estimates. That purpose should be written before the calculation because method follows the intended inference. The evidence base comprises coverage, timing, concepts and reporting incentives, each of which describes a different feature of educational opportunity. A result is useful only if a reader can identify the represented population, the reference period and the educational consequence attached to movement.[REF-13] [REF-16]

    In assessing reconciling conflicting sources, authorities must determine a defensible choice begins with the substantive education question and is then tested against alternative reasonable specifications. A proportionate conclusion must also recognise that if the headline conclusion changes materially, the range of results should be reported. Sensitivity does not make the exercise useless; it prevents a conventional choice from appearing inevitable. Documentation should also state which observations are direct, which are estimated and which are unavailable. Missing values should remain missing unless an explicit estimation method and its effect are shown. The principal methodological questions concern a documented comparison of population and variable definitions. Choices on these matters are not neutral presentation details. They determine how strongly one component, place or population can influence the conclusion.[REF-17] [REF-19]

    Institutional action on reconciling conflicting sources should be tested against a disparity can narrow because the better-served group deteriorates, while a composite score can improve although a protected minimum worsens. The evidence must therefore clarify how direction must therefore be joined to adequacy. Comparison should also examine absolute numbers: a smaller rate in a growing population may still correspond to more learners without the secured opportunity. No result should be described as equitable solely because one relative measure improved. The main interpretive danger is averaging incompatible estimates into an apparently precise figure. Avoiding it requires the underlying counts and distributions to remain visible beside any summary. Analysts should check whether an apparently favourable result arose through changed coverage, population movement, reclassification or concentration on cases nearest a threshold.[REF-21] [REF-23]

    The practical standard for reconciling conflicting sources concerns every response should name the expected population reach and the later observation that will test it. A proportionate conclusion must also recognise that otherwise a disparity can generate activity without demonstrating that conditions changed for the intended learners. The policy user is statistical authorities issuing one bounded account with visible uncertainty. Evidence should lead to a stated decision class: immediate removal of an access barrier, redistribution of staff or finance, further investigation, amendment of a classification, or evaluation of an existing measure. The certainty required depends upon the consequence. Credible indications of severe exclusion can justify protective action before exact prevalence is known, whereas durable allocation formulas require review as evidence improves.[REF-02] [REF-04]

    51

    Monitoring marginalisation during severe disruption

    The central question in monitoring marginalisation during severe disruption is a result is useful only if a reader can identify the represented population, the reference period and the educational consequence attached to movement. The public account remains incomplete unless it explains how the national total supplies context, while the distribution tests whether the total is shared. Where the measure condenses several observations, its construction must remain open to reconstruction from the underlying values. Where it separates groups, classifications must be lawful, meaningful and sufficiently stable for the comparison. The analytical purpose is maintaining useful distributional evidence when populations and services move rapidly. That purpose should be written before the calculation because method follows the intended inference. The evidence base comprises rapid counts, restored administrative returns and household evidence, each of which describes a different feature of educational opportunity.[REF-17] [REF-19]

    Review of monitoring marginalisation during severe disruption is credible only where it explains missing values should remain missing unless an explicit estimation method and its effect are shown. This matters because the principal methodological questions concern dated estimates, revision practice and minimal essential classifications. Choices on these matters are not neutral presentation details. They determine how strongly one component, place or population can influence the conclusion. A defensible choice begins with the substantive education question and is then tested against alternative reasonable specifications. If the headline conclusion changes materially, the range of results should be reported. Sensitivity does not make the exercise useless; it prevents a conventional choice from appearing inevitable. Documentation should also state which observations are direct, which are estimated and which are unavailable.[REF-21] [REF-23]

    For monitoring marginalisation during severe disruption, the material distinction is between avoiding it requires the underlying counts and distributions to remain visible beside any summary. The institutional consequence follows from whether analysts should check whether an apparently favourable result arose through changed coverage, population movement, reclassification or concentration on cases nearest a threshold. A disparity can narrow because the better-served group deteriorates, while a composite score can improve although a protected minimum worsens. Direction must therefore be joined to adequacy. Comparison should also examine absolute numbers: a smaller rate in a growing population may still correspond to more learners without the secured opportunity. No result should be described as equitable solely because one relative measure improved. The main interpretive danger is using an unstable emergency denominator to assert durable improvement.[REF-23] [REF-24]

    In assessing monitoring marginalisation during severe disruption, authorities must determine every response should name the expected population reach and the later observation that will test it. The resulting interpretation should show why otherwise a disparity can generate activity without demonstrating that conditions changed for the intended learners. The policy user is authorities protecting access while rebuilding regular statistics. Evidence should lead to a stated decision class: immediate removal of an access barrier, redistribution of staff or finance, further investigation, amendment of a classification, or evaluation of an existing measure. The certainty required depends upon the consequence. Credible indications of severe exclusion can justify protective action before exact prevalence is known, whereas durable allocation formulas require review as evidence improves.[REF-05] [REF-06]

    52

    Communicating uncertainty without losing urgency

    For communicating uncertainty without losing urgency, the material distinction is between a result is useful only if a reader can identify the represented population, the reference period and the educational consequence attached to movement. For the learners concerned, the decisive consideration is whether the national total supplies context, while the distribution tests whether the total is shared. Where the measure condenses several observations, its construction must remain open to reconstruction from the underlying values. Where it separates groups, classifications must be lawful, meaningful and sufficiently stable for the comparison. The analytical purpose is explaining what is known strongly enough to justify action and what remains unresolved. That purpose should be written before the calculation because method follows the intended inference. The evidence base comprises point estimates, ranges, quality statements and missing populations, each of which describes a different feature of educational opportunity.[REF-21] [REF-23]

    A defensible account of communicating uncertainty without losing urgency distinguishes missing values should remain missing unless an explicit estimation method and its effect are shown. The evidence must therefore clarify how the principal methodological questions concern plain institutional language joined to exact metadata. Choices on these matters are not neutral presentation details. They determine how strongly one component, place or population can influence the conclusion. A defensible choice begins with the substantive education question and is then tested against alternative reasonable specifications. If the headline conclusion changes materially, the range of results should be reported. Sensitivity does not make the exercise useless; it prevents a conventional choice from appearing inevitable. Documentation should also state which observations are direct, which are estimated and which are unavailable.[REF-23] [REF-24]

    Evidence concerning communicating uncertainty without losing urgency should establish a disparity can narrow because the better-served group deteriorates, while a composite score can improve although a protected minimum worsens. For the learners concerned, the decisive consideration is whether direction must therefore be joined to adequacy. Comparison should also examine absolute numbers: a smaller rate in a growing population may still correspond to more learners without the secured opportunity. No result should be described as equitable solely because one relative measure improved. The main interpretive danger is presenting caution as a reason for inaction or urgency as a reason for overstatement. Avoiding it requires the underlying counts and distributions to remain visible beside any summary. Analysts should check whether an apparently favourable result arose through changed coverage, population movement, reclassification or concentration on cases nearest a threshold.[REF-01] [REF-03]

    The practical standard for communicating uncertainty without losing urgency concerns otherwise a disparity can generate activity without demonstrating that conditions changed for the intended learners. For the learners concerned, the decisive consideration is whether the policy user is the public, affected communities and responsible decision-makers. Evidence should lead to a stated decision class: immediate removal of an access barrier, redistribution of staff or finance, further investigation, amendment of a classification, or evaluation of an existing measure. The certainty required depends upon the consequence. Credible indications of severe exclusion can justify protective action before exact prevalence is known, whereas durable allocation formulas require review as evidence improves. Every response should name the expected population reach and the later observation that will test it.[REF-08] [REF-12]

    53

    A national marginalisation profile

    Comparative interpretation of a national marginalisation profile depends upon a result is useful only if a reader can identify the represented population, the reference period and the educational consequence attached to movement. A proportionate conclusion must also recognise that the national total supplies context, while the distribution tests whether the total is shared. Where the measure condenses several observations, its construction must remain open to reconstruction from the underlying values. Where it separates groups, classifications must be lawful, meaningful and sufficiently stable for the comparison. The analytical purpose is assembling a concise recurring account beyond the national average. That purpose should be written before the calculation because method follows the intended inference. The evidence base comprises population, access, progression, learning, conditions, finance and unresolved evidence gaps, each of which describes a different feature of educational opportunity.[REF-23] [REF-24]

    a national marginalisation profile cannot be judged without identifying documentation should also state which observations are direct, which are estimated and which are unavailable. For the learners concerned, the decisive consideration is whether missing values should remain missing unless an explicit estimation method and its effect are shown. The principal methodological questions concern a stable core with context-specific distributions. Choices on these matters are not neutral presentation details. They determine how strongly one component, place or population can influence the conclusion. A defensible choice begins with the substantive education question and is then tested against alternative reasonable specifications. If the headline conclusion changes materially, the range of results should be reported. Sensitivity does not make the exercise useless; it prevents a conventional choice from appearing inevitable.[REF-01] [REF-03]

    a national marginalisation profile requires a decision about direction must therefore be joined to adequacy. For the learners concerned, the decisive consideration is whether comparison should also examine absolute numbers: a smaller rate in a growing population may still correspond to more learners without the secured opportunity. No result should be described as equitable solely because one relative measure improved. The main interpretive danger is creating an encyclopaedia of indicators without decision priority. Avoiding it requires the underlying counts and distributions to remain visible beside any summary. Analysts should check whether an apparently favourable result arose through changed coverage, population movement, reclassification or concentration on cases nearest a threshold. A disparity can narrow because the better-served group deteriorates, while a composite score can improve although a protected minimum worsens.[REF-02] [REF-04]

    a national marginalisation profile cannot be judged without identifying the certainty required depends upon the consequence. This matters because credible indications of severe exclusion can justify protective action before exact prevalence is known, whereas durable allocation formulas require review as evidence improves. Every response should name the expected population reach and the later observation that will test it. Otherwise a disparity can generate activity without demonstrating that conditions changed for the intended learners. The policy user is parliament, ministries, local authorities and communities reviewing educational equity. Evidence should lead to a stated decision class: immediate removal of an access barrier, redistribution of staff or finance, further investigation, amendment of a classification, or evaluation of an existing measure.[REF-13] [REF-16]

    Part XVII

    Extended disparity interpretation

    54

    Minimum comparison threshold

    minimum comparison threshold cannot be judged without identifying opportunity to learn, participation and exclusions remain material. This matters because a learning-disparity comparison requires a declared construct, represented population, assessment conditions, scale, uncertainty and distribution. National means should be accompanied by lower-tail, threshold and group evidence.[REF-02] [REF-03]

    55

    Within-system interpretation

    The central question in within-system interpretation is missing learners and non-participating schools must remain visible. A proportionate conclusion must also recognise that within-system gaps should preserve place, group and institutional context without assigning cause from identity. Counts, levels, absolute gaps and ratios answer different questions.[REF-01] [REF-06]

    56

    Between-system interpretation

    between-system interpretation cannot be judged without identifying harmonisation does not remove substantive system difference. The resulting interpretation should show why between-system comparison requires metadata tests for curriculum, age, language, sampling and assessment. Rank differences smaller than uncertainty should not support categorical conclusions.[REF-02] [REF-03]

    Part XVIII

    Extended analysis of learning disparities

    57

    Assessment participation and the represented learning population

    assessment participation and the represented learning population cannot be judged without identifying these routes have different meanings. The public account remains incomplete unless it explains how a result calculated only for participating learners may describe their performance accurately while failing to describe the educational system's full learner population. A learning distribution is defined partly by who participates in the assessment. The target population, eligible population, sampled population and assessed population should be reported separately. Learners can be absent because of illness, displacement, school non-attendance, language barriers, disability, conflict, administrative exclusion or ordinary sampling loss.[REF-01] [REF-02] [REF-03]

    assessment participation and the represented learning population requires a decision about similarly, replacement of inaccessible schools can preserve sample size while changing the population represented. A contrary reading would overlook that reports should describe replacement rules and show any material effect on geography or institutional type. Participation should therefore accompany every reported mean, proficiency share or percentile. School exclusion and within-school absence need separate observation where the design permits. If institutions outside the frame differ systematically from included schools, a high learner response rate within sampled schools cannot repair the coverage limitation.

    A defensible account of assessment participation and the represented learning population distinguishes where a numerical adjustment is not defensible, the limitation remains a substantive finding. The public account remains incomplete unless it explains how the absence of learning evidence for a population requiring education attention should not be interpreted as evidence of no disparity. Non-participation should not be assigned a score. Coding an absent learner as below threshold invents performance evidence; removing every absence without comment creates a different bias. Sensitivity analysis can examine plausible bounds or compare known characteristics of participants and non-participants.

    A defensible account of assessment participation and the represented learning population distinguishes exemption practices should be reported by reason and learner group. The evidence must therefore clarify how where an adapted form changes the construct materially, separate interpretation may be necessary. Where it removes an irrelevant barrier, results may belong in the common distribution. The report should explain that judgement rather than treat all adaptations as either incomparable or automatically identical. Accommodation and language arrangements affect participation and result validity. An assessment may permit attendance yet fail to elicit the intended construct if the format, communication or response mode is inaccessible.[REF-05] [REF-10] [REF-12]

    In assessing assessment participation and the represented learning population, authorities must determine a low response among remote schools may require field and service improvement; absence among out-of-school children may require a population-based study and re-entry action; assessment exclusion may require accessible design. This matters because these are not corrections to be made solely through statistical weighting. The responsible education body should identify which population remains unrepresented and when better evidence will be available. A national learning claim should be no broader than that coverage. Public interpretation should connect participation to policy.

    58

    Scale, threshold and distribution comparability

    A defensible account of scale, threshold and distribution comparability distinguishes a threshold adds a substantive judgement about the knowledge or capability learners should demonstrate; it should not be selected merely because it divides the sample conveniently. This matters because comparisons require a clear statement of what the assessment scale represents. A numerical score has meaning through the tasks, response model, scoring rules and population for which interpretation has been supported. Equal numerical differences should not be assumed to represent equal educational differences unless the scale warrants that inference.[REF-02] [REF-03]

    scale, threshold and distribution comparability cannot be judged without identifying the lower part of the distribution is especially important for minimum learning opportunity, while the upper part may reveal whether expansion has altered advanced performance. This matters because these summaries should remain tied to uncertainty and assessment coverage. Means provide one description of the centre and can conceal change elsewhere. The same mean can accompany a compressed distribution, a wide lower tail or polarisation. Reports should therefore consider percentiles, threshold shares and dispersion where technically sound.

    For scale, threshold and distribution comparability, the material distinction is between a change in the percentage above a threshold may reflect learning, scale revision, task composition or population change. The evidence must therefore clarify how if standards are reset, the break should be visible at the year of change and parallel results reported where possible. Descriptions such as basic, adequate or advanced should be treated as definitions under the assessment, not universal attributes of a learner or education system. Threshold comparisons need stable standard-setting and clear labels.

    scale, threshold and distribution comparability cannot be judged without identifying the permissible inference should be stated accordingly. A contrary reading would overlook that cross-language and cross-cultural comparability require evidence, not an assumption that translation has preserved difficulty and meaning. Task familiarity, curriculum exposure and response conventions can affect results. Review should examine translation, adaptation, differential item behaviour and opportunity to learn, while avoiding the claim that every detected difference invalidates the whole assessment. Some comparisons may remain defensible at a broad domain while narrower subscales do not.

    The central question in scale, threshold and distribution comparability is small rank movement can follow changes in participating systems or sampling variation. The resulting interpretation should show why public reporting should avoid categorical language when intervals overlap or scale linkage is weak. The useful question is whether the evidence indicates a material disparity requiring enquiry, what population is affected and which educational condition might be changed. League order does not answer those questions. A rank is usually less informative than the estimated difference, uncertainty and distribution.

    59

    Opportunity to learn and interpretation of achievement gaps

    In assessing opportunity to learn and interpretation of achievement gaps, authorities must determine opportunity does not determine performance completely, and its measurement is imperfect. The institutional consequence follows from whether it nevertheless prevents a learning gap from being attributed solely to learners or households when institutions supplied different educational conditions. Achievement evidence should be interpreted with opportunity to learn. This includes curriculum entitlement, content actually taught, instructional time, teacher availability, language, materials, attendance and access to support.[REF-01] [REF-13] [REF-14]

    The central question in opportunity to learn and interpretation of achievement gaps is teacher self-report may be influenced by recall or expectations; observation covers a short period; work samples are selective. The resulting interpretation should show why agreement across sources strengthens interpretation when dates and populations align. Contradictions can identify local variation or weak measurement and should not be resolved by choosing the account most favourable to the system. The official curriculum establishes intended opportunity, not delivered instruction. Teacher reports, schedules, classroom observation and learner work can add evidence, each with limitations.

    opportunity to learn and interpretation of achievement gaps requires a decision about a learner receiving more hours of poorly organised instruction does not necessarily have greater opportunity in the intended domain. A proportionate conclusion must also recognise that time is therefore one component of the explanatory evidence, not a conversion factor for predicted score. Instructional time should distinguish scheduled, delivered and attended time. Closures, teacher absence, shortened shifts and late entry can reduce delivered or usable time. Total hours also conceal subject allocation and teaching quality.

    The practical standard for opportunity to learn and interpretation of achievement gaps concerns facilities may exist but be inaccessible or unsafe. The public account remains incomplete unless it explains how these conditions should be examined at the level where the learning evidence was collected. National resource averages can obscure concentration of weak provision among the same learners whose scores form the lower tail. Resource indicators should be connected to use. Textbooks delivered to a school may not be available in the relevant language, grade or classroom. Teacher qualifications recorded administratively may not match subject assignment.

    Evidence concerning opportunity to learn and interpretation of achievement gaps should establish where learners received less curriculum exposure, a response may include additional teaching, staff deployment, accessible materials or revised pacing. A contrary reading would overlook that later assessment should test whether opportunity and learning changed. The original disparity and the remedial conditions should remain visible rather than being erased by a new cohort average. Policy conclusions should avoid treating opportunity indicators as excuses for low expectations. Their purpose is to identify conditions within public responsibility and to design support.

    60

    Decomposing disparities within education systems

    In assessing decomposing disparities within education systems, authorities must determine decomposition can describe where variation is concentrated, provided it is not treated as proof of cause. The resulting interpretation should show why a large between-school component may indicate segregation, resource distribution or residential pattern; it does not identify which mechanism operates. A large within-school component can coexist with institutional inequality and should not be read as evidence that schools are irrelevant. A national disparity can reflect differences between regions, schools, classrooms and learners.[REF-01] [REF-09] [REF-14]

    decomposing disparities within education systems cannot be judged without identifying standard errors and models should respect clustering. This matters because small schools and sparsely populated areas may require special treatment, but removal changes the population represented. Reports should state exclusions and avoid presenting a modelled residual as direct observation. The level of analysis should correspond to the sampling and decision structure. Learners are commonly nested within classes and schools, while policies may operate through districts or providers.

    In assessing decomposing disparities within education systems, authorities must determine adjusted measures can answer bounded questions but depend on variables and assumptions. For the learners concerned, the decisive consideration is whether they should not replace the unadjusted learner outcome or become a definitive quality rank. Adjustment for a condition influenced by the school can also remove part of the very effect under review. The purpose and causal assumptions need explanation. Group composition affects school comparisons. A raw school mean combines prior opportunity, intake, mobility, attendance and current teaching.

    decomposing disparities within education systems requires a decision about a small district with a severe gap may need urgent support even though it contributes little to national variance. The public account remains incomplete unless it explains how a populous area with a modest gap may represent many learners. Policy priority should consider educational severity, population, rights and feasibility rather than statistical contribution alone. Maps and rankings, where used elsewhere, should not expose small communities or imply a boundary creates the disparity. Geographical decomposition should preserve absolute numbers and service context.

    A defensible account of decomposing disparities within education systems distinguishes each hypothesis requires additional evidence. A contrary reading would overlook that the decomposition locates questions; it does not authorise blame. Follow-up should state the condition examined, action taken and later learning evidence. The appropriate outcome is a decision agenda. Between-region evidence may lead to allocation review; between-school evidence may lead to staffing, admissions or support enquiry; within-school evidence may lead to classroom, language or accessibility review.

    61

    Bounded conclusions between education systems

    The central question in bounded conclusions between education systems is the institutional consequence follows from whether between-system comparison serves public learning when it identifies patterns, plausible questions and alternative institutional arrangements. The public account remains incomplete unless it explains how it becomes misleading when harmonised labels conceal different programme structures, ages, curricula, languages, participation or assessment conditions. Metadata review should precede numerical comparison. Evidence concerning bounded conclusions between education systems should establish where a material difference cannot be reconciled, the systems may still be described separately without a common rank.[REF-02] [REF-03] [REF-16]

    In assessing bounded conclusions between education systems, authorities must determine within-system distributions can overlap substantially even where means differ. This matters because group composition and population coverage matter. An apparent national advantage may not extend to poor, rural, minority-language or disabled learners. Reports should place distributional evidence beside the system result and avoid using nationality as an explanation. Country and system averages should not be interpreted as attributes of every school or learner.

    bounded conclusions between education systems cannot be judged without identifying a linked scale can support trend if common items or other methods preserve meaning and security, but linkage error should accompany the estimate. The evidence must therefore clarify how a later higher score should not automatically be described as system improvement where the represented population changed materially. Temporal comparison requires stable linkage. Changes in curriculum, assessment mode, participation, sampling frame or system boundaries can create discontinuity.

    The central question in bounded conclusions between education systems is the comparison can identify an option, not guarantee its effect. The public account remains incomplete unless it explains how pilots should state the mechanism and review evidence. Adoption based solely on rank proximity or reputation substitutes imitation for analysis. Policy borrowing should attend to authority, capacity, sequence and context. A practice associated with high performance elsewhere may depend upon teacher preparation, finance, curriculum coherence or social conditions absent in the receiving system.

    The central question in bounded conclusions between education systems is interpretation identifies patterns and limitations. The public account remains incomplete unless it explains how policy consideration proposes further enquiry or action under national authority. Causal judgement requires additional design. This separation permits strong public concern about a learning disparity without false certainty about its source and protects education systems from both complacency and unsupported prescription. The final international statement should distinguish three levels. Recorded fact describes the observed results under stated methods.

    62

    Uncertainty, materiality and the duty to respond

    The practical standard for uncertainty, materiality and the duty to respond concerns the public report should identify which uncertainty could alter the decision and which does not affect the direction of urgent protection. A proportionate conclusion must also recognise that uncertainty should qualify a disparity claim without neutralising it. Sampling error, non-response, scale linkage, classification and model choice each affect the range of defensible conclusions. They should be described separately because they support different remedies. A larger sample may reduce sampling error; it does not correct systematic exclusion or an invalid construct. More decimal places cannot repair weak coverage.[REF-02] [REF-03]

    A defensible account of uncertainty, materiality and the duty to respond distinguishes additional diagnostic support can be introduced and reviewed more readily than a high-stakes classification of schools or learners. A contrary reading would overlook that statistical significance should not substitute for educational materiality. A small estimated difference may be precise but have limited practical consequence, while an uncertain large gap affecting a protected minimum may require immediate enquiry. Materiality should consider the knowledge or capability involved, the number of learners, distribution, duration and consequences for later progression. The threshold for action also depends upon reversibility.

    Institutional action on uncertainty, materiality and the duty to respond should be tested against where participation is selective, improve coverage and examine barriers. The resulting interpretation should show why where opportunity to learn differs, address time, teachers, curriculum, language, accessibility or materials. Where scale comparability is weak, improve the assessment before publishing ranks. Where a group disparity persists under several definitions, investigate institutional mechanisms without attributing cause to identity. Every response should name the competent body, intended population, resources and review date. A remedy should match the evidence.

    Comparative interpretation of uncertainty, materiality and the duty to respond depends upon reports should preserve the original estimate, method and limitation, then state why revision occurred. For the learners concerned, the decisive consideration is whether silent replacement makes apparent improvement impossible to distinguish from correction. A national evidence system demonstrates strength when it can acknowledge uncertainty, improve measurement and amend policy. The final test is whether learners receive better educational opportunity and whether remaining disparities continue to be visible rather than whether one annual figure becomes more favourable. Later evidence should be capable of changing the conclusion.

    Part XIX

    Targeted improvement planning

    63

    Defining a persistent learning gap

    A defensible account of defining a persistent learning gap distinguishes the baseline should preserve the underlying distribution and absolute numbers. A contrary reading would overlook that if the assessment population excludes learners most at risk, the gap is not adequately defined. The plan should state which observation is direct, which is estimated and which remains unknown. The improvement question concerns a sustained disparity in a declared learning domain, population and period. It should be stated with the learner population, educational domain, geography and evidence period. A national average is insufficient when the plan addresses a local or group disparity.[REF-01] [REF-02]

    Evidence concerning defining a persistent learning gap should establish authorities should examine curriculum, teacher availability, time, attendance, materials, assessment access and learner support. The institutional consequence follows from whether alternative explanations should remain open until evidence discriminates among them. This discipline prevents a targeted plan from attaching deficit to learners instead of changing institutions. The principal error is a fluctuating score or one cohort difference being treated as proof of persistence. Avoiding it requires evidence on both learning and the opportunity supplied. Achievement differences may be associated with poverty, language, disability or residence, but those characteristics are not instructional mechanisms.[REF-03] [REF-06]

    defining a persistent learning gap cannot be judged without identifying observation and work samples can explain classroom conditions without estimating national prevalence. The resulting interpretation should show why learner and teacher accounts can identify barriers. Agreement adds confidence after dates and definitions align; contradiction should guide further enquiry rather than selective reporting. The required evidence includes baseline, assessment coverage, uncertainty, distribution and opportunity to learn. Each source should be used within scope. Administrative records can describe staffing and participation while omitting non-enrolled learners. Assessment provides bounded learning evidence subject to coverage and validity.[REF-08] [REF-09]

    Evidence concerning defining a persistent learning gap should establish the governing action is to select a gap that is educationally material and within public influence. The evidence must therefore clarify how implementation should identify who acts, with what authority, resources and deadline. Dependencies should be sequenced. Teacher guidance without planning time, materials without accessible use, or tutoring without safe attendance cannot deliver the expected mechanism. Review of defining a persistent learning gap is credible only where it explains local adaptation should remain possible within a common substantive condition. The evidence must therefore clarify how any departure should be recorded with its reason and expected learner consequence.[REF-13] [REF-14]

    defining a persistent learning gap cannot be judged without identifying a plan can improve its average by reaching learners closest to a threshold while leaving those farthest behind. The resulting interpretation should show why targets should therefore include the least-served position and protection against exclusion. Small groups require confidentiality and careful precision, not disappearance from review. Distributional review should ask who is eligible, offered support, participates, receives the intended intensity and demonstrates a later response. Gender, household resources, disability, language, residence and prior opportunity may intersect.[REF-17] [REF-19]

    Comparative interpretation of defining a persistent learning gap depends upon leadership should route obstacles to bodies able to change staffing, finance, curriculum or assessment rather than leaving every correction to the classroom. The resulting interpretation should show why professional capability is central. Teachers need subject knowledge, worked examples, diagnostic interpretation and protected time to collaborate. Moderation should examine evidence and reasoning rather than force identical decisions for unlike cases. Staff workload and turnover should be monitored. A plan relying on a few exceptional individuals is not institutionally secure.[REF-23] [REF-24]

    In assessing defining a persistent learning gap, authorities must determine a favourable later result may reflect population or assessment change and should be tested against the baseline metadata. The institutional consequence follows from whether where the intervention is ineffective, adaptation or cessation is responsible improvement. Where benefit depends on temporary support, institutionalisation requires recurrent finance and ordinary ownership. Completion is demonstrated by stronger learning opportunity and a functioning correction route, not by the end of a project. The public account should distinguish authorisation, delivery, use and learning consequence. It should report limitations, adverse effects and unresolved learners.[REF-01] [REF-02]

    64

    From diagnosis to an intervention hypothesis

    For from diagnosis to an intervention hypothesis, the material distinction is between it should be stated with the learner population, educational domain, geography and evidence period. The public account remains incomplete unless it explains how a national average is insufficient when the plan addresses a local or group disparity. The baseline should preserve the underlying distribution and absolute numbers. If the assessment population excludes learners most at risk, the gap is not adequately defined. The plan should state which observation is direct, which is estimated and which remains unknown. The improvement question concerns an explicit account of the condition expected to change learning.[REF-03] [REF-06]

    Institutional action on from diagnosis to an intervention hypothesis should be tested against achievement differences may be associated with poverty, language, disability or residence, but those characteristics are not instructional mechanisms. The institutional consequence follows from whether authorities should examine curriculum, teacher availability, time, attendance, materials, assessment access and learner support. Alternative explanations should remain open until evidence discriminates among them. This discipline prevents a targeted plan from attaching deficit to learners instead of changing institutions. The principal error is group identity or low performance being mistaken for a causal explanation. Avoiding it requires evidence on both learning and the opportunity supplied.[REF-08] [REF-09]

    Public responsibility for from diagnosis to an intervention hypothesis begins with assessment provides bounded learning evidence subject to coverage and validity. A proportionate conclusion must also recognise that observation and work samples can explain classroom conditions without estimating national prevalence. Learner and teacher accounts can identify barriers. Agreement adds confidence after dates and definitions align; contradiction should guide further enquiry rather than selective reporting. The required evidence includes curriculum exposure, teaching practice, language, time, materials and support. Each source should be used within scope. Administrative records can describe staffing and participation while omitting non-enrolled learners.[REF-13] [REF-14]

    The central question in from diagnosis to an intervention hypothesis is any departure should be recorded with its reason and expected learner consequence. The resulting interpretation should show why the governing action is to state the mechanism and plausible alternatives before choosing activity. Implementation should identify who acts, with what authority, resources and deadline. Dependencies should be sequenced. Teacher guidance without planning time, materials without accessible use, or tutoring without safe attendance cannot deliver the expected mechanism. Local adaptation should remain possible within a common substantive condition.[REF-17] [REF-19]

    65

    Designing the targeted plan

    Evidence concerning designing the targeted plan should establish the baseline should preserve the underlying distribution and absolute numbers. A proportionate conclusion must also recognise that if the assessment population excludes learners most at risk, the gap is not adequately defined. The plan should state which observation is direct, which is estimated and which remains unknown. The improvement question concerns a bounded sequence connecting resources and actions to learner-facing change. It should be stated with the learner population, educational domain, geography and evidence period. A national average is insufficient when the plan addresses a local or group disparity.[REF-08] [REF-09]

    In assessing designing the targeted plan, authorities must determine this discipline prevents a targeted plan from attaching deficit to learners instead of changing institutions. A contrary reading would overlook that the principal error is a list of activities replacing a coherent implementation logic. Avoiding it requires evidence on both learning and the opportunity supplied. Achievement differences may be associated with poverty, language, disability or residence, but those characteristics are not instructional mechanisms. Authorities should examine curriculum, teacher availability, time, attendance, materials, assessment access and learner support. Alternative explanations should remain open until evidence discriminates among them.[REF-13] [REF-14]

    In assessing designing the targeted plan, authorities must determine each source should be used within scope. A proportionate conclusion must also recognise that administrative records can describe staffing and participation while omitting non-enrolled learners. Assessment provides bounded learning evidence subject to coverage and validity. Observation and work samples can explain classroom conditions without estimating national prevalence. Learner and teacher accounts can identify barriers. Agreement adds confidence after dates and definitions align; contradiction should guide further enquiry rather than selective reporting. The required evidence includes responsibility, staff capability, learner support, milestones and correction.[REF-17] [REF-19]

    The practical standard for designing the targeted plan concerns implementation should identify who acts, with what authority, resources and deadline. The public account remains incomplete unless it explains how dependencies should be sequenced. Teacher guidance without planning time, materials without accessible use, or tutoring without safe attendance cannot deliver the expected mechanism. Local adaptation should remain possible within a common substantive condition. Any departure should be recorded with its reason and expected learner consequence. The governing action is to choose a feasible intensity and protect the common educational entitlement.[REF-23] [REF-24]

    66

    Resourcing equitable implementation

    resourcing equitable implementation cannot be judged without identifying the plan should state which observation is direct, which is estimated and which remains unknown. The public account remains incomplete unless it explains how the improvement question concerns the staff, time, materials, accessibility and finance required by the selected mechanism. It should be stated with the learner population, educational domain, geography and evidence period. A national average is insufficient when the plan addresses a local or group disparity. The baseline should preserve the underlying distribution and absolute numbers. If the assessment population excludes learners most at risk, the gap is not adequately defined.[REF-13] [REF-14]

    The central question in resourcing equitable implementation is authorities should examine curriculum, teacher availability, time, attendance, materials, assessment access and learner support. The public account remains incomplete unless it explains how alternative explanations should remain open until evidence discriminates among them. This discipline prevents a targeted plan from attaching deficit to learners instead of changing institutions. The principal error is weak institutions being expected to implement with the same nominal allocation. Avoiding it requires evidence on both learning and the opportunity supplied. Achievement differences may be associated with poverty, language, disability or residence, but those characteristics are not instructional mechanisms.[REF-17] [REF-19]

    The practical standard for resourcing equitable implementation concerns observation and work samples can explain classroom conditions without estimating national prevalence. A contrary reading would overlook that learner and teacher accounts can identify barriers. Agreement adds confidence after dates and definitions align; contradiction should guide further enquiry rather than selective reporting. The required evidence includes recurrent cost, teacher workload, additional need and household burden. Each source should be used within scope. Administrative records can describe staffing and participation while omitting non-enrolled learners. Assessment provides bounded learning evidence subject to coverage and validity.[REF-23] [REF-24]

    The central question in resourcing equitable implementation is any departure should be recorded with its reason and expected learner consequence. The resulting interpretation should show why the governing action is to direct greater support where barriers and implementation costs are greater. Implementation should identify who acts, with what authority, resources and deadline. Dependencies should be sequenced. Teacher guidance without planning time, materials without accessible use, or tutoring without safe attendance cannot deliver the expected mechanism. Local adaptation should remain possible within a common substantive condition.[REF-01] [REF-02]

    67

    Monitoring reach, quality and learning response

    Public responsibility for monitoring reach, quality and learning response begins with the plan should state which observation is direct, which is estimated and which remains unknown. This matters because the improvement question concerns evidence that the intended learners received the intervention as designed and benefited educationally. It should be stated with the learner population, educational domain, geography and evidence period. A national average is insufficient when the plan addresses a local or group disparity. The baseline should preserve the underlying distribution and absolute numbers. If the assessment population excludes learners most at risk, the gap is not adequately defined.[REF-17] [REF-19]

    The practical standard for monitoring reach, quality and learning response concerns authorities should examine curriculum, teacher availability, time, attendance, materials, assessment access and learner support. The public account remains incomplete unless it explains how alternative explanations should remain open until evidence discriminates among them. This discipline prevents a targeted plan from attaching deficit to learners instead of changing institutions. The principal error is participation counts being treated as learning evidence. Avoiding it requires evidence on both learning and the opportunity supplied. Achievement differences may be associated with poverty, language, disability or residence, but those characteristics are not instructional mechanisms.[REF-23] [REF-24]

    The practical standard for monitoring reach, quality and learning response concerns learner and teacher accounts can identify barriers. For the learners concerned, the decisive consideration is whether agreement adds confidence after dates and definitions align; contradiction should guide further enquiry rather than selective reporting. The required evidence includes eligibility, offer, take-up, dosage, teaching quality, work and assessment. Each source should be used within scope. Administrative records can describe staffing and participation while omitting non-enrolled learners. Assessment provides bounded learning evidence subject to coverage and validity. Observation and work samples can explain classroom conditions without estimating national prevalence.[REF-01] [REF-02]

    For monitoring reach, quality and learning response, the material distinction is between teacher guidance without planning time, materials without accessible use, or tutoring without safe attendance cannot deliver the expected mechanism. The evidence must therefore clarify how local adaptation should remain possible within a common substantive condition. Any departure should be recorded with its reason and expected learner consequence. The governing action is to combine timely implementation evidence with valid learning review. Implementation should identify who acts, with what authority, resources and deadline. Dependencies should be sequenced.[REF-03] [REF-06]

    68

    Adaptation, institutionalisation and exit

    In assessing adaptation, institutionalisation and exit, authorities must determine a national average is insufficient when the plan addresses a local or group disparity. A proportionate conclusion must also recognise that the baseline should preserve the underlying distribution and absolute numbers. If the assessment population excludes learners most at risk, the gap is not adequately defined. The plan should state which observation is direct, which is estimated and which remains unknown. The improvement question concerns reasoned decisions to continue, change, scale or end the plan. It should be stated with the learner population, educational domain, geography and evidence period.[REF-23] [REF-24]

    A defensible account of adaptation, institutionalisation and exit distinguishes this discipline prevents a targeted plan from attaching deficit to learners instead of changing institutions. A contrary reading would overlook that the principal error is temporary measures persisting without benefit or disappearing before durable capability exists. Avoiding it requires evidence on both learning and the opportunity supplied. Achievement differences may be associated with poverty, language, disability or residence, but those characteristics are not instructional mechanisms. Authorities should examine curriculum, teacher availability, time, attendance, materials, assessment access and learner support. Alternative explanations should remain open until evidence discriminates among them.[REF-01] [REF-02]

    Comparative interpretation of adaptation, institutionalisation and exit depends upon agreement adds confidence after dates and definitions align; contradiction should guide further enquiry rather than selective reporting. The evidence must therefore clarify how the required evidence includes thresholds, adverse effects, unresolved cases, recurrent ownership and later evidence. Each source should be used within scope. Administrative records can describe staffing and participation while omitting non-enrolled learners. Assessment provides bounded learning evidence subject to coverage and validity. Observation and work samples can explain classroom conditions without estimating national prevalence. Learner and teacher accounts can identify barriers.[REF-03] [REF-06]

    adaptation, institutionalisation and exit cannot be judged without identifying any departure should be recorded with its reason and expected learner consequence. For the learners concerned, the decisive consideration is whether the governing action is to retain useful capability while ending ineffective or inequitable arrangements. Implementation should identify who acts, with what authority, resources and deadline. Dependencies should be sequenced. Teacher guidance without planning time, materials without accessible use, or tutoring without safe attendance cannot deliver the expected mechanism. Local adaptation should remain possible within a common substantive condition.[REF-08] [REF-09]

    Part XX

    National translation after adoption of the 2030 Agenda

    69

    Fixing the contemporaneous institutional baseline

    Institutional planning should begin from the position established by 4 December 2017. The Incheon Declaration expressed the education community's commitment to inclusive and equitable quality education and lifelong learning; Addis established a wider financing framework; the General Assembly adopted the 2030 Agenda; and the Education 2030 Framework for Action supplied an implementation reference available by the cut-off. These texts provide direction without proving that national or institutional delivery has occurred.[REF-25] [REF-26] [REF-27] [REF-28]

    fixing the contemporaneous institutional baseline requires a decision about each country should identify which obligations and targets can be acted upon under existing authority, which require legislative or administrative change, and which depend on clarification through later competent decisions. This matters because unsettled detail should be recorded as such rather than filled by anticipation. The distinction between adoption and implementation is essential. Adoption establishes an agreed direction and permits governments to begin alignment. It does not prove that national law, plans, budgets, information or delivery arrangements already satisfy the commitment.[REF-07] [REF-12] [REF-27]

    In assessing fixing the contemporaneous institutional baseline, authorities must determine this avoids two risks: abandoning useful evidence because terminology changed, and claiming continuity where a new target has broader scope or different educational substance. For the learners concerned, the decisive consideration is whether the baseline should preserve existing education commitments and national evidence. The new agenda does not erase the right to education, the unfinished Education for All undertaking, programme structures or established statistical series. National authorities should map the newly adopted targets against those instruments and against the education stages, populations and institutions already in law and plans.[REF-08] [REF-09] [REF-25]

    Public responsibility for fixing the contemporaneous institutional baseline begins with the register should be revised openly as competent bodies act. The evidence must therefore clarify how a dated commitment register can support institutional accuracy. For each relevant proposition, it should record the adopting body, date, legal or policy status, national authority, present implementing instrument and unresolved question. This is not an administrative inventory for its own sake. It prevents a proposed measure from being represented as an obligation, a declaration from being treated as proof of delivery, and a later decision from being projected backwards.[REF-17] [REF-19] [REF-27]

    Public responsibility for fixing the contemporaneous institutional baseline begins with a precise account strengthens credibility because it makes clear which choices belong to national democratic and administrative processes and which follow directly from adopted commitments. This matters because it also establishes a reliable point from which later implementation can be judged. Public communication should use the same discipline. Governments can state that the 2030 Agenda has been adopted and that national alignment is commencing. They should not state that every indicator, national milestone or implementation mechanism has already been internationally settled.[REF-25] [REF-26] [REF-27]

    70

    Selecting priorities without narrowing the commitment

    selecting priorities without narrowing the commitment requires a decision about a first improvement priority should be chosen because evidence shows a serious and remediable break, not because other elements have ceased to matter. A contrary reading would overlook that the breadth of Goal 4 requires sequencing, not selective abandonment. A national plan cannot improve every condition simultaneously, yet it should retain a complete map of early childhood, primary and secondary education, technical and vocational learning, tertiary participation, adult learning, relevant skills, equality, literacy, learning environments, scholarships and teachers as they appear in the adopted targets.[REF-25] [REF-27]

    In assessing selecting priorities without narrowing the commitment, authorities must determine a priority that scores highly on visibility but weakly on consequence or equity should be reconsidered. For the learners concerned, the decisive consideration is whether the reasons for selection should be published beside the evidence and known limitations. Priority selection should apply four tests. The entitlement test asks which population and educational condition are at stake. The consequence test asks the scale and severity of the denial. The actionability test asks whether a competent authority has a plausible means of change. The equity test asks whether the measure will reach learners farthest from the secured opportunity.[REF-01] [REF-03] [REF-12]

    The practical standard for selecting priorities without narrowing the commitment concerns it should also distinguish poor data from satisfactory conditions. The evidence must therefore clarify how where the least-served population is weakly observed, strengthening coverage can itself become an immediate priority while urgent service evidence supports proportionate protection. National averages should not determine the sequence alone. A moderate national gap may conceal acute failure in one district or population; a large aggregate shortfall may require broad system expansion alongside targeted support. Analysis should retain the national level, absolute number affected, subnational and group distributions and the minimum educational floor.[REF-02] [REF-06] [REF-24]

    Public responsibility for selecting priorities without narrowing the commitment begins with adult learning may require flexible provision, recognition and learner support. The public account remains incomplete unless it explains how the plan should make these dependencies visible and decide which must precede or accompany the selected measure. Otherwise a high-level commitment can be converted into an isolated activity unable to change the learner-facing condition. Dependencies should influence sequence. An assessment reform cannot improve learning without curriculum alignment, teacher capability and participation. Expansion of secondary places may depend on primary completion, trained staff and facilities.[REF-03] [REF-16] [REF-18]

    Evidence concerning selecting priorities without narrowing the commitment should establish the complete commitment map should remain in view so that repeated concentration on one readily measured target does not produce silent neglect of lifelong learning, equality, educational quality or populations outside formal schooling. The evidence must therefore clarify how selection should remain revisable. New evidence may show that the diagnosed mechanism was wrong, that another population is more severely affected or that implementation capacity is insufficient. Revision is not a retreat from ambition when reasons and consequences are published. It is a condition of responsible improvement.[REF-08] [REF-25] [REF-27]

    71

    From global target to national improvement proposition

    The central question in from global target to national improvement proposition is the narrower proposition remains connected to the universal commitment and should not be mistaken for its completion. For the learners concerned, the decisive consideration is whether a national improvement proposition should translate a broad target into a bounded statement of change. It should name the population, present educational condition, institutional mechanism, responsible body, resources, time and evidence of success. For example, a commitment to equitable quality education is too broad to guide one implementation decision; a proposition to improve regular attendance for a defined remote population through transport, staffing and calendar changes can be examined and corrected.[REF-07] [REF-12] [REF-27]

    A defensible account of from global target to national improvement proposition distinguishes a low completion rate may reflect late entry, repetition, household cost, distance, school safety, language, disability exclusion, teacher shortage or unreliable records. This matters because these mechanisms require different responses. A government should compare administrative, household, assessment and local service evidence and state where inference remains uncertain. Consultation with teachers, learners and communities can identify mechanisms, but it should not replace representative population evidence when prevalence is claimed. Diagnosis should precede instrument choice.[REF-01] [REF-03] [REF-24]

    The central question in from global target to national improvement proposition is additional materials will not improve learning if teachers lack time or knowledge to use them; professional guidance will not improve attendance where transport is decisive; a new indicator will not correct exclusion without authority and resources. For the learners concerned, the decisive consideration is whether the plan should identify necessary dependencies and foreseeable adverse effects. It should specify which part of the hypothesis is established by evidence and which remains to be tested during implementation. The intervention hypothesis should explain how the proposed measure changes the barrier.[REF-06] [REF-12] [REF-16]

    A defensible account of from global target to national improvement proposition distinguishes adaptation is legitimate where it makes the commitment operational and comparable over time. A proportionate conclusion must also recognise that it is not legitimate where it narrows the entitled population, lowers expectations for disadvantaged groups or converts learning into attendance alone. The public record should show the relationship between the global target, national definition and selected measure, including any material difference from an existing series. National adaptation should preserve educational substance. Targets may need national definitions for programme levels, age groups, language and institutional responsibility.[REF-02] [REF-10] [REF-25]

    In assessing from global target to national improvement proposition, authorities must determine authorities should state what evidence would justify continuation, expansion, adaptation or cessation and when that decision will occur. This matters because a pilot that continues because it attracts support rather than because it changes the intended condition is not an improvement method. Conversely, an intervention should not be abandoned merely because early outcomes are uncertain where delivery has not reached intended intensity. Review must distinguish theory failure, implementation failure, measurement weakness and insufficient time. The proposition should end with a decision rule.[REF-08] [REF-11] [REF-23]

    72

    Aligning authority, finance and professional capability

    aligning authority, finance and professional capability cannot be judged without identifying one public owner should remain answerable for whether the learner-facing condition changes. The institutional consequence follows from whether implementation requires a chain of competent authority. National policy may set the priority, but regional administrations, municipalities, schools, training institutions or other bodies may control staffing, facilities and learner support. The plan should allocate each function to the body able to perform it and identify escalation where local authority is insufficient. Coordination should not allow responsibility to become diffuse.[REF-12] [REF-17] [REF-19]

    Comparative interpretation of aligning authority, finance and professional capability depends upon staff preparation, salaries, accessible materials, transport, maintenance, guidance, assessment, evidence and review may all be necessary. The institutional consequence follows from whether the Addis Ababa Action Agenda places national action within a broader financing context, but an international commitment does not supply a national cost estimate. Authorities should identify recurrent and capital requirements, the source and timing of funds, distribution rules and the conditions for continuity after temporary support. Finance should cover the complete intervention rather than visible start-up items.[REF-15] [REF-26]

    aligning authority, finance and professional capability cannot be judged without identifying equal per-learner funding can reproduce inequality where remoteness, disability, language, insecurity or weak infrastructure makes adequate provision more expensive. The public account remains incomplete unless it explains how formulae should state which need factors are recognised and should be checked against actual receipt. Announced expenditure is not proof of delivery: review should trace authorization, transfer, institutional use and learner consequence. Where households continue to bear material cost, formal fee policy should not be represented as full accessibility. Allocation should respond to unequal cost and starting capacity.[REF-06] [REF-10] [REF-13]

    aligning authority, finance and professional capability requires a decision about professional learning should provide subject substance, practical examples, time for collaboration and a route to support. The resulting interpretation should show why staffing and turnover should be examined in the locations expected to implement first. A plan dependent on exceptional individuals or uncompensated workload is not institutionally secure and can deepen disparity between strong and weak institutions. Teachers and institutional leaders require capability proportionate to the change. A new curriculum, assessment or inclusion expectation needs more than notification.[REF-01] [REF-18]

    Institutional action on aligning authority, finance and professional capability should be tested against its success should be judged by national capability and equitable learner opportunity, not the duration or visibility of the supporting activity. This matters because international cooperation should strengthen ordinary national capacity. External support may finance initial expansion, evidence, technical work or regional learning, but roles, conditions and exit should be clear. Parallel activities and reporting arrangements can fragment public authority. Assistance should align with a national improvement proposition, use compatible records and establish how essential functions will enter recurrent provision.[REF-22] [REF-26] [REF-27]

    73

    Monitoring delivery, reach and educational consequence

    Comparative interpretation of monitoring delivery, reach and educational consequence depends upon these stages should not be collapsed. A proportionate conclusion must also recognise that a budget can be executed without materials arriving, a programme can operate without reaching disadvantaged learners, and participation can rise without improvement in learning or progression. Monitoring should follow the causal sequence of the improvement proposition. Inputs show whether resources and staff were available; delivery evidence shows whether the measure operated; reach shows which eligible learners participated and with what intensity; educational evidence shows whether the intended condition changed.[REF-03] [REF-06] [REF-12]

    monitoring delivery, reach and educational consequence requires a decision about apparent improvement caused by revised population estimates, wider institutional reporting or changed assessment participation should be separated from educational change. The resulting interpretation should show why comparable trends are valuable, but continuity should not be asserted where concepts differ materially. The baseline should retain numerator, denominator, population, date, geography, definition and exclusions. Where a new target requires a new measure, authorities should preserve the preceding series and identify any break.[REF-02] [REF-23] [REF-24]

    In assessing monitoring delivery, reach and educational consequence, authorities must determine small numbers may require controlled access or combined reporting, but the affected population and public responsibility should not disappear. This matters because equity monitoring should show eligibility, offer, take-up, attendance, completion and outcome for relevant groups and places. A plan can improve its average by reaching learners already closest to the desired condition. It should therefore report the least-served position and protect against exclusion or deterioration. Disaggregation must remain lawful, meaningful and safe.[REF-01] [REF-10] [REF-13]

    In assessing monitoring delivery, reach and educational consequence, authorities must determine a favourable mean among tested learners cannot represent those outside school or absent from assessment. The resulting interpretation should show why classroom observation, work samples and learner accounts can illuminate mechanism without establishing national prevalence. Agreement across sources strengthens a conclusion only after definitions and dates align; disagreement should guide investigation rather than selective reporting. Learning evidence requires population coverage and opportunity to learn. Assessment results should identify the domain, eligible population, participation and exclusions.[REF-03] [REF-24]

    Review of monitoring delivery, reach and educational consequence is credible only where it explains a concise account should state what was implemented, who received it, what changed, what remains uncertain, which adverse effects occurred and whether the measure will continue, adapt, expand or cease. The resulting interpretation should show why the next review date and responsible authority should be visible. This makes monitoring a means of correction rather than an obligation to produce favourable figures. The adopted agenda gains national credibility when evidence can alter action and expose populations who remain underserved. Public reporting should connect the result to a decision.[REF-08] [REF-25] [REF-27]

    74

    First national decisions after adoption

    The practical standard for first national decisions after adoption concerns governments can designate a competent coordinating authority, preserve existing sector responsibilities, assemble a dated commitment register and commission a baseline review. The public account remains incomplete unless it explains how the review should map every adopted education target against national law, plans, budgets and evidence. It should identify urgent gaps and unsettled definitions separately. This creates a disciplined bridge between global adoption and national action. The immediate national decision is to establish governance for alignment without pretending that implementation detail is complete.[REF-25] [REF-27]

    Comparative interpretation of first national decisions after adoption depends upon where credible evidence reveals severe exclusion or harm, protective action need not await a perfect estimate. A contrary reading would overlook that longer-term allocation, however, should be reviewed as coverage improves. The plan should guard against choosing only learners and institutions most likely to produce rapid favourable results. A first priority should be small enough for accountable action and important enough to change educational opportunity. It should name the population and condition, state why it takes precedence, and preserve the wider commitment map.[REF-01] [REF-06] [REF-12]

    first national decisions after adoption cannot be judged without identifying the national authority should state the source, timing and distribution of finance and how external cooperation relates to ordinary provision. A proportionate conclusion must also recognise that a financing gap should be described rather than hidden through reduced educational substance or transfer of cost to poor households. The first budget decision should identify recurrent implications. Temporary finance may permit testing, but teachers, learner support, accessible facilities, maintenance and evidence cannot be sustained by an announcement.[REF-15] [REF-26]

    For first national decisions after adoption, the material distinction is between it can state that the principal education, financing and implementation texts named in this report were available by the cut-off. The public account remains incomplete unless it explains how it should also state that later indicator settlements, technical revisions and results are outside the record. This protects the difference between an adopted commitment, a national policy choice, an institutional plan and later statistical clarification. Accurate status is a condition of accountable planning, not a reason for delay. The first public report should be candid about chronology.[REF-25] [REF-26] [REF-27] [REF-28]

    A defensible account of first national decisions after adoption distinguishes the transition from global commitment to national improvement is complete only when ordinary institutions can sustain the changed condition and correct foreseeable departures; the end of a project or reporting period is not evidence of that result. This matters because the first review should test whether institutions learned, not merely whether a plan was issued. It should examine authority, delivery, reach, professional capability, finance, learner experience and early educational consequence. It should record adverse findings and adapt the measure where the hypothesis or delivery proves weak.[REF-08] [REF-12] [REF-23]

    Part XXI

    Institutional improvement plans aligned with Education 2030

    75

    Institutional mandate and scope

    institutional mandate and scope cannot be judged without identifying the plan should therefore map each selected Education 2030 priority to the institutional function that can materially influence it and identify dependencies requiring action at another level. This matters because an institutional improvement plan should begin by identifying the authority under which the institution acts and the population for whom it is responsible. A school, training provider, university, local administration or other body does not implement the whole global agenda alone. It contributes within its mandate, while national authorities retain responsibilities for law, finance, curriculum, workforce and equitable system provision.[REF-12] [REF-25] [REF-28]

    In assessing institutional mandate and scope, authorities must determine an institution may seek more regular participation, stronger foundational learning, safer and more accessible facilities, better transition, improved teacher support or a more equitable adult-learning offer. A contrary reading would overlook that the statement should name the eligible population, current condition, intended change and period. A global target supplies direction, but the institutional proposition must be narrow enough for authority, resources and evidence to be assigned. Scope should be expressed as a learner-facing condition rather than a broad aspiration.[REF-07] [REF-27] [REF-28]

    For institutional mandate and scope, the material distinction is between selecting one priority does not authorize deterioration in another essential condition. For the learners concerned, the decisive consideration is whether safeguards should state which minimums must be protected during implementation. Institutional plans should preserve the relationship among access, equity and quality. A plan concerned with learning cannot ignore assessment participation or opportunity to learn; a plan concerned with enrolment cannot treat registration as regular attendance; a plan concerned with completion should examine the educational substance and recognised value of the programme.[REF-01] [REF-03] [REF-09]

    The central question in institutional mandate and scope is the choice of timetable, professional-learning arrangement, local partnership or diagnostic method may permit adaptation. The evidence must therefore clarify how this distinction allows institutions to respond to context without treating national variation or resource constraint as authority to lower the substantive entitlement of disadvantaged learners. Any requested exception should identify the legal basis, expected consequence and competent approving body. The plan should distinguish obligations from discretionary methods. Rights, national requirements and adopted policy establish conditions that the institution must respect.[REF-08] [REF-10] [REF-28]

    The practical standard for institutional mandate and scope concerns participation can reveal barriers and test whether proposed measures are workable; it should not transfer public responsibility to those affected by weak provision. The public account remains incomplete unless it explains how the decision record should show how evidence and consultation altered the priority. Where the institution lacks authority over a material dependency, the plan should route the issue to the responsible authority and retain it as an unresolved risk rather than quietly narrow the objective. Governance should identify one accountable owner and the roles of staff, learners, communities and partner bodies.[REF-17] [REF-19] [REF-25]

    76

    Diagnostic review and priority selection

    Evidence concerning diagnostic review and priority selection should establish the purpose is to locate the first material break rather than compile every available statistic. The evidence must therefore clarify how administrative records describe registered learners and delivery; household or community evidence may reveal people outside the institution; assessment and work samples describe bounded aspects of learning. Diagnostic review should begin with the institution's eligible population and educational course. Entry, attendance, progression, completion, learning and transition should be mapped together with staffing, instructional time, facilities, accessibility and learner support.[REF-02] [REF-03] [REF-24]

    In assessing diagnostic review and priority selection, authorities must determine small groups require caution about precision and disclosure, but the service barrier should remain visible. For the learners concerned, the decisive consideration is whether where population estimates are weak, qualitative and case evidence may justify immediate correction without being represented as a prevalence estimate. The review should compare level, distribution and trend. A stable institutional mean can conceal deterioration among one programme or learner group. Sex, household constraint, disability, language, residence, displacement and prior opportunity may be relevant, subject to lawful collection and protection.[REF-06] [REF-10] [REF-13]

    The central question in diagnostic review and priority selection is low attendance may be associated with transport, cost, safety, calendar, health, discrimination or inaccessible instruction. The resulting interpretation should show why weak learning may reflect interrupted participation, limited curriculum exposure, teacher absence, language, assessment design or inadequate support. A descriptive difference cannot rank these explanations. The institution should gather the smallest additional evidence needed for a practical decision and should identify explanations requiring action outside its authority. Diagnosis should distinguish symptom from mechanism.[REF-01] [REF-05] [REF-12]

    For diagnostic review and priority selection, the material distinction is between a condition affecting fewer learners may warrant immediate action if the consequence is severe or violates a minimum entitlement. The institutional consequence follows from whether conversely, a large but weakly measured difference may first require stronger evidence. The plan should explain why the selected priority takes precedence, which learners are expected to benefit and how the rest of the institutional responsibility remains monitored. Ease of measurement alone is not a sufficient selection criterion. Priority selection should consider severity, scale, equity, feasibility and dependency.[REF-07] [REF-16] [REF-28]

    diagnostic review and priority selection requires a decision about a baseline reconstructed after results are known is vulnerable to selective interpretation. For the learners concerned, the decisive consideration is whether the plan should approve the baseline and revision rule before judging change. The baseline should be preserved at the point of decision. Population, definitions, source completeness, date, uncertainty and any exclusions should travel with the starting value. Where the institution changes its record system or assessment during implementation, overlap evidence should be retained and breaks marked.[REF-03] [REF-23] [REF-24]

    77

    Improvement proposition and implementation design

    improvement proposition and implementation design cannot be judged without identifying activities should not be listed without the reason they are expected to alter learner opportunity. The resulting interpretation should show why the improvement proposition should state how a defined measure is expected to change the diagnosed condition. It should connect an action to an institutional mechanism: additional instructional support to a documented curriculum gap; revised attendance arrangements to a known timing or transport barrier; accessible materials and assessment to exclusion of learners with disabilities; professional support to weak subject teaching.[REF-10] [REF-12] [REF-28]

    Comparative interpretation of improvement proposition and implementation design depends upon where a dependency lies with another authority, agreement and escalation should precede large-scale implementation. A proportionate conclusion must also recognise that dependencies should be sequenced. Teacher guidance may require curriculum clarification, planning time, examples and leadership support. New materials require procurement, accessible formats, distribution and maintenance. Learner support may require eligibility rules, communication and protection of personal information. A plan should identify the latest point at which a missing dependency can be corrected without compromising delivery.[REF-17] [REF-18] [REF-25]

    The practical standard for improvement proposition and implementation design concerns attendance at one professional meeting, receipt of materials or registration in support does not prove sustained use. A proportionate conclusion must also recognise that monitoring should capture actual exposure and reasons for incomplete delivery while avoiding burdens so heavy that they displace teaching or service. Implementation intensity should be specified. An institution should know which learners or staff receive the measure, how frequently, for how long and with what expected standard. Without intensity, non-response cannot be distinguished from weak delivery.[REF-03] [REF-06] [REF-16]

    The practical standard for improvement proposition and implementation design concerns it becomes inequitable where learners facing disadvantage receive reduced content, weaker staff or an unrecognised route. The institutional consequence follows from whether equity should be built into eligibility, outreach and adaptation. An apparently equal offer may be unusable because of cost, time, language, disability, safety or location. The plan should test which learners can take up the measure and provide the additional condition required for substantive access. Local adaptation is desirable where it removes a barrier while protecting the common educational objective.[REF-08] [REF-13] [REF-14]

    improvement proposition and implementation design cannot be judged without identifying the plan should identify indicators or qualitative evidence that would reveal these effects and the authority able to intervene. A contrary reading would overlook that improvement is not established by movement of the selected measure if another essential condition deteriorates. Foreseeable adverse effects should be recorded. Additional assessment can narrow curriculum or increase exclusion; targeted grouping can stigmatise learners; extended time can increase household costs; intensive support for one cohort can divert staff from another.[REF-01] [REF-11] [REF-23]

    78

    Resources and professional capability

    In assessing resources and professional capability, authorities must determine an unfunded dependency should remain visible in the risk statement. The public account remains incomplete unless it explains how the resource plan should cover the complete recurrent service: personnel, preparation, instructional time, accessible materials, facilities, learner support, assessment, coordination, evidence and review. Capital expenditure or a temporary grant may be necessary but cannot establish sustained capability by itself. The institution should identify which costs fall within its budget, which require higher-level allocation and which are presently unfunded.[REF-15] [REF-26] [REF-28]

    In assessing resources and professional capability, authorities must determine programmes serving learners with greater need may require more staff time, specialist support, transport or accessible materials. The evidence must therefore clarify how equal departmental or per-learner amounts can reproduce unequal opportunity. Allocation reasons should be documented and reviewed against actual receipt and use. Additional finance should be justified by the common educational condition it protects, not by a permanent assumption of lower capability among the affected learners. Distributional finance matters within institutions as well as across systems.[REF-06] [REF-10] [REF-13]

    In assessing resources and professional capability, authorities must determine one briefing does not establish readiness for sustained change. For the learners concerned, the decisive consideration is whether the plan should monitor participation, usable learning, workload, turnover and access to later assistance. Professional capability is a central implementation condition. Teachers, trainers and institutional leaders need subject knowledge, diagnostic skill, practical examples and time to collaborate. Where the plan alters curriculum, assessment or inclusion arrangements, staff should have supported opportunities to understand the educational reasoning.[REF-01] [REF-18] [REF-25]

    For resources and professional capability, the material distinction is between conversely, higher-level constraints should not prevent local correction that is lawful and feasible. This matters because the plan should identify decision thresholds, escalation routes and the time within which the responsible body will respond. Leadership should create a route from classroom or service evidence to institutional and system correction. Staff should not be held responsible for changing transport, staffing establishment, qualification rules or infrastructure beyond their authority.[REF-12] [REF-17] [REF-19]

    In assessing resources and professional capability, authorities must determine universities, civil society, employers or international bodies may contribute knowledge, facilities or finance, but programme status, data protection and continuity should be clear. A contrary reading would overlook that parallel arrangements can create fragmented standards or reporting burdens. Agreements should state the educational purpose, authority, resource duration, records, safeguards and handover. The measure of partnership is whether the institution can sustain equitable quality and correction after exceptional support ends. External partnership should strengthen ordinary capability.[REF-22] [REF-26] [REF-28]

    79

    Evidence, adaptation and institutionalisation

    In assessing evidence, adaptation and institutionalisation, authorities must determine the review schedule should collect evidence proportionate to each claim and identify the population represented. The resulting interpretation should show why a favourable result should not be broadened beyond the learners, content and period observed. Review should distinguish authorization, delivery, reach, educational quality and consequence. Approval of the plan establishes authority; expenditure demonstrates a resource transaction; implementation records show activity; learner evidence shows participation or change. None automatically proves the next.[REF-03] [REF-16] [REF-24]

    Comparative interpretation of evidence, adaptation and institutionalisation depends upon non-participation can reveal communication, cost, timing, safety or accessibility barriers. This matters because the review should disaggregate where lawful and precise and should retain absolute numbers. An improved average does not establish that learners farthest from the baseline benefited. The plan should define a minimum floor or least-served test alongside its aggregate objective. Reach should be examined from eligibility through offer, take-up, participation and intended intensity.[REF-01] [REF-06] [REF-13]

    For evidence, adaptation and institutionalisation, the material distinction is between agreement strengthens the conclusion only after dates and definitions align. A proportionate conclusion must also recognise that contradiction should initiate enquiry into coverage, implementation or measurement rather than selection of the most favourable source. Educational consequence may require several forms of evidence. Assessment can show a bounded learning domain; attendance and completion records show participation; observation and work samples can explain teaching and learner activity; interviews can identify experience and mechanism.[REF-02] [REF-12] [REF-23]

    Comparative interpretation of evidence, adaptation and institutionalisation depends upon ending an ineffective measure is responsible improvement when learner safeguards and alternative provision are addressed. The resulting interpretation should show why adaptation should follow a stated decision rule. Weak outcomes may reflect an incorrect hypothesis, incomplete delivery, insufficient intensity, adverse conditions, measurement weakness or inadequate time. These explanations call for different action. The review should record which is supported, what changes and how the baseline or comparison remains interpretable.[REF-08] [REF-11] [REF-28]

    evidence, adaptation and institutionalisation cannot be judged without identifying completion means that the institution can sustain the improved learner-facing condition and correct foreseeable departure, not that the plan period has ended. The institutional consequence follows from whether institutionalisation requires ordinary authority, recurrent finance, professional ownership and a functioning correction route. Continued activity after a pilot is not sufficient if it depends on exceptional staff or external funds. The final decision should state which function enters ordinary provision, which limitation remains, who owns it and when evidence will next be examined.[REF-17] [REF-25] [REF-26]

    Part XXII

    Accountability design for educational improvement

    80

    A balanced evidence set

    An accountability arrangement should begin with the educational claim and assemble the smallest balanced evidence set capable of testing it. Participation, progression, completion, learning, teacher and facility conditions, resources and learner experience are related but not interchangeable. A single headline may support communication, but the underlying components should remain available so that strong performance in one dimension cannot conceal exclusion or failure in another.[REF-01] [REF-03] [REF-12]

    The population represented by each item should be explicit. Administrative records describe registered learners and institutions; assessments describe eligible participants under stated conditions; household sources may represent people outside provision; observation and interviews explain mechanisms without necessarily establishing prevalence. Apparent agreement adds confidence only after dates and definitions align. Contradiction should guide enquiry rather than selection of the preferred value.[REF-02] [REF-23] [REF-24]

    The balanced set should include distribution. National or institutional averages can improve while a smaller group deteriorates or remains below an educational minimum. Sex, household resources, disability, language, residence, migration and other relevant classifications should be selected for a public decision and protected against unsafe disclosure. Group identity should not be treated as the cause of an observed difference.[REF-05] [REF-10] [REF-13]

    Context should inform interpretation without excusing a failed entitlement. Prior opportunity, population movement, institutional mandate and resource conditions may explain why results differ and which response is feasible. They should not be used to normalize low expectations for disadvantaged learners. Reports should distinguish direct result, contextual condition, plausible mechanism and accountable action.[REF-06] [REF-14]

    The evidence set should be reviewed for burden and behavioural effect. Collection that repeatedly consumes teaching time or duplicates records can weaken the service it seeks to improve. A measure that dominates consequences can narrow attention. Authorities should retain measures because they support a real decision and remove or redesign those whose cost or distortion exceeds their public value.[REF-08] [REF-17]

    81

    Proportionate consequences and fair process

    Consequences should follow the nature, severity and certainty of the finding. A verified safety, discrimination or legality failure can require immediate protection while fuller explanation proceeds. A recurring learning weakness may call for support, diagnostic review and a time-bound plan. An unstable or low-precision difference may warrant stronger evidence. Treating every adverse value alike weakens fairness and directs effort away from the mechanism.[REF-05] [REF-12]

    Institutions should receive the population, definition, source, calculation, uncertainty and proposed interpretation before a consequential decision. They should be able to correct factual error and supply relevant missing evidence. This is not a right to suppress or rewrite a valid finding. The deciding body should record the response, its evaluation and reasons.[REF-08] [REF-09]

    Authority and control matter. A school cannot recruit teachers beyond an authorized establishment, alter national curriculum or construct facilities without finance. Accountability should assign each dependency to the body capable of remedy rather than convert a system constraint into a local sanction. Local practice remains reviewable, but the responsibility chain should reflect actual powers.[REF-17] [REF-19]

    Consequences can produce exclusion incentives. If results depend heavily on tested populations or completion, institutions may discourage admission, reclassify learners or narrow participation. Monitoring should include entry, movement, exclusions, assessment participation and unresolved destinations. The existence of an incentive does not prove misconduct; it requires design safeguards and independent review.[REF-01] [REF-03]

    Appeal and correction should be timely and accessible. A consequential error can affect reputation, resources and learners. The record should state the decision, evidence, remedy, further appeal and whether earlier public information will be corrected. Where immediate protective action was necessary, the affected institution or individual should still receive fair subsequent review.[REF-10] [REF-16]

    82

    Support attached to accountability

    An adverse finding should lead to a support and improvement decision, not merely classification. The response should name the educational condition, responsible authority, action, resource, delivery date and evidence of operation. Additional plans, meetings or guidance are not themselves improvement. The review should verify that staff and learners received the intended condition and that it addressed the diagnosed barrier.[REF-07] [REF-12]

    Teacher-facing recommendations should include subject or pedagogical substance, examples, planning time and follow-up. General training can be too remote from the observed problem. Professional support should respect workload and use moderated evidence rather than impose one script for unlike classrooms. Serious professional failures require fair procedures, but average results alone do not establish individual misconduct.[REF-15] [REF-18]

    Resource recommendations should identify recurrent as well as initial cost. Materials require accessible use and replacement; facilities require maintenance; learner support requires eligibility and continuity; additional staff require authorized posts. An institution should not be judged for failing to deliver a measure whose necessary resource was never transferred. The funding body should remain accountable for timing and adequacy.[REF-17] [REF-19]

    Support should be distributed according to need. Equal offers can reinforce inequality where institutions begin with different staffing, infrastructure or learner barriers. Additional support should protect a common substantive standard, not mark a community as inherently weak. Allocation and receipt should be reported, and later review should show whether the least-served population benefited.[REF-06] [REF-13]

    External assistance should build ordinary institutional capability. A temporary expert or grant can help diagnose and test a response, but the plan should state how knowledge, finance and ownership transfer. A successful intervention dependent on exceptional support is not yet institutionalized.[REF-26] [REF-28]

    83

    Public reporting without distortion

    Public accountability information should state the exact claim, population, period and uncertainty. It should avoid comprehensive labels such as “good” or “failing” when evidence concerns one outcome. Rankings should not be produced where populations and opportunity are unlike or differences are statistically or educationally immaterial. Components and context are more useful for improvement than a precise ordinal position.[REF-03] [REF-08]

    Composite measures require particular caution. Weighting permits one dimension to compensate for another and may hide essential failure. If used, the components, weights, missing-value treatment and sensitivity should be public. Safety, access and minimum quality should remain separately visible where failure cannot be ethically offset by performance elsewhere.[REF-05] [REF-12]

    Privacy and small numbers constrain detail. Suppression, aggregation, controlled access and narrative accounts can protect learners, but the affected population and institutional action should not disappear. Repeated releases and combinations with other information should be considered when assessing disclosure risk.[REF-10] [REF-13]

    Communication should distinguish observation, explanation and attribution. A disparity can justify action without proving one cause. Improvement after an intervention can support a contribution account without establishing sole effect. The strongest defensible wording should be used; overstating evidence can later undermine confidence in valid findings.[REF-01] [REF-16]

    Corrections should reach the audience of the original release. Replacing a file silently is insufficient. The notice should state what changed, why, which conclusions are affected and whether a decision should be reconsidered. Preserving revision history demonstrates professional integrity and supports later learning.[REF-08] [REF-09]

    84

    Candour and institutional learning

    Improvement depends upon the ability to reveal weakness without assuming that every adverse finding proves negligence. Institutions should be expected to maintain accurate evidence, report serious harm and respond honestly. Authorities should distinguish candour from performance and should not create incentives to delay recording difficult cases or exclude populations from the denominator.[REF-08] [REF-17]

    Self-evaluation provides local knowledge and rapid feedback but should not be treated as independent assurance. External review provides challenge and comparability but may miss mechanism or context. A constructive arrangement uses each for its strengths and records disagreement. Institutions should retain evidence that contradicts the preferred account.[REF-02] [REF-23]

    Learning requires a trace from finding to response. The record should show the original criterion, evidence, interpretation, selected action, resource, implementation, reach, outcome and review. When a decision changes, the reason should remain accessible. This prevents repeated initiatives from being represented as cumulative improvement without verified continuity.[REF-03] [REF-24]

    Leaders should protect professional discussion. Teachers and learners need safe channels to explain barriers and adverse effects. Confidentiality and safeguarding remain applicable. Participation does not make every account equally probative, but exclusion of those affected can produce unworkable or unsafe recommendations.[REF-10] [REF-15]

    External scrutiny should test the reasoning chain, distribution, proportionality and correction. It should not reward polished documentation or the absence of reported problems. A credible institution can state what it does not know and what lies outside its authority, provided it pursues the appropriate evidence and escalation.[REF-12] [REF-19]

    85

    Verification of improvement and exit

    Follow-up should return to the original population and criterion. It should distinguish authorization, delivery, use, reach and educational consequence. A response can be fully implemented yet ineffective; a promising result can arise before full delivery through other change. Each state supports a different judgement.[REF-01] [REF-03]

    Comparability should be protected. If assessment, population, programme or collection changed, a before-and-after difference may not represent educational change. Overlap evidence and sensitivity can sometimes support interpretation. Otherwise the institution should report current status and break the trend rather than manufacture continuity.[REF-16] [REF-23]

    Equity verification should identify who received the measure, its intensity and the later condition. Improvement in an average does not establish benefit for the intended or least-served group. The review should examine non-participation and adverse effects and should not count people who left the institution as successful without a known destination.[REF-06] [REF-13]

    Scaling requires evidence on capability, cost, context and durability. A favourable pilot does not show that ordinary institutions can reproduce the result. Expansion should preserve core educational mechanisms while allowing adaptation and should monitor whether weaker institutions receive necessary support.[REF-17] [REF-28]

    Exit from a plan should be a reasoned decision. The measure may become ordinary provision, adapt, continue under defined conditions or cease. Completion is demonstrated when the improved condition is sustained and a correction route functions, not when a project period or external review ends.[REF-12] [REF-19]

    Part XXIII

    Interpreting purpose, proportionality and attribution

    86

    The public purpose of accountability

    Accountability is a relationship in which an actor has a defined responsibility, must provide an account to a forum entitled to examine it, and may face a response grounded in that examination. In education, the relationship is never adequately described by the publication of results alone. It requires a prior duty, a competent forum, an intelligible account and an available remedy. A result without an identified duty may inform public debate, but it does not by itself establish who has failed. A duty without an accessible account leaves learners and the public unable to judge whether the obligation has been discharged.[REF-08] [REF-27] [REF-29]

    The legitimate purposes are plural. Accountability can protect rights, ensure lawful use of public resources, expose unequal treatment, maintain professional standards, support democratic deliberation and stimulate correction. These purposes overlap but do not permit instruments to be exchanged without examination. Financial regularity does not demonstrate educational effectiveness. A test score does not establish compliance with every public duty. Satisfaction evidence does not displace safeguarding or curriculum requirements. The authority should therefore state which purpose an instrument serves and what decision it can support.[REF-07] [REF-12] [REF-29]

    Learners and families are rights-holders and participants, not merely sources of performance data. Their capacity to demand an account varies with information, language, legal status, disability, poverty and social power. An arrangement that relies exclusively on individual choice or complaint can leave the least powerful with the weakest remedy. Public authorities retain responsibility for ensuring that information is accessible, that complaints can be made safely, and that structural failures are examined even when no individual complaint has been lodged.[REF-08] [REF-10] [REF-29]

    Professional responsibility is also indispensable. Teachers and institutional leaders exercise judgement that cannot be reduced to compliance with a detailed instruction. They should explain the basis of consequential decisions, use relevant evidence, protect learners and participate in peer scrutiny. Professional autonomy is not immunity from account; it is authority bounded by knowledge, ethics and public obligation. Conversely, an accountability regime that treats professional judgement as mere obedience can suppress adaptation to learner need and weaken ownership of correction.[REF-15] [REF-18] [REF-29]

    Democratic accountability requires that elected and administrative authorities account for the conditions they control. National standards, finance, teacher establishments, admissions rules, assessment design and data systems shape what institutions can deliver. Concentrating scrutiny on schools while leaving these system decisions outside view misstates responsibility. The public account should follow the allocation of legal authority and resources from central government through intermediate bodies to institutions, contractors and other authorised providers.[REF-17] [REF-19] [REF-29]

    87

    Mapping responsibility before judgement

    Responsibility should be mapped before evidence is converted into consequence. The mapping should identify the substantive duty, the actor bearing it, the powers and resources available, the period covered, the body entitled to review performance and the remedy within that body's competence. Several actors may hold connected duties. The existence of shared responsibility does not justify vague collective blame; it requires the contribution of each actor to be specified.[REF-27] [REF-28] [REF-29]

    Responsibility should be distinguished from influence. A teacher influences learning but does not control prior opportunity, household conditions, staffing policy, infrastructure or the design of national examinations. A ministry influences local delivery but cannot attribute every classroom event directly to a single policy. Accountability should concern the reasonable discharge of an assigned duty under actual authority, including whether a responsible actor identified constraints, used available powers and escalated dependencies.[REF-01] [REF-17] [REF-29]

    Delegation does not extinguish public duty. Where private, community, international or civil-society bodies deliver education with public authority or finance, the competent public body should define standards, assure access, monitor performance and provide remedy. The provider must account for its own obligations, while the state must account for regulation and oversight. A contract can allocate tasks; it cannot by itself remove the public responsibility to secure the right to education.[REF-08] [REF-26] [REF-29]

    Parents, learners and communities have responsibilities that should not be enlarged into explanations for public failure. Attendance, participation and care matter, but poverty, displacement, unsafe travel, inaccessible facilities or unaffordable costs can constrain the exercise of individual responsibility. An account should examine whether the public arrangement made participation genuinely possible before treating non-participation as voluntary. Where families contribute information or local oversight, they require accessible procedures and protection against retaliation.[REF-06] [REF-10] [REF-29]

    International actors should account for commitments, finance, advice and the effects of their conditions, while national authorities remain responsible for national policy. Aid volatility, fragmented projects and parallel reporting can weaken institutional capacity. The relevant account should show alignment with national priorities, predictability, additional burden, distribution and whether external support strengthens ordinary public systems. A favourable project result does not settle whether the wider financing relationship was equitable or sustainable.[REF-26] [REF-28] [REF-29]

    88

    Proportionality of scrutiny and response

    Proportionality connects the importance of the duty, seriousness of possible harm, strength of evidence, degree of control and burden of the response. It does not mean that serious failure should receive a mild response. It means that scrutiny and consequence must be suitable for the public purpose and no more restrictive or punitive than justified by the evidence and available remedy. Urgent protection can be proportionate before causal certainty is complete when credible evidence indicates immediate risk.[REF-08] [REF-10] [REF-29]

    Low-stakes diagnostic evidence may appropriately trigger enquiry or support. High-stakes consequences require stronger safeguards: stable definitions, sufficient coverage, attention to uncertainty, verification of material facts and an opportunity to respond. A measure developed for system monitoring should not be attached to an individual sanction merely because it is available. The new use changes the incentive, required precision and fairness assessment.[REF-02] [REF-03] [REF-29]

    Frequency should also be proportionate. Continuous reporting can appear rigorous while consuming instructional and administrative capacity. The interval should reflect how quickly the condition can change, when a responsible authority can act, and when evidence can show a meaningful difference. Safeguarding events may require immediate reporting; curriculum implementation or learning progression may require a longer period. Repeated measurement before action is possible can create noise and encourage superficial response.[REF-12] [REF-17]

    Consequences should be reversible where uncertainty is material. Temporary support, supervised correction or a bounded improvement requirement may protect learners while evidence develops. Irreversible closure, exclusion, dismissal or reputational labelling demands a stronger basis and fair process. Where an urgent interim measure is necessary, the authority should state its duration, review test and route for challenge.[REF-08] [REF-10] [REF-29]

    Proportionality applies to disclosure. Public access is important, but release of small-group or sensitive information can expose learners, stigmatise communities or undermine trust. The reporting body should publish enough to support scrutiny while using aggregation, suppression or controlled access where necessary. The restriction should protect legitimate interests without concealing the existence of a disparity or the authority's response.[REF-10] [REF-13]

    89

    The limits of attribution

    Attribution is a causal judgement, not a synonym for observed association. A difference between institutions, groups or periods can arise from prior opportunity, population composition, measurement, implementation, contextual change or chance as well as from the action under examination. Accountability may still require a response to an observed condition, but the response should not allege a cause that the evidence cannot sustain.[REF-01] [REF-02] [REF-29]

    Education outcomes are jointly produced over time. Learning reflects cumulative opportunity across homes, communities, institutions and public systems. Completion reflects admission, progression, affordability, support and recognition. Employment or civic outcomes lie still farther from the control of a single institution. The further the outcome from the actor's decision, the more explicit the contribution pathway and alternative explanations must be.[REF-03] [REF-12] [REF-29]

    Value-added or contextual adjustment can improve a comparison but does not create complete control for unobserved difference. Model choices, missing data, measurement error and mobility affect estimates. Adjusted results should therefore be presented with uncertainty and sensitivity, and should not be described as the pure effect of an institution or teacher. Statistical sophistication cannot replace a credible account of mechanism, implementation and population.[REF-02] [REF-23]

    An accountability body should distinguish four statements: the condition occurred; the actor held a relevant duty; an act or omission contributed to the condition; and a particular remedy is justified. Evidence sufficient for the first may be insufficient for the third. A duty can require investigation or correction even when sole causation is not established. This distinction preserves action without sacrificing fairness.[REF-08] [REF-29]

    Improvement following an intervention should likewise be interpreted cautiously. The change may be consistent with the intended mechanism, but contemporaneous policy, staffing, population or assessment changes may also contribute. Stronger attribution can be supported by a clear time sequence, verified implementation, comparison, dose or exposure evidence, mechanism evidence and the absence of plausible rival explanations. Even then, the claim should be confined to the observed population and period.[REF-16] [REF-24]

    90

    Safeguards, remedy and institutional learning

    Fair process strengthens rather than obstructs accountability. The affected actor should know the duty, evidence, interpretation and proposed consequence; be able to correct factual error and submit relevant evidence; receive a reasoned decision; and have access to review. These protections do not permit delay in preventing serious harm or allow an institution to veto a valid finding. They make the account more accurate, legitimate and usable.[REF-08] [REF-10]

    Remedy should correspond to the violated interest and competent authority. Learners may need immediate protection, restoration of access, additional teaching, accessible support, correction of a record or financial redress under applicable law. Institutions may need staffing, finance, technical assistance or a change in system rule. Public explanation alone is rarely a complete remedy where an educational entitlement remains unmet.[REF-17] [REF-19] [REF-29]

    Collective and individual accountability should be connected carefully. A system weakness may require national correction even when individual misconduct is absent. Individual misconduct may require a fair professional or legal procedure without branding an entire institution. Aggregate outcomes should not be used as automatic proof against an individual; nor should division of labour allow serious collective failure to escape review.[REF-15] [REF-29]

    Accountability should preserve the capacity to learn from failure. If every adverse finding automatically produces severe punishment, institutions may conceal uncertainty, avoid difficult learners or limit experimentation. The regime should distinguish candid identification followed by competent correction from concealment, repeated neglect or bad faith. The distinction depends upon verified action and learner protection, not upon accepting self-description.[REF-12] [REF-29]

    The final review should ask whether the duty was clarified, the account reached an authorised forum, evidence was fairly interpreted, consequence was proportionate, remedy reached affected learners and recurrence is less likely. Closure should not be declared merely because a response document was accepted. A completed accountability relationship leaves a verifiable change, a reasoned decision or an explicit unresolved duty assigned to the body capable of acting.[REF-03] [REF-28]

    References

    1. REF-01

      Education for All Global Monitoring Report Team. Reaching the Marginalized — EFA Global Monitoring Report 2010. 2010.

      Principal contemporaneous analysis of intersecting disadvantage and education marginalisation.

      https://unesdoc.unesco.org/ark:/48223/pf0000186606
    2. REF-02

      UNESCO Institute for Statistics. Global Education Digest 2010: Comparing Education Statistics Across the World. 2010.

      Comparative education statistics, definitions and limitations.

      https://uis.unesco.org/sites/default/files/documents/global-education-digest-2010-comparing-education-statistics-across-the-world-en.pdf
    3. REF-03

      UNESCO Institute for Statistics. Education Indicators: Technical Guidelines. 2009.

      Definitions and interpretation of participation, progression, completion and resource indicators.

      https://uis.unesco.org/sites/default/files/documents/education-indicators-technical-guidelines-en_0.pdf
    4. REF-04

      United Nations. The Millennium Development Goals Report 2010. 2010.

      Global and regional monitoring of primary education, gender, poverty and related development conditions.

      https://www.un.org/millenniumgoals/pdf/MDG%20Report%202010%20En%20r15%20-low%20res%2020100615%20-.pdf
    5. REF-05

      United Nations Development Programme. Human Development Report 2010: The Real Wealth of Nations — Pathways to Human Development. 2010.

      Distribution-sensitive human development concepts and evidence available before the cut-off.

      https://hdr.undp.org/content/human-development-report-2010
    6. REF-06

      UNICEF. Progress for Children: Achieving the MDGs with Equity, Number 9. 2010.

      Equity-focused child indicators and comparison between population groups.

      https://www.unicef.org/reports/progress-children-no-9
    7. REF-07

      World Education Forum. The Dakar Framework for Action: Education for All — Meeting Our Collective Commitments. 2000.

      Commitments to equitable access, quality, measurable outcomes and accountable national planning.

      https://unesdoc.unesco.org/ark:/48223/pf0000121147
    8. REF-08

      United Nations General Assembly. Convention on the Rights of the Child. 1989.

      Rights concerning non-discrimination, identity, education and development.

      https://www.ohchr.org/en/instruments-mechanisms/instruments/convention-rights-child
    9. REF-09

      United Nations Committee on Economic, Social and Cultural Rights. General Comment No. 13: The Right to Education. 1999.

      Interpretation of availability, accessibility, acceptability and adaptability.

      https://www.refworld.org/legal/general/cescr/1999/en/37937
    10. REF-10

      United Nations General Assembly. Convention on the Rights of Persons with Disabilities. 2006.

      Non-discrimination, accessibility, inclusive education and disability data safeguards.

      https://www.ohchr.org/en/instruments-mechanisms/instruments/convention-rights-persons-disabilities
    11. REF-11

      United Nations. Guiding Principles on Internal Displacement. 1998.

      Principles relevant to protection, documentation, education and non-discrimination of displaced persons.

      https://www.ohchr.org/en/special-procedures/sr-internally-displaced-persons/international-standards
    12. REF-12

      UNESCO and UNICEF. A Human Rights-Based Approach to Education for All. 2007.

      Rights-based planning, equality, participation, accountability and education quality.

      https://unesdoc.unesco.org/ark:/48223/pf0000154861
    13. REF-13

      Education for All Global Monitoring Report Team. Overcoming Inequality: Why Governance Matters — EFA Global Monitoring Report 2009. 2008.

      Governance, finance and unequal educational opportunity.

      https://unesdoc.unesco.org/ark:/48223/pf0000177683
    14. REF-14

      World Bank. World Development Report 2006: Equity and Development. 2005.

      Concepts of unequal opportunity, institutions and equitable public action.

      https://documents.worldbank.org/curated/en/435331468127174418/pdf/322040World0Development0Report02006.pdf
    15. REF-15

      World Bank. Safeguarding Education During Economic Crisis. 2009.

      Risks to budgets, households, participation and long-term human development during economic crisis.

      https://documents1.worldbank.org/curated/en/489131468340200911/pdf/485120WP0Avert10Box338912B01PUBLIC1.pdf
    16. REF-16

      Organisation for Economic Co-operation and Development. Education at a Glance 2010: OECD Indicators. 2010.

      Comparative participation, progression, expenditure and outcomes evidence with system-level metadata.

      https://doi.org/10.1787/eag-2010-en
    17. REF-17

      European Commission. Europe 2020: A Strategy for Smart, Sustainable and Inclusive Growth. 2010.

      Contemporaneous European policy context for education, inclusion, employment and headline indicators.

      https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:52010DC2020
    18. REF-18

      European Commission. Youth on the Move: An Initiative to Unleash the Potential of Young People to Achieve Smart, Sustainable and Inclusive Growth in the European Union. 2010.

      European education, mobility, attainment and youth inclusion policy context.

      https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:52010DC0477
    19. REF-19

      European Commission. A Renewed Commitment to Social Europe: Reinforcing the Open Method of Coordination for Social Protection and Social Inclusion. 2008.

      Social inclusion monitoring, common objectives and context-sensitive indicators.

      https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:52008DC0418
    20. REF-20

      United Nations General Assembly. Resolution 64/250: Assistance to Haiti in the Aftermath of the Recent Earthquake. 2010.

      Contemporaneous recognition of humanitarian and reconstruction needs and national leadership.

      https://undocs.org/A/RES/64/250
    21. REF-21

      United Nations Office for the Coordination of Humanitarian Affairs. Haiti Revised Humanitarian Appeal. 2010.

      Displacement, service disruption and humanitarian education context.

      https://reliefweb.int/report/haiti/haiti-revised-humanitarian-appeal-2010
    22. REF-22

      European Commission. European Union Response to the Earthquake in Haiti. 2010.

      European humanitarian and recovery support, coordination and Haitian ownership.

      https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:52010DC0056
    23. REF-23

      United Nations Economic and Social Council. Principles and Recommendations for Population and Housing Censuses, Revision 2. 2008.

      Official principles for population coverage, definitions, classifications and data quality.

      https://unstats.un.org/unsd/demographic-social/Standards-and-Methods/files/Principles_and_Recommendations/Population-and-Housing-Censuses/Series_M67Rev2-E.pdf
    24. REF-24

      United Nations Statistics Division. Designing Household Survey Samples: Practical Guidelines. 2005.

      Sample design, estimation, precision and non-response guidance.

      https://unstats.un.org/unsd/demographic/sources/surveys/Handbook23June05.pdf
    25. REF-25

      World Education Forum 2015. Incheon Declaration: Education 2030 — Towards Inclusive and Equitable Quality Education and Lifelong Learning for All. 2015.

      Education commitments adopted at Incheon and their contemporaneous institutional status.

      https://unesdoc.unesco.org/ark:/48223/pf0000233137
    26. REF-26

      United Nations General Assembly. Addis Ababa Action Agenda of the Third International Conference on Financing for Development. 2015.

      Adopted global financing framework and relevant principles for domestic public finance and international cooperation.

      https://undocs.org/A/RES/69/313
    27. REF-27

      United Nations General Assembly. Transforming Our World: The 2030 Agenda for Sustainable Development. 2015.

      Agenda adopted on 25 September 2015, including Goal 4 and its education targets, as available at the evidence cut-off.

      https://undocs.org/A/RES/70/1
    28. REF-28

      UNESCO. Education 2030: Incheon Declaration and Framework for Action for the Implementation of Sustainable Development Goal 4. 2015.

      Framework for Action available before the cut-off and relevant to institutional planning, coordination, equity and review.

      https://unesdoc.unesco.org/ark:/48223/pf0000245656
    29. REF-29

      Global Education Monitoring Report Team. Accountability in Education: Meeting Our Commitments — Global Education Monitoring Report 2017/8. 2017.

      Principal contemporaneous global analysis of the purposes, actors, instruments and risks of accountability in education.

      https://unesdoc.unesco.org/ark:/48223/pf0000259338