ICEQC-R-2017-04 — School-Level Performance Measures and the Risk of Misclassification cover

Informe de Investigación Temática

ICEQC-R-2017-04 — School-Level Performance Measures and the Risk of Misclassification

A global comparative indicator study of measure validity, uncertainty, incentives and proportionate accountability

Fecha de publicación
Categoría de investigación
Investigación de datos e indicadores
Informe arquetipo
Estudio comparativo de indicadores
Ámbito geográfico
Global
Fecha límite para la presentación de pruebas
Organismo responsable
Dirección de Investigación y Políticas de ICEQC
ICEQC-R-2017-04 — School-Level Performance Measures and the Risk of Misclassification cover

Publication record

This is the controlled English edition. Evidence and institutional status are stated as at the evidence cut-off date.

Executive summary

School-level performance measures can support diagnosis, resource allocation and improvement, but a classification is not a neutral description. It places a school above or below a threshold, into a category or in a rank, often with consequences for staff, learners and community confidence. Misclassification occurs when the observed evidence, threshold or model places a school in a category that does not represent the underlying condition relevant to the decision.

This report begins with purpose. Achievement, progress, attendance, completion, school climate and service conditions answer different questions. A measure fit for descriptive reporting may be unfit for sanction. A test score may describe assessed learning but cannot represent safeguarding, curriculum breadth or inclusion. A composite can conceal severe weakness through compensation and can change ranks when normalisation or weights change. Every measure therefore requires a conceptual model, defined school and learner population, reference period, source coverage and permissible inference.

Fair comparison requires context but not lower expectations. Prior attainment, poverty, language, migration, disability, school size and resources affect observed outcomes and estimate stability. Context should help diagnose barriers and assign system responsibility; it should not predict away the learning of disadvantaged learners or excuse denial of a minimum standard. Reports should retain absolute outcome levels, progress evidence and opportunity-to-learn conditions.

Statistical uncertainty is central. Sampling, non-response, measurement reliability, threshold choice, model specification and year-to-year volatility can move schools across categories. Schools near a boundary need a review band and corroborating evidence. Small schools require counts, intervals and cautious use of pooled years. Missing assessment or survey data should be profiled, because strategic absence or exclusion can improve apparent performance.

High-stakes measures alter behaviour. Curriculum may narrow; admission, transfer or disciplinary decisions may protect results; records may substitute compliance for service; and low labels may accelerate staff loss and family avoidance. Accountability should distinguish conditions controlled by the school from finance, staffing or access duties controlled by higher authorities. Classification should lead to a dated diagnosis, support and later review, not become a permanent identity.

The recommended public account presents component levels and distributions, coverage, uncertainty, context and limitations before any category. Decision rules are published in advance. High-stakes classifications require independent technical and contextual review, time-bound appeal and continuing learner protection. Errors and revisions are visible. The account uses only official evidence available by 1 April 2017 and does not apply later accountability frameworks retrospectively.

Key findings

  • A school measure must be designed for a stated decision and cannot automatically support every accountability use.
  • Achievement, progress, participation, completion, climate and service conditions should remain distinct components.
  • Context informs diagnosis and allocation but should not lower the common educational entitlement.
  • School size, missingness, reliability, thresholds, model choice and annual volatility create material classification error.
  • Schools close to category boundaries require uncertainty bands and corroborating evidence.
  • High stakes can narrow curriculum, encourage selection or exclusion and substitute record compliance for education.
  • Accountability should separate school-controlled practice from system-controlled staffing, finance and access conditions.
  • Classification should trigger proportionate support, review and remedy and should be retired when the diagnosed condition changes.

Scope and method

This global comparative indicator study addresses ministries, inspectorates, statistical offices, school authorities and communities using school-level measures. It covers public reporting, diagnosis, support, resource decisions and higher-stakes categories. It does not supply a universal school score or endorse league tables.

Evidence is restricted to official United Nations, UNESCO and European institutional material available by 1 April 2017. It combines statistical principles, education indicator guidance, school-evaluation evidence, rights instruments and contemporaneous monitoring. Later accountability reports, models and results are excluded.

Part I

Purpose, unit and permissible claim

1

Defining the school as the reporting unit

Defining the school as the reporting unit defines the classification question for the school unit. For defining the school as the reporting unit, the relevant population or units are institutions, campuses, programmes and shifts grouped as one school, and the direct evidence concerns unit boundary. In examining defining the school as the reporting unit, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on defining the school as the reporting unit, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of defining the school as the reporting unit, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-01] [REF-05] [REF-06] [REF-16]

The principal misclassification risk in defining the school as the reporting unit is that multi-site, selective or shared programmes are combined differently across schools. In examining defining the school as the reporting unit, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on defining the school as the reporting unit, authorities should identify how learners, events and institutions enter or leave the school unit, whether missingness is concentrated, and which assumptions drive the result. For classification of defining the school as the reporting unit, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-01] [REF-05]

The minimum measurement response for defining the school as the reporting unit is to publish the legal and operational unit and every inclusion or split rule. Within evidence on defining the school as the reporting unit, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of defining the school as the reporting unit, it should also state the decision rule, consequence and evidence required for review. In interpreting defining the school as the reporting unit, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-05] [REF-06]

Source fitness for the school unit depends on coverage. For classification of defining the school as the reporting unit, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting defining the school as the reporting unit, for defining the school as the reporting unit, each source should retain its principal error. For decisions about defining the school as the reporting unit, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-06] [REF-16]

Equity is part of defining the school as the reporting unit. In interpreting defining the school as the reporting unit, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about defining the school as the reporting unit, group levels and population shares should remain visible beside any school average. For defining the school as the reporting unit, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining defining the school as the reporting unit, if part of institutions, campuses, programmes and shifts grouped as one school is absent, the likely effect on unit boundary and classification should be reported.[REF-01] [REF-16]

Interpretation of unit boundary should separate observation, explanation and attribution. For decisions about defining the school as the reporting unit, for defining the school as the reporting unit, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For defining the school as the reporting unit, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining defining the school as the reporting unit, consequences should be proportionate to evidence strength. Within evidence on defining the school as the reporting unit, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-05] [REF-06] [REF-16]

Accountability completes defining the school as the reporting unit. For defining the school as the reporting unit, a material finding about the school unit should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining defining the school as the reporting unit, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on defining the school as the reporting unit, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-01] [REF-05] [REF-06] [REF-16]

2

Separating description, diagnosis and judgement

Separating description, diagnosis and judgement defines the classification question for the measure purpose. For separating description, diagnosis and judgement, the relevant population or units are schools and users receiving different forms of evidence, and the direct evidence concerns descriptive, diagnostic and evaluative use. In examining separating description, diagnosis and judgement, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on separating description, diagnosis and judgement, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of separating description, diagnosis and judgement, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-01] [REF-09] [REF-10] [REF-16]

The principal misclassification risk in separating description, diagnosis and judgement is that a descriptive result becomes an unsupported judgement of institutional quality. In examining separating description, diagnosis and judgement, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on separating description, diagnosis and judgement, authorities should identify how learners, events and institutions enter or leave the measure purpose, whether missingness is concentrated, and which assumptions drive the result. For classification of separating description, diagnosis and judgement, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-01] [REF-09]

The minimum measurement response for separating description, diagnosis and judgement is to state the decision, inference and evidence threshold before calculation. Within evidence on separating description, diagnosis and judgement, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of separating description, diagnosis and judgement, it should also state the decision rule, consequence and evidence required for review. In interpreting separating description, diagnosis and judgement, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-09] [REF-10]

Source fitness for the measure purpose depends on coverage. For classification of separating description, diagnosis and judgement, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting separating description, diagnosis and judgement, for separating description, diagnosis and judgement, each source should retain its principal error. For decisions about separating description, diagnosis and judgement, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-10] [REF-16]

Equity is part of separating description, diagnosis and judgement. In interpreting separating description, diagnosis and judgement, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about separating description, diagnosis and judgement, group levels and population shares should remain visible beside any school average. For separating description, diagnosis and judgement, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining separating description, diagnosis and judgement, if part of schools and users receiving different forms of evidence is absent, the likely effect on descriptive, diagnostic and evaluative use and classification should be reported.[REF-01] [REF-16]

Interpretation of descriptive, diagnostic and evaluative use should separate observation, explanation and attribution. For decisions about separating description, diagnosis and judgement, for separating description, diagnosis and judgement, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For separating description, diagnosis and judgement, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining separating description, diagnosis and judgement, consequences should be proportionate to evidence strength. Within evidence on separating description, diagnosis and judgement, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-09] [REF-10] [REF-16]

Accountability completes separating description, diagnosis and judgement. For separating description, diagnosis and judgement, a material finding about the measure purpose should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining separating description, diagnosis and judgement, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on separating description, diagnosis and judgement, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-01] [REF-09] [REF-10] [REF-16]

3

Defining performance beyond a single outcome

Defining performance beyond a single outcome defines the classification question for the performance construct. For defining performance beyond a single outcome, the relevant population or units are learners and schools with multiple education duties, and the direct evidence concerns access, learning, wellbeing and progression. In examining defining performance beyond a single outcome, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on defining performance beyond a single outcome, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of defining performance beyond a single outcome, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-01] [REF-02] [REF-12] [REF-13]

The principal misclassification risk in defining performance beyond a single outcome is that one tested outcome stands for all school quality and public responsibility. In examining defining performance beyond a single outcome, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on defining performance beyond a single outcome, authorities should identify how learners, events and institutions enter or leave the performance construct, whether missingness is concentrated, and which assumptions drive the result. For classification of defining performance beyond a single outcome, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-01] [REF-02]

The minimum measurement response for defining performance beyond a single outcome is to use a coherent limited construct and retain component results. Within evidence on defining performance beyond a single outcome, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of defining performance beyond a single outcome, it should also state the decision rule, consequence and evidence required for review. In interpreting defining performance beyond a single outcome, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-02] [REF-12]

Source fitness for the performance construct depends on coverage. For classification of defining performance beyond a single outcome, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting defining performance beyond a single outcome, for defining performance beyond a single outcome, each source should retain its principal error. For decisions about defining performance beyond a single outcome, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-12] [REF-13]

Equity is part of defining performance beyond a single outcome. In interpreting defining performance beyond a single outcome, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about defining performance beyond a single outcome, group levels and population shares should remain visible beside any school average. For defining performance beyond a single outcome, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining defining performance beyond a single outcome, if part of learners and schools with multiple education duties is absent, the likely effect on access, learning, wellbeing and progression and classification should be reported.[REF-01] [REF-13]

Interpretation of access, learning, wellbeing and progression should separate observation, explanation and attribution. For decisions about defining performance beyond a single outcome, for defining performance beyond a single outcome, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For defining performance beyond a single outcome, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining defining performance beyond a single outcome, consequences should be proportionate to evidence strength. Within evidence on defining performance beyond a single outcome, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-02] [REF-12] [REF-13]

Accountability completes defining performance beyond a single outcome. For defining performance beyond a single outcome, a material finding about the performance construct should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining defining performance beyond a single outcome, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on defining performance beyond a single outcome, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-01] [REF-02] [REF-12] [REF-13]

4

Reference period and cohort alignment

Reference period and cohort alignment defines the classification question for the performance period. For reference period and cohort alignment, the relevant population or units are learners observed across school years and stages, and the direct evidence concerns annual or cohort result. In examining reference period and cohort alignment, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on reference period and cohort alignment, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of reference period and cohort alignment, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-05] [REF-06] [REF-17] [REF-19]

The principal misclassification risk in reference period and cohort alignment is that current resources are compared with outcomes produced by earlier pupils or mixed cohorts. In examining reference period and cohort alignment, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on reference period and cohort alignment, authorities should identify how learners, events and institutions enter or leave the performance period, whether missingness is concentrated, and which assumptions drive the result. For classification of reference period and cohort alignment, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-05] [REF-06]

The minimum measurement response for reference period and cohort alignment is to align exposure, cohort, event and outcome dates and disclose lag. Within evidence on reference period and cohort alignment, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of reference period and cohort alignment, it should also state the decision rule, consequence and evidence required for review. In interpreting reference period and cohort alignment, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-06] [REF-17]

Source fitness for the performance period depends on coverage. For classification of reference period and cohort alignment, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting reference period and cohort alignment, for reference period and cohort alignment, each source should retain its principal error. For decisions about reference period and cohort alignment, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-17] [REF-19]

Equity is part of reference period and cohort alignment. In interpreting reference period and cohort alignment, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about reference period and cohort alignment, group levels and population shares should remain visible beside any school average. For reference period and cohort alignment, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining reference period and cohort alignment, if part of learners observed across school years and stages is absent, the likely effect on annual or cohort result and classification should be reported.[REF-05] [REF-19]

Interpretation of annual or cohort result should separate observation, explanation and attribution. For decisions about reference period and cohort alignment, for reference period and cohort alignment, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For reference period and cohort alignment, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining reference period and cohort alignment, consequences should be proportionate to evidence strength. Within evidence on reference period and cohort alignment, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-06] [REF-17] [REF-19]

Accountability completes reference period and cohort alignment. For reference period and cohort alignment, a material finding about the performance period should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining reference period and cohort alignment, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on reference period and cohort alignment, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-05] [REF-06] [REF-17] [REF-19]

5

Minimum standards and relative ranks

Minimum standards and relative ranks defines the classification question for the classification threshold. For minimum standards and relative ranks, the relevant population or units are schools compared with standards and one another, and the direct evidence concerns absolute level and relative position. In examining minimum standards and relative ranks, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on minimum standards and relative ranks, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of minimum standards and relative ranks, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-02] [REF-10] [REF-18] [REF-21]

The principal misclassification risk in minimum standards and relative ranks is that rank differences replace substantive adequacy or a low national standard normalises weak performance. In examining minimum standards and relative ranks, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on minimum standards and relative ranks, authorities should identify how learners, events and institutions enter or leave the classification threshold, whether missingness is concentrated, and which assumptions drive the result. For classification of minimum standards and relative ranks, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-02] [REF-10]

The minimum measurement response for minimum standards and relative ranks is to report threshold attainment, distance and uncertainty beside rank. Within evidence on minimum standards and relative ranks, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of minimum standards and relative ranks, it should also state the decision rule, consequence and evidence required for review. In interpreting minimum standards and relative ranks, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-10] [REF-18]

Source fitness for the classification threshold depends on coverage. For classification of minimum standards and relative ranks, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting minimum standards and relative ranks, for minimum standards and relative ranks, each source should retain its principal error. For decisions about minimum standards and relative ranks, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-18] [REF-21]

Equity is part of minimum standards and relative ranks. In interpreting minimum standards and relative ranks, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about minimum standards and relative ranks, group levels and population shares should remain visible beside any school average. For minimum standards and relative ranks, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining minimum standards and relative ranks, if part of schools compared with standards and one another is absent, the likely effect on absolute level and relative position and classification should be reported.[REF-02] [REF-21]

Interpretation of absolute level and relative position should separate observation, explanation and attribution. For decisions about minimum standards and relative ranks, for minimum standards and relative ranks, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For minimum standards and relative ranks, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining minimum standards and relative ranks, consequences should be proportionate to evidence strength. Within evidence on minimum standards and relative ranks, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-10] [REF-18] [REF-21]

Accountability completes minimum standards and relative ranks. For minimum standards and relative ranks, a material finding about the classification threshold should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining minimum standards and relative ranks, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on minimum standards and relative ranks, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-02] [REF-10] [REF-18] [REF-21]

6

High-stakes use and proportional evidence

High-stakes use and proportional evidence defines the classification question for the decision proportionality. For high-stakes use and proportional evidence, the relevant population or units are schools, staff and learners affected by classification, and the direct evidence concerns evidence strength and consequence. In examining high-stakes use and proportional evidence, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on high-stakes use and proportional evidence, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of high-stakes use and proportional evidence, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-09] [REF-10] [REF-16] [REF-20]

The principal misclassification risk in high-stakes use and proportional evidence is that weak or unstable estimates trigger sanctions, closure or reputational harm. In examining high-stakes use and proportional evidence, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on high-stakes use and proportional evidence, authorities should identify how learners, events and institutions enter or leave the decision proportionality, whether missingness is concentrated, and which assumptions drive the result. For classification of high-stakes use and proportional evidence, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-09] [REF-10]

The minimum measurement response for high-stakes use and proportional evidence is to raise evidence, review and appeal requirements with the severity of consequence. Within evidence on high-stakes use and proportional evidence, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of high-stakes use and proportional evidence, it should also state the decision rule, consequence and evidence required for review. In interpreting high-stakes use and proportional evidence, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-10] [REF-16]

Source fitness for the decision proportionality depends on coverage. For classification of high-stakes use and proportional evidence, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting high-stakes use and proportional evidence, for high-stakes use and proportional evidence, each source should retain its principal error. For decisions about high-stakes use and proportional evidence, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-16] [REF-20]

Equity is part of high-stakes use and proportional evidence. In interpreting high-stakes use and proportional evidence, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about high-stakes use and proportional evidence, group levels and population shares should remain visible beside any school average. For high-stakes use and proportional evidence, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining high-stakes use and proportional evidence, if part of schools, staff and learners affected by classification is absent, the likely effect on evidence strength and consequence and classification should be reported.[REF-09] [REF-20]

Interpretation of evidence strength and consequence should separate observation, explanation and attribution. For decisions about high-stakes use and proportional evidence, for high-stakes use and proportional evidence, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For high-stakes use and proportional evidence, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining high-stakes use and proportional evidence, consequences should be proportionate to evidence strength. Within evidence on high-stakes use and proportional evidence, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-10] [REF-16] [REF-20]

Accountability completes high-stakes use and proportional evidence. For high-stakes use and proportional evidence, a material finding about the decision proportionality should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining high-stakes use and proportional evidence, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on high-stakes use and proportional evidence, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-09] [REF-10] [REF-16] [REF-20]

Part II

Measures, sources and construct validity

7

Achievement levels and score distributions

Achievement levels and score distributions defines the classification question for the achievement measure. For achievement levels and score distributions, the relevant population or units are eligible assessed learners and those excluded or absent, and the direct evidence concerns score level, distribution and threshold. In examining achievement levels and score distributions, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on achievement levels and score distributions, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of achievement levels and score distributions, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-02] [REF-05] [REF-17] [REF-18]

The principal misclassification risk in achievement levels and score distributions is that mean scores conceal the learning floor, exclusions and internal inequality. In examining achievement levels and score distributions, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on achievement levels and score distributions, authorities should identify how learners, events and institutions enter or leave the achievement measure, whether missingness is concentrated, and which assumptions drive the result. For classification of achievement levels and score distributions, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-02] [REF-05]

The minimum measurement response for achievement levels and score distributions is to publish coverage, distribution, threshold shares and uncertainty. Within evidence on achievement levels and score distributions, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of achievement levels and score distributions, it should also state the decision rule, consequence and evidence required for review. In interpreting achievement levels and score distributions, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-05] [REF-17]

Source fitness for the achievement measure depends on coverage. For classification of achievement levels and score distributions, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting achievement levels and score distributions, for achievement levels and score distributions, each source should retain its principal error. For decisions about achievement levels and score distributions, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-17] [REF-18]

Equity is part of achievement levels and score distributions. In interpreting achievement levels and score distributions, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about achievement levels and score distributions, group levels and population shares should remain visible beside any school average. For achievement levels and score distributions, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining achievement levels and score distributions, if part of eligible assessed learners and those excluded or absent is absent, the likely effect on score level, distribution and threshold and classification should be reported.[REF-02] [REF-18]

Interpretation of score level, distribution and threshold should separate observation, explanation and attribution. For decisions about achievement levels and score distributions, for achievement levels and score distributions, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For achievement levels and score distributions, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining achievement levels and score distributions, consequences should be proportionate to evidence strength. Within evidence on achievement levels and score distributions, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-05] [REF-17] [REF-18]

Accountability completes achievement levels and score distributions. For achievement levels and score distributions, a material finding about the achievement measure should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining achievement levels and score distributions, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on achievement levels and score distributions, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-02] [REF-05] [REF-17] [REF-18]

8

Learner progress and value-added claims

Learner progress and value-added claims defines the classification question for the progress measure. For learner progress and value-added claims, the relevant population or units are cohorts with comparable prior and later evidence, and the direct evidence concerns change conditional on prior attainment. In examining learner progress and value-added claims, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on learner progress and value-added claims, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of learner progress and value-added claims, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-01] [REF-08] [REF-17] [REF-21]

The principal misclassification risk in learner progress and value-added claims is that missing prior scores, mobility and model choice create apparent school effects. In examining learner progress and value-added claims, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on learner progress and value-added claims, authorities should identify how learners, events and institutions enter or leave the progress measure, whether missingness is concentrated, and which assumptions drive the result. For classification of learner progress and value-added claims, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-01] [REF-08]

The minimum measurement response for learner progress and value-added claims is to state cohort coverage, assumptions, sensitivity and limits on causal interpretation. Within evidence on learner progress and value-added claims, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of learner progress and value-added claims, it should also state the decision rule, consequence and evidence required for review. In interpreting learner progress and value-added claims, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-08] [REF-17]

Source fitness for the progress measure depends on coverage. For classification of learner progress and value-added claims, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting learner progress and value-added claims, for learner progress and value-added claims, each source should retain its principal error. For decisions about learner progress and value-added claims, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-17] [REF-21]

Equity is part of learner progress and value-added claims. In interpreting learner progress and value-added claims, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about learner progress and value-added claims, group levels and population shares should remain visible beside any school average. For learner progress and value-added claims, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining learner progress and value-added claims, if part of cohorts with comparable prior and later evidence is absent, the likely effect on change conditional on prior attainment and classification should be reported.[REF-01] [REF-21]

Interpretation of change conditional on prior attainment should separate observation, explanation and attribution. For decisions about learner progress and value-added claims, for learner progress and value-added claims, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For learner progress and value-added claims, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining learner progress and value-added claims, consequences should be proportionate to evidence strength. Within evidence on learner progress and value-added claims, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-08] [REF-17] [REF-21]

Accountability completes learner progress and value-added claims. For learner progress and value-added claims, a material finding about the progress measure should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining learner progress and value-added claims, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on learner progress and value-added claims, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-01] [REF-08] [REF-17] [REF-21]

9

Attendance, exclusion and participation

Attendance, exclusion and participation defines the classification question for the participation measure. For attendance, exclusion and participation, the relevant population or units are enrolled and eligible learners using the school, and the direct evidence concerns attendance intensity and exclusion. In examining attendance, exclusion and participation, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on attendance, exclusion and participation, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of attendance, exclusion and participation, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-04] [REF-11] [REF-14] [REF-18]

The principal misclassification risk in attendance, exclusion and participation is that attendance records omit never-enrolled learners or disciplinary exclusion raises averages. In examining attendance, exclusion and participation, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on attendance, exclusion and participation, authorities should identify how learners, events and institutions enter or leave the participation measure, whether missingness is concentrated, and which assumptions drive the result. For classification of attendance, exclusion and participation, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-04] [REF-11]

The minimum measurement response for attendance, exclusion and participation is to separate eligibility, enrolment, presence, exclusion and sustained participation. Within evidence on attendance, exclusion and participation, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of attendance, exclusion and participation, it should also state the decision rule, consequence and evidence required for review. In interpreting attendance, exclusion and participation, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-11] [REF-14]

Source fitness for the participation measure depends on coverage. For classification of attendance, exclusion and participation, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting attendance, exclusion and participation, for attendance, exclusion and participation, each source should retain its principal error. For decisions about attendance, exclusion and participation, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-14] [REF-18]

Equity is part of attendance, exclusion and participation. In interpreting attendance, exclusion and participation, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about attendance, exclusion and participation, group levels and population shares should remain visible beside any school average. For attendance, exclusion and participation, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining attendance, exclusion and participation, if part of enrolled and eligible learners using the school is absent, the likely effect on attendance intensity and exclusion and classification should be reported.[REF-04] [REF-18]

Interpretation of attendance intensity and exclusion should separate observation, explanation and attribution. For decisions about attendance, exclusion and participation, for attendance, exclusion and participation, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For attendance, exclusion and participation, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining attendance, exclusion and participation, consequences should be proportionate to evidence strength. Within evidence on attendance, exclusion and participation, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-11] [REF-14] [REF-18]

Accountability completes attendance, exclusion and participation. For attendance, exclusion and participation, a material finding about the participation measure should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining attendance, exclusion and participation, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on attendance, exclusion and participation, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-04] [REF-11] [REF-14] [REF-18]

10

Completion, retention and transition

Completion, retention and transition defines the classification question for the pathway measure. For completion, retention and transition, the relevant population or units are cohorts approaching programme completion, and the direct evidence concerns retention, completion and destination. In examining completion, retention and transition, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on completion, retention and transition, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of completion, retention and transition, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-05] [REF-06] [REF-18] [REF-24]

The principal misclassification risk in completion, retention and transition is that cross-sectional counts reward selective retention or omit transfers and late completers. In examining completion, retention and transition, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on completion, retention and transition, authorities should identify how learners, events and institutions enter or leave the pathway measure, whether missingness is concentrated, and which assumptions drive the result. For classification of completion, retention and transition, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-05] [REF-06]

The minimum measurement response for completion, retention and transition is to trace defined cohorts, exits, returns, credentials and sustained destinations. Within evidence on completion, retention and transition, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of completion, retention and transition, it should also state the decision rule, consequence and evidence required for review. In interpreting completion, retention and transition, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-06] [REF-18]

Source fitness for the pathway measure depends on coverage. For classification of completion, retention and transition, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting completion, retention and transition, for completion, retention and transition, each source should retain its principal error. For decisions about completion, retention and transition, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-18] [REF-24]

Equity is part of completion, retention and transition. In interpreting completion, retention and transition, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about completion, retention and transition, group levels and population shares should remain visible beside any school average. For completion, retention and transition, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining completion, retention and transition, if part of cohorts approaching programme completion is absent, the likely effect on retention, completion and destination and classification should be reported.[REF-05] [REF-24]

Interpretation of retention, completion and destination should separate observation, explanation and attribution. For decisions about completion, retention and transition, for completion, retention and transition, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For completion, retention and transition, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining completion, retention and transition, consequences should be proportionate to evidence strength. Within evidence on completion, retention and transition, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-06] [REF-18] [REF-24]

Accountability completes completion, retention and transition. For completion, retention and transition, a material finding about the pathway measure should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining completion, retention and transition, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on completion, retention and transition, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-05] [REF-06] [REF-18] [REF-24]

11

School climate, safety and learner voice

School climate, safety and learner voice defines the classification question for the climate measure. For school climate, safety and learner voice, the relevant population or units are learners and staff experiencing school conditions, and the direct evidence concerns safety, belonging and participation. In examining school climate, safety and learner voice, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on school climate, safety and learner voice, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of school climate, safety and learner voice, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-10] [REF-12] [REF-14] [REF-15]

The principal misclassification risk in school climate, safety and learner voice is that response fear, non-response and socially preferred answers create false assurance. In examining school climate, safety and learner voice, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on school climate, safety and learner voice, authorities should identify how learners, events and institutions enter or leave the climate measure, whether missingness is concentrated, and which assumptions drive the result. For classification of school climate, safety and learner voice, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-10] [REF-12]

The minimum measurement response for school climate, safety and learner voice is to use confidential multi-source evidence and report response and subgroup patterns. Within evidence on school climate, safety and learner voice, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of school climate, safety and learner voice, it should also state the decision rule, consequence and evidence required for review. In interpreting school climate, safety and learner voice, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-12] [REF-14]

Source fitness for the climate measure depends on coverage. For classification of school climate, safety and learner voice, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting school climate, safety and learner voice, for school climate, safety and learner voice, each source should retain its principal error. For decisions about school climate, safety and learner voice, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-14] [REF-15]

Equity is part of school climate, safety and learner voice. In interpreting school climate, safety and learner voice, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about school climate, safety and learner voice, group levels and population shares should remain visible beside any school average. For school climate, safety and learner voice, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining school climate, safety and learner voice, if part of learners and staff experiencing school conditions is absent, the likely effect on safety, belonging and participation and classification should be reported.[REF-10] [REF-15]

Interpretation of safety, belonging and participation should separate observation, explanation and attribution. For decisions about school climate, safety and learner voice, for school climate, safety and learner voice, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For school climate, safety and learner voice, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining school climate, safety and learner voice, consequences should be proportionate to evidence strength. Within evidence on school climate, safety and learner voice, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-12] [REF-14] [REF-15]

Accountability completes school climate, safety and learner voice. For school climate, safety and learner voice, a material finding about the climate measure should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining school climate, safety and learner voice, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on school climate, safety and learner voice, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-10] [REF-12] [REF-14] [REF-15]

12

Composite indices and hidden trade-offs

Composite indices and hidden trade-offs defines the classification question for the composite performance measure. For composite indices and hidden trade-offs, the relevant population or units are schools summarised across component indicators, and the direct evidence concerns weighted aggregate and components. In examining composite indices and hidden trade-offs, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on composite indices and hidden trade-offs, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of composite indices and hidden trade-offs, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-01] [REF-19] [REF-20] [REF-21]

The principal misclassification risk in composite indices and hidden trade-offs is that normalisation and weights compensate severe weakness and make rank sensitive to arbitrary choices. In examining composite indices and hidden trade-offs, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on composite indices and hidden trade-offs, authorities should identify how learners, events and institutions enter or leave the composite performance measure, whether missingness is concentrated, and which assumptions drive the result. For classification of composite indices and hidden trade-offs, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-01] [REF-19]

The minimum measurement response for composite indices and hidden trade-offs is to publish conceptual model, components, weights, sensitivity and non-compensable floors. Within evidence on composite indices and hidden trade-offs, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of composite indices and hidden trade-offs, it should also state the decision rule, consequence and evidence required for review. In interpreting composite indices and hidden trade-offs, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-19] [REF-20]

Source fitness for the composite performance measure depends on coverage. For classification of composite indices and hidden trade-offs, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting composite indices and hidden trade-offs, for composite indices and hidden trade-offs, each source should retain its principal error. For decisions about composite indices and hidden trade-offs, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-20] [REF-21]

Equity is part of composite indices and hidden trade-offs. In interpreting composite indices and hidden trade-offs, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about composite indices and hidden trade-offs, group levels and population shares should remain visible beside any school average. For composite indices and hidden trade-offs, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining composite indices and hidden trade-offs, if part of schools summarised across component indicators is absent, the likely effect on weighted aggregate and components and classification should be reported.[REF-01] [REF-21]

Interpretation of weighted aggregate and components should separate observation, explanation and attribution. For decisions about composite indices and hidden trade-offs, for composite indices and hidden trade-offs, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For composite indices and hidden trade-offs, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining composite indices and hidden trade-offs, consequences should be proportionate to evidence strength. Within evidence on composite indices and hidden trade-offs, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-19] [REF-20] [REF-21]

Accountability completes composite indices and hidden trade-offs. For composite indices and hidden trade-offs, a material finding about the composite performance measure should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining composite indices and hidden trade-offs, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on composite indices and hidden trade-offs, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-01] [REF-19] [REF-20] [REF-21]

Part III

Intake, context and fair comparison

13

Prior attainment and starting points

Prior attainment and starting points defines the classification question for the intake baseline. For prior attainment and starting points, the relevant population or units are learners entering with varied prior opportunity, and the direct evidence concerns prior attainment and later result. In examining prior attainment and starting points, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on prior attainment and starting points, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of prior attainment and starting points, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-02] [REF-04] [REF-17] [REF-21]

The principal misclassification risk in prior attainment and starting points is that raw outcomes classify schools largely by initial intake rather than education received. In examining prior attainment and starting points, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on prior attainment and starting points, authorities should identify how learners, events and institutions enter or leave the intake baseline, whether missingness is concentrated, and which assumptions drive the result. For classification of prior attainment and starting points, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-02] [REF-04]

The minimum measurement response for prior attainment and starting points is to report starting distributions and progress while retaining absolute learning levels. Within evidence on prior attainment and starting points, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of prior attainment and starting points, it should also state the decision rule, consequence and evidence required for review. In interpreting prior attainment and starting points, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-04] [REF-17]

Source fitness for the intake baseline depends on coverage. For classification of prior attainment and starting points, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting prior attainment and starting points, for prior attainment and starting points, each source should retain its principal error. For decisions about prior attainment and starting points, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-17] [REF-21]

Equity is part of prior attainment and starting points. In interpreting prior attainment and starting points, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about prior attainment and starting points, group levels and population shares should remain visible beside any school average. For prior attainment and starting points, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining prior attainment and starting points, if part of learners entering with varied prior opportunity is absent, the likely effect on prior attainment and later result and classification should be reported.[REF-02] [REF-21]

Interpretation of prior attainment and later result should separate observation, explanation and attribution. For decisions about prior attainment and starting points, for prior attainment and starting points, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For prior attainment and starting points, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining prior attainment and starting points, consequences should be proportionate to evidence strength. Within evidence on prior attainment and starting points, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-04] [REF-17] [REF-21]

Accountability completes prior attainment and starting points. For prior attainment and starting points, a material finding about the intake baseline should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining prior attainment and starting points, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on prior attainment and starting points, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-02] [REF-04] [REF-17] [REF-21]

14

Poverty and household-resource composition

Poverty and household-resource composition defines the classification question for the socioeconomic context. For poverty and household-resource composition, the relevant population or units are learners across household resources and material barriers, and the direct evidence concerns school intake and outcome gradient. In examining poverty and household-resource composition, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on poverty and household-resource composition, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of poverty and household-resource composition, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-03] [REF-04] [REF-07] [REF-11]

The principal misclassification risk in poverty and household-resource composition is that school location or free-meal proxies misclassify household disadvantage and become excuses for low standards. In examining poverty and household-resource composition, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on poverty and household-resource composition, authorities should identify how learners, events and institutions enter or leave the socioeconomic context, whether missingness is concentrated, and which assumptions drive the result. For classification of poverty and household-resource composition, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-03] [REF-04]

The minimum measurement response for poverty and household-resource composition is to state the measure, missingness and diagnostic use and preserve the common floor. Within evidence on poverty and household-resource composition, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of poverty and household-resource composition, it should also state the decision rule, consequence and evidence required for review. In interpreting poverty and household-resource composition, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-04] [REF-07]

Source fitness for the socioeconomic context depends on coverage. For classification of poverty and household-resource composition, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting poverty and household-resource composition, for poverty and household-resource composition, each source should retain its principal error. For decisions about poverty and household-resource composition, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-07] [REF-11]

Equity is part of poverty and household-resource composition. In interpreting poverty and household-resource composition, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about poverty and household-resource composition, group levels and population shares should remain visible beside any school average. For poverty and household-resource composition, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining poverty and household-resource composition, if part of learners across household resources and material barriers is absent, the likely effect on school intake and outcome gradient and classification should be reported.[REF-03] [REF-11]

Interpretation of school intake and outcome gradient should separate observation, explanation and attribution. For decisions about poverty and household-resource composition, for poverty and household-resource composition, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For poverty and household-resource composition, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining poverty and household-resource composition, consequences should be proportionate to evidence strength. Within evidence on poverty and household-resource composition, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-04] [REF-07] [REF-11]

Accountability completes poverty and household-resource composition. For poverty and household-resource composition, a material finding about the socioeconomic context should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining poverty and household-resource composition, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on poverty and household-resource composition, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-03] [REF-04] [REF-07] [REF-11]

15

Language, migration and refugee mobility

Language, migration and refugee mobility defines the classification question for the mobility context. For language, migration and refugee mobility, the relevant population or units are newly arrived, refugee and multilingual learners, and the direct evidence concerns entry date, language support and continuity. In examining language, migration and refugee mobility, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on language, migration and refugee mobility, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of language, migration and refugee mobility, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-22] [REF-23] [REF-24] [REF-13]

The principal misclassification risk in language, migration and refugee mobility is that schools serving mobile learners are penalised for partial exposure or excluded learners vanish from cohorts. In examining language, migration and refugee mobility, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on language, migration and refugee mobility, authorities should identify how learners, events and institutions enter or leave the mobility context, whether missingness is concentrated, and which assumptions drive the result. For classification of language, migration and refugee mobility, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-22] [REF-23]

The minimum measurement response for language, migration and refugee mobility is to report exposure time, mobility, support, coverage and portable progression. Within evidence on language, migration and refugee mobility, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of language, migration and refugee mobility, it should also state the decision rule, consequence and evidence required for review. In interpreting language, migration and refugee mobility, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-23] [REF-24]

Source fitness for the mobility context depends on coverage. For classification of language, migration and refugee mobility, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting language, migration and refugee mobility, for language, migration and refugee mobility, each source should retain its principal error. For decisions about language, migration and refugee mobility, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-24] [REF-13]

Equity is part of language, migration and refugee mobility. In interpreting language, migration and refugee mobility, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about language, migration and refugee mobility, group levels and population shares should remain visible beside any school average. For language, migration and refugee mobility, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining language, migration and refugee mobility, if part of newly arrived, refugee and multilingual learners is absent, the likely effect on entry date, language support and continuity and classification should be reported.[REF-22] [REF-13]

Interpretation of entry date, language support and continuity should separate observation, explanation and attribution. For decisions about language, migration and refugee mobility, for language, migration and refugee mobility, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For language, migration and refugee mobility, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining language, migration and refugee mobility, consequences should be proportionate to evidence strength. Within evidence on language, migration and refugee mobility, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-23] [REF-24] [REF-13]

Accountability completes language, migration and refugee mobility. For language, migration and refugee mobility, a material finding about the mobility context should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining language, migration and refugee mobility, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on language, migration and refugee mobility, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-22] [REF-23] [REF-24] [REF-13]

16

Disability and accommodation

Disability and accommodation defines the classification question for the disability context. For disability and accommodation, the relevant population or units are learners with different functional and support requirements, and the direct evidence concerns participation, accommodation and learning. In examining disability and accommodation, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on disability and accommodation, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of disability and accommodation, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-10] [REF-11] [REF-15] [REF-20]

The principal misclassification risk in disability and accommodation is that diagnosis categories are incomparable or exclusions raise school averages. In examining disability and accommodation, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on disability and accommodation, authorities should identify how learners, events and institutions enter or leave the disability context, whether missingness is concentrated, and which assumptions drive the result. For classification of disability and accommodation, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-10] [REF-11]

The minimum measurement response for disability and accommodation is to report accessibility, accommodation, coverage and substantive progress without lowering entitlement. Within evidence on disability and accommodation, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of disability and accommodation, it should also state the decision rule, consequence and evidence required for review. In interpreting disability and accommodation, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-11] [REF-15]

Source fitness for the disability context depends on coverage. For classification of disability and accommodation, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting disability and accommodation, for disability and accommodation, each source should retain its principal error. For decisions about disability and accommodation, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-15] [REF-20]

Equity is part of disability and accommodation. In interpreting disability and accommodation, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about disability and accommodation, group levels and population shares should remain visible beside any school average. For disability and accommodation, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining disability and accommodation, if part of learners with different functional and support requirements is absent, the likely effect on participation, accommodation and learning and classification should be reported.[REF-10] [REF-20]

Interpretation of participation, accommodation and learning should separate observation, explanation and attribution. For decisions about disability and accommodation, for disability and accommodation, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For disability and accommodation, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining disability and accommodation, consequences should be proportionate to evidence strength. Within evidence on disability and accommodation, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-11] [REF-15] [REF-20]

Accountability completes disability and accommodation. For disability and accommodation, a material finding about the disability context should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining disability and accommodation, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on disability and accommodation, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-10] [REF-11] [REF-15] [REF-20]

17

School size and estimate instability

School size and estimate instability defines the classification question for the small-school context. For school size and estimate instability, the relevant population or units are schools with small cohorts or rare outcomes, and the direct evidence concerns sample size and year-to-year variation. In examining school size and estimate instability, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on school size and estimate instability, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of school size and estimate instability, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-08] [REF-09] [REF-19] [REF-21]

The principal misclassification risk in school size and estimate instability is that a few pupils or events cause large rank movement interpreted as school change. In examining school size and estimate instability, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on school size and estimate instability, authorities should identify how learners, events and institutions enter or leave the small-school context, whether missingness is concentrated, and which assumptions drive the result. For classification of school size and estimate instability, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-08] [REF-09]

The minimum measurement response for school size and estimate instability is to publish counts, intervals, pooled evidence and limits on classification. Within evidence on school size and estimate instability, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of school size and estimate instability, it should also state the decision rule, consequence and evidence required for review. In interpreting school size and estimate instability, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-09] [REF-19]

Source fitness for the small-school context depends on coverage. For classification of school size and estimate instability, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting school size and estimate instability, for school size and estimate instability, each source should retain its principal error. For decisions about school size and estimate instability, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-19] [REF-21]

Equity is part of school size and estimate instability. In interpreting school size and estimate instability, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about school size and estimate instability, group levels and population shares should remain visible beside any school average. For school size and estimate instability, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining school size and estimate instability, if part of schools with small cohorts or rare outcomes is absent, the likely effect on sample size and year-to-year variation and classification should be reported.[REF-08] [REF-21]

Interpretation of sample size and year-to-year variation should separate observation, explanation and attribution. For decisions about school size and estimate instability, for school size and estimate instability, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For school size and estimate instability, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining school size and estimate instability, consequences should be proportionate to evidence strength. Within evidence on school size and estimate instability, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-09] [REF-19] [REF-21]

Accountability completes school size and estimate instability. For school size and estimate instability, a material finding about the small-school context should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining school size and estimate instability, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on school size and estimate instability, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-08] [REF-09] [REF-19] [REF-21]

18

Resources and opportunity to learn

Resources and opportunity to learn defines the classification question for the service context. For resources and opportunity to learn, the relevant population or units are schools with different staffing, time, facilities and support, and the direct evidence concerns learner exposure to educational conditions. In examining resources and opportunity to learn, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on resources and opportunity to learn, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of resources and opportunity to learn, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-01] [REF-02] [REF-03] [REF-18]

The principal misclassification risk in resources and opportunity to learn is that context adjustment removes resource inequality from accountability rather than identifying it. In examining resources and opportunity to learn, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on resources and opportunity to learn, authorities should identify how learners, events and institutions enter or leave the service context, whether missingness is concentrated, and which assumptions drive the result. For classification of resources and opportunity to learn, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-01] [REF-02]

The minimum measurement response for resources and opportunity to learn is to report service conditions separately and use them to assign system responsibility. Within evidence on resources and opportunity to learn, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of resources and opportunity to learn, it should also state the decision rule, consequence and evidence required for review. In interpreting resources and opportunity to learn, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-02] [REF-03]

Source fitness for the service context depends on coverage. For classification of resources and opportunity to learn, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting resources and opportunity to learn, for resources and opportunity to learn, each source should retain its principal error. For decisions about resources and opportunity to learn, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-03] [REF-18]

Equity is part of resources and opportunity to learn. In interpreting resources and opportunity to learn, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about resources and opportunity to learn, group levels and population shares should remain visible beside any school average. For resources and opportunity to learn, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining resources and opportunity to learn, if part of schools with different staffing, time, facilities and support is absent, the likely effect on learner exposure to educational conditions and classification should be reported.[REF-01] [REF-18]

Interpretation of learner exposure to educational conditions should separate observation, explanation and attribution. For decisions about resources and opportunity to learn, for resources and opportunity to learn, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For resources and opportunity to learn, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining resources and opportunity to learn, consequences should be proportionate to evidence strength. Within evidence on resources and opportunity to learn, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-02] [REF-03] [REF-18]

Accountability completes resources and opportunity to learn. For resources and opportunity to learn, a material finding about the service context should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining resources and opportunity to learn, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on resources and opportunity to learn, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-01] [REF-02] [REF-03] [REF-18]

Part IV

Statistical error and misclassification

19

Sampling error and confidence around school estimates

Sampling error and confidence around school estimates defines the classification question for the sampling uncertainty. For sampling error and confidence around school estimates, the relevant population or units are schools represented by assessed samples or surveyed respondents, and the direct evidence concerns interval around the estimate. In examining sampling error and confidence around school estimates, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on sampling error and confidence around school estimates, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of sampling error and confidence around school estimates, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-08] [REF-09] [REF-17] [REF-19]

The principal misclassification risk in sampling error and confidence around school estimates is that point estimates create precise ranks unsupported by the sample. In examining sampling error and confidence around school estimates, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on sampling error and confidence around school estimates, authorities should identify how learners, events and institutions enter or leave the sampling uncertainty, whether missingness is concentrated, and which assumptions drive the result. For classification of sampling error and confidence around school estimates, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-08] [REF-09]

The minimum measurement response for sampling error and confidence around school estimates is to publish intervals, effective samples and probability of material difference. Within evidence on sampling error and confidence around school estimates, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of sampling error and confidence around school estimates, it should also state the decision rule, consequence and evidence required for review. In interpreting sampling error and confidence around school estimates, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-09] [REF-17]

Source fitness for the sampling uncertainty depends on coverage. For classification of sampling error and confidence around school estimates, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting sampling error and confidence around school estimates, for sampling error and confidence around school estimates, each source should retain its principal error. For decisions about sampling error and confidence around school estimates, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-17] [REF-19]

Equity is part of sampling error and confidence around school estimates. In interpreting sampling error and confidence around school estimates, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about sampling error and confidence around school estimates, group levels and population shares should remain visible beside any school average. For sampling error and confidence around school estimates, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining sampling error and confidence around school estimates, if part of schools represented by assessed samples or surveyed respondents is absent, the likely effect on interval around the estimate and classification should be reported.[REF-08] [REF-19]

Interpretation of interval around the estimate should separate observation, explanation and attribution. For decisions about sampling error and confidence around school estimates, for sampling error and confidence around school estimates, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For sampling error and confidence around school estimates, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining sampling error and confidence around school estimates, consequences should be proportionate to evidence strength. Within evidence on sampling error and confidence around school estimates, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-09] [REF-17] [REF-19]

Accountability completes sampling error and confidence around school estimates. For sampling error and confidence around school estimates, a material finding about the sampling uncertainty should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining sampling error and confidence around school estimates, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on sampling error and confidence around school estimates, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-08] [REF-09] [REF-17] [REF-19]

20

Non-response, absence and selective coverage

Non-response, absence and selective coverage defines the classification question for the coverage error. For non-response, absence and selective coverage, the relevant population or units are learners missing from tests, surveys or records, and the direct evidence concerns participation and missing pattern. In examining non-response, absence and selective coverage, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on non-response, absence and selective coverage, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of non-response, absence and selective coverage, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-08] [REF-11] [REF-17] [REF-19]

The principal misclassification risk in non-response, absence and selective coverage is that low participation or strategic absence improves apparent school performance. In examining non-response, absence and selective coverage, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on non-response, absence and selective coverage, authorities should identify how learners, events and institutions enter or leave the coverage error, whether missingness is concentrated, and which assumptions drive the result. For classification of non-response, absence and selective coverage, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-08] [REF-11]

The minimum measurement response for non-response, absence and selective coverage is to publish target population, response, exclusions and sensitivity to missing outcomes. Within evidence on non-response, absence and selective coverage, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of non-response, absence and selective coverage, it should also state the decision rule, consequence and evidence required for review. In interpreting non-response, absence and selective coverage, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-11] [REF-17]

Source fitness for the coverage error depends on coverage. For classification of non-response, absence and selective coverage, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting non-response, absence and selective coverage, for non-response, absence and selective coverage, each source should retain its principal error. For decisions about non-response, absence and selective coverage, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-17] [REF-19]

Equity is part of non-response, absence and selective coverage. In interpreting non-response, absence and selective coverage, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about non-response, absence and selective coverage, group levels and population shares should remain visible beside any school average. For non-response, absence and selective coverage, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining non-response, absence and selective coverage, if part of learners missing from tests, surveys or records is absent, the likely effect on participation and missing pattern and classification should be reported.[REF-08] [REF-19]

Interpretation of participation and missing pattern should separate observation, explanation and attribution. For decisions about non-response, absence and selective coverage, for non-response, absence and selective coverage, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For non-response, absence and selective coverage, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining non-response, absence and selective coverage, consequences should be proportionate to evidence strength. Within evidence on non-response, absence and selective coverage, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-11] [REF-17] [REF-19]

Accountability completes non-response, absence and selective coverage. For non-response, absence and selective coverage, a material finding about the coverage error should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining non-response, absence and selective coverage, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on non-response, absence and selective coverage, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-08] [REF-11] [REF-17] [REF-19]

21

Measurement error and score reliability

Measurement error and score reliability defines the classification question for the measurement reliability. For measurement error and score reliability, the relevant population or units are learners completing finite tasks or questionnaires, and the direct evidence concerns observed and underlying construct. In examining measurement error and score reliability, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on measurement error and score reliability, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of measurement error and score reliability, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-02] [REF-17] [REF-19] [REF-21]

The principal misclassification risk in measurement error and score reliability is that short instruments and rater differences are ignored in school classification. In examining measurement error and score reliability, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on measurement error and score reliability, authorities should identify how learners, events and institutions enter or leave the measurement reliability, whether missingness is concentrated, and which assumptions drive the result. For classification of measurement error and score reliability, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-02] [REF-17]

The minimum measurement response for measurement error and score reliability is to state reliability, scoring quality and classification uncertainty. Within evidence on measurement error and score reliability, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of measurement error and score reliability, it should also state the decision rule, consequence and evidence required for review. In interpreting measurement error and score reliability, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-17] [REF-19]

Source fitness for the measurement reliability depends on coverage. For classification of measurement error and score reliability, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting measurement error and score reliability, for measurement error and score reliability, each source should retain its principal error. For decisions about measurement error and score reliability, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-19] [REF-21]

Equity is part of measurement error and score reliability. In interpreting measurement error and score reliability, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about measurement error and score reliability, group levels and population shares should remain visible beside any school average. For measurement error and score reliability, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining measurement error and score reliability, if part of learners completing finite tasks or questionnaires is absent, the likely effect on observed and underlying construct and classification should be reported.[REF-02] [REF-21]

Interpretation of observed and underlying construct should separate observation, explanation and attribution. For decisions about measurement error and score reliability, for measurement error and score reliability, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For measurement error and score reliability, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining measurement error and score reliability, consequences should be proportionate to evidence strength. Within evidence on measurement error and score reliability, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-17] [REF-19] [REF-21]

Accountability completes measurement error and score reliability. For measurement error and score reliability, a material finding about the measurement reliability should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining measurement error and score reliability, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on measurement error and score reliability, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-02] [REF-17] [REF-19] [REF-21]

22

Threshold selection and schools near the boundary

Threshold selection and schools near the boundary defines the classification question for the classification boundary. For threshold selection and schools near the boundary, the relevant population or units are schools close to a performance cut point, and the direct evidence concerns distance and crossing probability. In examining threshold selection and schools near the boundary, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on threshold selection and schools near the boundary, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of threshold selection and schools near the boundary, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-09] [REF-10] [REF-16] [REF-21]

The principal misclassification risk in threshold selection and schools near the boundary is that minor noise changes category and triggers discontinuous consequences. In examining threshold selection and schools near the boundary, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on threshold selection and schools near the boundary, authorities should identify how learners, events and institutions enter or leave the classification boundary, whether missingness is concentrated, and which assumptions drive the result. For classification of threshold selection and schools near the boundary, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-09] [REF-10]

The minimum measurement response for threshold selection and schools near the boundary is to create review bands and require corroborating evidence near thresholds. Within evidence on threshold selection and schools near the boundary, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of threshold selection and schools near the boundary, it should also state the decision rule, consequence and evidence required for review. In interpreting threshold selection and schools near the boundary, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-10] [REF-16]

Source fitness for the classification boundary depends on coverage. For classification of threshold selection and schools near the boundary, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting threshold selection and schools near the boundary, for threshold selection and schools near the boundary, each source should retain its principal error. For decisions about threshold selection and schools near the boundary, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-16] [REF-21]

Equity is part of threshold selection and schools near the boundary. In interpreting threshold selection and schools near the boundary, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about threshold selection and schools near the boundary, group levels and population shares should remain visible beside any school average. For threshold selection and schools near the boundary, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining threshold selection and schools near the boundary, if part of schools close to a performance cut point is absent, the likely effect on distance and crossing probability and classification should be reported.[REF-09] [REF-21]

Interpretation of distance and crossing probability should separate observation, explanation and attribution. For decisions about threshold selection and schools near the boundary, for threshold selection and schools near the boundary, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For threshold selection and schools near the boundary, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining threshold selection and schools near the boundary, consequences should be proportionate to evidence strength. Within evidence on threshold selection and schools near the boundary, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-10] [REF-16] [REF-21]

Accountability completes threshold selection and schools near the boundary. For threshold selection and schools near the boundary, a material finding about the classification boundary should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining threshold selection and schools near the boundary, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on threshold selection and schools near the boundary, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-09] [REF-10] [REF-16] [REF-21]

23

Model specification and sensitivity

Model specification and sensitivity defines the classification question for the model uncertainty. For model specification and sensitivity, the relevant population or units are schools classified by adjusted or composite estimates, and the direct evidence concerns result under plausible alternatives. In examining model specification and sensitivity, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on model specification and sensitivity, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of model specification and sensitivity, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-08] [REF-19] [REF-20] [REF-21]

The principal misclassification risk in model specification and sensitivity is that one model, covariate set or weight is presented as uniquely correct. In examining model specification and sensitivity, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on model specification and sensitivity, authorities should identify how learners, events and institutions enter or leave the model uncertainty, whether missingness is concentrated, and which assumptions drive the result. For classification of model specification and sensitivity, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-08] [REF-19]

The minimum measurement response for model specification and sensitivity is to publish alternate specifications, influential assumptions and stable conclusions. Within evidence on model specification and sensitivity, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of model specification and sensitivity, it should also state the decision rule, consequence and evidence required for review. In interpreting model specification and sensitivity, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-19] [REF-20]

Source fitness for the model uncertainty depends on coverage. For classification of model specification and sensitivity, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting model specification and sensitivity, for model specification and sensitivity, each source should retain its principal error. For decisions about model specification and sensitivity, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-20] [REF-21]

Equity is part of model specification and sensitivity. In interpreting model specification and sensitivity, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about model specification and sensitivity, group levels and population shares should remain visible beside any school average. For model specification and sensitivity, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining model specification and sensitivity, if part of schools classified by adjusted or composite estimates is absent, the likely effect on result under plausible alternatives and classification should be reported.[REF-08] [REF-21]

Interpretation of result under plausible alternatives should separate observation, explanation and attribution. For decisions about model specification and sensitivity, for model specification and sensitivity, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For model specification and sensitivity, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining model specification and sensitivity, consequences should be proportionate to evidence strength. Within evidence on model specification and sensitivity, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-19] [REF-20] [REF-21]

Accountability completes model specification and sensitivity. For model specification and sensitivity, a material finding about the model uncertainty should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining model specification and sensitivity, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on model specification and sensitivity, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-08] [REF-19] [REF-20] [REF-21]

24

Year-to-year volatility and regression to the mean

Year-to-year volatility and regression to the mean defines the classification question for the temporal instability. For year-to-year volatility and regression to the mean, the relevant population or units are schools with extreme or changing annual results, and the direct evidence concerns trend and persistence. In examining year-to-year volatility and regression to the mean, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on year-to-year volatility and regression to the mean, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of year-to-year volatility and regression to the mean, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-05] [REF-09] [REF-16] [REF-19]

The principal misclassification risk in year-to-year volatility and regression to the mean is that one exceptional year is treated as durable quality or decline. In examining year-to-year volatility and regression to the mean, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on year-to-year volatility and regression to the mean, authorities should identify how learners, events and institutions enter or leave the temporal instability, whether missingness is concentrated, and which assumptions drive the result. For classification of year-to-year volatility and regression to the mean, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-05] [REF-09]

The minimum measurement response for year-to-year volatility and regression to the mean is to use multiple years where fit, retain current evidence and test persistence. Within evidence on year-to-year volatility and regression to the mean, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of year-to-year volatility and regression to the mean, it should also state the decision rule, consequence and evidence required for review. In interpreting year-to-year volatility and regression to the mean, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-09] [REF-16]

Source fitness for the temporal instability depends on coverage. For classification of year-to-year volatility and regression to the mean, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting year-to-year volatility and regression to the mean, for year-to-year volatility and regression to the mean, each source should retain its principal error. For decisions about year-to-year volatility and regression to the mean, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-16] [REF-19]

Equity is part of year-to-year volatility and regression to the mean. In interpreting year-to-year volatility and regression to the mean, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about year-to-year volatility and regression to the mean, group levels and population shares should remain visible beside any school average. For year-to-year volatility and regression to the mean, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining year-to-year volatility and regression to the mean, if part of schools with extreme or changing annual results is absent, the likely effect on trend and persistence and classification should be reported.[REF-05] [REF-19]

Interpretation of trend and persistence should separate observation, explanation and attribution. For decisions about year-to-year volatility and regression to the mean, for year-to-year volatility and regression to the mean, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For year-to-year volatility and regression to the mean, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining year-to-year volatility and regression to the mean, consequences should be proportionate to evidence strength. Within evidence on year-to-year volatility and regression to the mean, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-09] [REF-16] [REF-19]

Accountability completes year-to-year volatility and regression to the mean. For year-to-year volatility and regression to the mean, a material finding about the temporal instability should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining year-to-year volatility and regression to the mean, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on year-to-year volatility and regression to the mean, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-05] [REF-09] [REF-16] [REF-19]

Part V

Consequences, incentives and equity harms

25

Curriculum narrowing and teaching to the measure

Curriculum narrowing and teaching to the measure defines the classification question for the curricular incentive. For curriculum narrowing and teaching to the measure, the relevant population or units are learners entitled to a broad curriculum, and the direct evidence concerns tested and untested learning. In examining curriculum narrowing and teaching to the measure, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on curriculum narrowing and teaching to the measure, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of curriculum narrowing and teaching to the measure, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-01] [REF-02] [REF-12] [REF-17]

The principal misclassification risk in curriculum narrowing and teaching to the measure is that high stakes shift time toward tested content and away from practical, civic or creative learning. In examining curriculum narrowing and teaching to the measure, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on curriculum narrowing and teaching to the measure, authorities should identify how learners, events and institutions enter or leave the curricular incentive, whether missingness is concentrated, and which assumptions drive the result. For classification of curriculum narrowing and teaching to the measure, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-01] [REF-02]

The minimum measurement response for curriculum narrowing and teaching to the measure is to monitor breadth, time and learner work and limit stakes of narrow measures. Within evidence on curriculum narrowing and teaching to the measure, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of curriculum narrowing and teaching to the measure, it should also state the decision rule, consequence and evidence required for review. In interpreting curriculum narrowing and teaching to the measure, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-02] [REF-12]

Source fitness for the curricular incentive depends on coverage. For classification of curriculum narrowing and teaching to the measure, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting curriculum narrowing and teaching to the measure, for curriculum narrowing and teaching to the measure, each source should retain its principal error. For decisions about curriculum narrowing and teaching to the measure, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-12] [REF-17]

Equity is part of curriculum narrowing and teaching to the measure. In interpreting curriculum narrowing and teaching to the measure, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about curriculum narrowing and teaching to the measure, group levels and population shares should remain visible beside any school average. For curriculum narrowing and teaching to the measure, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining curriculum narrowing and teaching to the measure, if part of learners entitled to a broad curriculum is absent, the likely effect on tested and untested learning and classification should be reported.[REF-01] [REF-17]

Interpretation of tested and untested learning should separate observation, explanation and attribution. For decisions about curriculum narrowing and teaching to the measure, for curriculum narrowing and teaching to the measure, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For curriculum narrowing and teaching to the measure, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining curriculum narrowing and teaching to the measure, consequences should be proportionate to evidence strength. Within evidence on curriculum narrowing and teaching to the measure, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-02] [REF-12] [REF-17]

Accountability completes curriculum narrowing and teaching to the measure. For curriculum narrowing and teaching to the measure, a material finding about the curricular incentive should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining curriculum narrowing and teaching to the measure, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on curriculum narrowing and teaching to the measure, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-01] [REF-02] [REF-12] [REF-17]

26

Selective admission, exclusion and transfer

Selective admission, exclusion and transfer defines the classification question for the selection incentive. For selective admission, exclusion and transfer, the relevant population or units are applicants and current learners who may lower results, and the direct evidence concerns entry, exclusion and cohort change. In examining selective admission, exclusion and transfer, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on selective admission, exclusion and transfer, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of selective admission, exclusion and transfer, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-04] [REF-10] [REF-14] [REF-16]

The principal misclassification risk in selective admission, exclusion and transfer is that schools protect classifications by discouraging admission or moving learners before measurement. In examining selective admission, exclusion and transfer, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on selective admission, exclusion and transfer, authorities should identify how learners, events and institutions enter or leave the selection incentive, whether missingness is concentrated, and which assumptions drive the result. For classification of selective admission, exclusion and transfer, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-04] [REF-10]

The minimum measurement response for selective admission, exclusion and transfer is to review admission, transfer, exclusion and missing cohorts alongside outcomes. Within evidence on selective admission, exclusion and transfer, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of selective admission, exclusion and transfer, it should also state the decision rule, consequence and evidence required for review. In interpreting selective admission, exclusion and transfer, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-10] [REF-14]

Source fitness for the selection incentive depends on coverage. For classification of selective admission, exclusion and transfer, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting selective admission, exclusion and transfer, for selective admission, exclusion and transfer, each source should retain its principal error. For decisions about selective admission, exclusion and transfer, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-14] [REF-16]

Equity is part of selective admission, exclusion and transfer. In interpreting selective admission, exclusion and transfer, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about selective admission, exclusion and transfer, group levels and population shares should remain visible beside any school average. For selective admission, exclusion and transfer, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining selective admission, exclusion and transfer, if part of applicants and current learners who may lower results is absent, the likely effect on entry, exclusion and cohort change and classification should be reported.[REF-04] [REF-16]

Interpretation of entry, exclusion and cohort change should separate observation, explanation and attribution. For decisions about selective admission, exclusion and transfer, for selective admission, exclusion and transfer, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For selective admission, exclusion and transfer, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining selective admission, exclusion and transfer, consequences should be proportionate to evidence strength. Within evidence on selective admission, exclusion and transfer, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-10] [REF-14] [REF-16]

Accountability completes selective admission, exclusion and transfer. For selective admission, exclusion and transfer, a material finding about the selection incentive should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining selective admission, exclusion and transfer, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on selective admission, exclusion and transfer, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-04] [REF-10] [REF-14] [REF-16]

27

Record manipulation and metric substitution

Record manipulation and metric substitution defines the classification question for the recording incentive. For record manipulation and metric substitution, the relevant population or units are staff reporting events tied to classification, and the direct evidence concerns data authenticity and educational purpose. In examining record manipulation and metric substitution, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on record manipulation and metric substitution, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of record manipulation and metric substitution, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-09] [REF-19] [REF-20] [REF-21]

The principal misclassification risk in record manipulation and metric substitution is that reported compliance replaces real service or event coding changes to improve the indicator. In examining record manipulation and metric substitution, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on record manipulation and metric substitution, authorities should identify how learners, events and institutions enter or leave the recording incentive, whether missingness is concentrated, and which assumptions drive the result. For classification of record manipulation and metric substitution, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-09] [REF-19]

The minimum measurement response for record manipulation and metric substitution is to validate records, compare sources and preserve correction without punitive concealment. Within evidence on record manipulation and metric substitution, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of record manipulation and metric substitution, it should also state the decision rule, consequence and evidence required for review. In interpreting record manipulation and metric substitution, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-19] [REF-20]

Source fitness for the recording incentive depends on coverage. For classification of record manipulation and metric substitution, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting record manipulation and metric substitution, for record manipulation and metric substitution, each source should retain its principal error. For decisions about record manipulation and metric substitution, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-20] [REF-21]

Equity is part of record manipulation and metric substitution. In interpreting record manipulation and metric substitution, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about record manipulation and metric substitution, group levels and population shares should remain visible beside any school average. For record manipulation and metric substitution, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining record manipulation and metric substitution, if part of staff reporting events tied to classification is absent, the likely effect on data authenticity and educational purpose and classification should be reported.[REF-09] [REF-21]

Interpretation of data authenticity and educational purpose should separate observation, explanation and attribution. For decisions about record manipulation and metric substitution, for record manipulation and metric substitution, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For record manipulation and metric substitution, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining record manipulation and metric substitution, consequences should be proportionate to evidence strength. Within evidence on record manipulation and metric substitution, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-19] [REF-20] [REF-21]

Accountability completes record manipulation and metric substitution. For record manipulation and metric substitution, a material finding about the recording incentive should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining record manipulation and metric substitution, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on record manipulation and metric substitution, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-09] [REF-19] [REF-20] [REF-21]

28

Stigma and reputational feedback

Stigma and reputational feedback defines the classification question for the classification stigma. For stigma and reputational feedback, the relevant population or units are learners, staff and communities associated with a school label, and the direct evidence concerns public perception and participation. In examining stigma and reputational feedback, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on stigma and reputational feedback, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of stigma and reputational feedback, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-10] [REF-13] [REF-16] [REF-18]

The principal misclassification risk in stigma and reputational feedback is that a low category accelerates staff loss, avoidance and resource decline. In examining stigma and reputational feedback, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on stigma and reputational feedback, authorities should identify how learners, events and institutions enter or leave the classification stigma, whether missingness is concentrated, and which assumptions drive the result. For classification of stigma and reputational feedback, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-10] [REF-13]

The minimum measurement response for stigma and reputational feedback is to use careful language, uncertainty, support and a visible improvement route. Within evidence on stigma and reputational feedback, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of stigma and reputational feedback, it should also state the decision rule, consequence and evidence required for review. In interpreting stigma and reputational feedback, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-13] [REF-16]

Source fitness for the classification stigma depends on coverage. For classification of stigma and reputational feedback, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting stigma and reputational feedback, for stigma and reputational feedback, each source should retain its principal error. For decisions about stigma and reputational feedback, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-16] [REF-18]

Equity is part of stigma and reputational feedback. In interpreting stigma and reputational feedback, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about stigma and reputational feedback, group levels and population shares should remain visible beside any school average. For stigma and reputational feedback, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining stigma and reputational feedback, if part of learners, staff and communities associated with a school label is absent, the likely effect on public perception and participation and classification should be reported.[REF-10] [REF-18]

Interpretation of public perception and participation should separate observation, explanation and attribution. For decisions about stigma and reputational feedback, for stigma and reputational feedback, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For stigma and reputational feedback, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining stigma and reputational feedback, consequences should be proportionate to evidence strength. Within evidence on stigma and reputational feedback, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-13] [REF-16] [REF-18]

Accountability completes stigma and reputational feedback. For stigma and reputational feedback, a material finding about the classification stigma should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining stigma and reputational feedback, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on stigma and reputational feedback, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-10] [REF-13] [REF-16] [REF-18]

29

Unequal sanction and system responsibility

Unequal sanction and system responsibility defines the classification question for the accountability allocation. For unequal sanction and system responsibility, the relevant population or units are schools serving disadvantaged learners under system constraints, and the direct evidence concerns school and authority contribution. In examining unequal sanction and system responsibility, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on unequal sanction and system responsibility, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of unequal sanction and system responsibility, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-03] [REF-04] [REF-10] [REF-24]

The principal misclassification risk in unequal sanction and system responsibility is that school sanctions substitute for correcting finance, staffing or access failures controlled elsewhere. In examining unequal sanction and system responsibility, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on unequal sanction and system responsibility, authorities should identify how learners, events and institutions enter or leave the accountability allocation, whether missingness is concentrated, and which assumptions drive the result. For classification of unequal sanction and system responsibility, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-03] [REF-04]

The minimum measurement response for unequal sanction and system responsibility is to separate institutional practice from system conditions and assign each owner. Within evidence on unequal sanction and system responsibility, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of unequal sanction and system responsibility, it should also state the decision rule, consequence and evidence required for review. In interpreting unequal sanction and system responsibility, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-04] [REF-10]

Source fitness for the accountability allocation depends on coverage. For classification of unequal sanction and system responsibility, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting unequal sanction and system responsibility, for unequal sanction and system responsibility, each source should retain its principal error. For decisions about unequal sanction and system responsibility, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-10] [REF-24]

Equity is part of unequal sanction and system responsibility. In interpreting unequal sanction and system responsibility, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about unequal sanction and system responsibility, group levels and population shares should remain visible beside any school average. For unequal sanction and system responsibility, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining unequal sanction and system responsibility, if part of schools serving disadvantaged learners under system constraints is absent, the likely effect on school and authority contribution and classification should be reported.[REF-03] [REF-24]

Interpretation of school and authority contribution should separate observation, explanation and attribution. For decisions about unequal sanction and system responsibility, for unequal sanction and system responsibility, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For unequal sanction and system responsibility, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining unequal sanction and system responsibility, consequences should be proportionate to evidence strength. Within evidence on unequal sanction and system responsibility, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-04] [REF-10] [REF-24]

Accountability completes unequal sanction and system responsibility. For unequal sanction and system responsibility, a material finding about the accountability allocation should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining unequal sanction and system responsibility, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on unequal sanction and system responsibility, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-03] [REF-04] [REF-10] [REF-24]

30

Improvement support versus punitive ranking

Improvement support versus punitive ranking defines the classification question for the classification response. For improvement support versus punitive ranking, the relevant population or units are schools identified with material needs, and the direct evidence concerns diagnosis, assistance and later evidence. In examining improvement support versus punitive ranking, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on improvement support versus punitive ranking, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of improvement support versus punitive ranking, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-01] [REF-16] [REF-18] [REF-24]

The principal misclassification risk in improvement support versus punitive ranking is that ranking generates publicity without targeted support or verified change. In examining improvement support versus punitive ranking, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on improvement support versus punitive ranking, authorities should identify how learners, events and institutions enter or leave the classification response, whether missingness is concentrated, and which assumptions drive the result. For classification of improvement support versus punitive ranking, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-01] [REF-16]

The minimum measurement response for improvement support versus punitive ranking is to link classification to diagnostic review, resources, milestones and proportionate escalation. Within evidence on improvement support versus punitive ranking, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of improvement support versus punitive ranking, it should also state the decision rule, consequence and evidence required for review. In interpreting improvement support versus punitive ranking, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-16] [REF-18]

Source fitness for the classification response depends on coverage. For classification of improvement support versus punitive ranking, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting improvement support versus punitive ranking, for improvement support versus punitive ranking, each source should retain its principal error. For decisions about improvement support versus punitive ranking, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-18] [REF-24]

Equity is part of improvement support versus punitive ranking. In interpreting improvement support versus punitive ranking, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about improvement support versus punitive ranking, group levels and population shares should remain visible beside any school average. For improvement support versus punitive ranking, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining improvement support versus punitive ranking, if part of schools identified with material needs is absent, the likely effect on diagnosis, assistance and later evidence and classification should be reported.[REF-01] [REF-24]

Interpretation of diagnosis, assistance and later evidence should separate observation, explanation and attribution. For decisions about improvement support versus punitive ranking, for improvement support versus punitive ranking, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For improvement support versus punitive ranking, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining improvement support versus punitive ranking, consequences should be proportionate to evidence strength. Within evidence on improvement support versus punitive ranking, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-16] [REF-18] [REF-24]

Accountability completes improvement support versus punitive ranking. For improvement support versus punitive ranking, a material finding about the classification response should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining improvement support versus punitive ranking, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on improvement support versus punitive ranking, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-01] [REF-16] [REF-18] [REF-24]

Part VI

Responsible classification, reporting and remedy

31

A minimum school performance evidence set

A minimum school performance evidence set defines the classification question for the evidence set. For a minimum school performance evidence set, the relevant population or units are schools and communities needing a balanced account, and the direct evidence concerns levels, progress, participation, pathways and conditions. In examining a minimum school performance evidence set, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on a minimum school performance evidence set, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of a minimum school performance evidence set, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-01] [REF-02] [REF-16] [REF-18]

The principal misclassification risk in a minimum school performance evidence set is that selective indicators create an incomplete but authoritative label. In examining a minimum school performance evidence set, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on a minimum school performance evidence set, authorities should identify how learners, events and institutions enter or leave the evidence set, whether missingness is concentrated, and which assumptions drive the result. For classification of a minimum school performance evidence set, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-01] [REF-02]

The minimum measurement response for a minimum school performance evidence set is to publish a compact stable set with components, distributions and coverage. Within evidence on a minimum school performance evidence set, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of a minimum school performance evidence set, it should also state the decision rule, consequence and evidence required for review. In interpreting a minimum school performance evidence set, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-02] [REF-16]

Source fitness for the evidence set depends on coverage. For classification of a minimum school performance evidence set, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting a minimum school performance evidence set, for a minimum school performance evidence set, each source should retain its principal error. For decisions about a minimum school performance evidence set, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-16] [REF-18]

Equity is part of a minimum school performance evidence set. In interpreting a minimum school performance evidence set, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about a minimum school performance evidence set, group levels and population shares should remain visible beside any school average. For a minimum school performance evidence set, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining a minimum school performance evidence set, if part of schools and communities needing a balanced account is absent, the likely effect on levels, progress, participation, pathways and conditions and classification should be reported.[REF-01] [REF-18]

Interpretation of levels, progress, participation, pathways and conditions should separate observation, explanation and attribution. For decisions about a minimum school performance evidence set, for a minimum school performance evidence set, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For a minimum school performance evidence set, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining a minimum school performance evidence set, consequences should be proportionate to evidence strength. Within evidence on a minimum school performance evidence set, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-02] [REF-16] [REF-18]

Accountability completes a minimum school performance evidence set. For a minimum school performance evidence set, a material finding about the evidence set should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining a minimum school performance evidence set, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on a minimum school performance evidence set, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-01] [REF-02] [REF-16] [REF-18]

32

Pre-stated decision and classification rules

Pre-stated decision and classification rules defines the classification question for the classification rule. For pre-stated decision and classification rules, the relevant population or units are schools subject to categories and consequences, and the direct evidence concerns threshold, evidence and response. In examining pre-stated decision and classification rules, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on pre-stated decision and classification rules, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of pre-stated decision and classification rules, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-09] [REF-16] [REF-19] [REF-20]

The principal misclassification risk in pre-stated decision and classification rules is that rules change after results or hidden judgement alters category. In examining pre-stated decision and classification rules, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on pre-stated decision and classification rules, authorities should identify how learners, events and institutions enter or leave the classification rule, whether missingness is concentrated, and which assumptions drive the result. For classification of pre-stated decision and classification rules, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-09] [REF-16]

The minimum measurement response for pre-stated decision and classification rules is to publish formula, threshold, uncertainty band, corroboration and consequence in advance. Within evidence on pre-stated decision and classification rules, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of pre-stated decision and classification rules, it should also state the decision rule, consequence and evidence required for review. In interpreting pre-stated decision and classification rules, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-16] [REF-19]

Source fitness for the classification rule depends on coverage. For classification of pre-stated decision and classification rules, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting pre-stated decision and classification rules, for pre-stated decision and classification rules, each source should retain its principal error. For decisions about pre-stated decision and classification rules, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-19] [REF-20]

Equity is part of pre-stated decision and classification rules. In interpreting pre-stated decision and classification rules, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about pre-stated decision and classification rules, group levels and population shares should remain visible beside any school average. For pre-stated decision and classification rules, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining pre-stated decision and classification rules, if part of schools subject to categories and consequences is absent, the likely effect on threshold, evidence and response and classification should be reported.[REF-09] [REF-20]

Interpretation of threshold, evidence and response should separate observation, explanation and attribution. For decisions about pre-stated decision and classification rules, for pre-stated decision and classification rules, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For pre-stated decision and classification rules, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining pre-stated decision and classification rules, consequences should be proportionate to evidence strength. Within evidence on pre-stated decision and classification rules, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-16] [REF-19] [REF-20]

Accountability completes pre-stated decision and classification rules. For pre-stated decision and classification rules, a material finding about the classification rule should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining pre-stated decision and classification rules, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on pre-stated decision and classification rules, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-09] [REF-16] [REF-19] [REF-20]

33

Independent technical and contextual review

Independent technical and contextual review defines the classification question for the classification review. For independent technical and contextual review, the relevant population or units are schools near thresholds or facing high stakes, and the direct evidence concerns verification and contextual evidence. In examining independent technical and contextual review, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on independent technical and contextual review, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of independent technical and contextual review, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-09] [REF-10] [REF-16] [REF-20]

The principal misclassification risk in independent technical and contextual review is that automated categorisation ignores data error, exceptional events or source limitations. In examining independent technical and contextual review, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on independent technical and contextual review, authorities should identify how learners, events and institutions enter or leave the classification review, whether missingness is concentrated, and which assumptions drive the result. For classification of independent technical and contextual review, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-09] [REF-10]

The minimum measurement response for independent technical and contextual review is to require bounded independent review without permitting evidence-free discretion. Within evidence on independent technical and contextual review, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of independent technical and contextual review, it should also state the decision rule, consequence and evidence required for review. In interpreting independent technical and contextual review, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-10] [REF-16]

Source fitness for the classification review depends on coverage. For classification of independent technical and contextual review, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting independent technical and contextual review, for independent technical and contextual review, each source should retain its principal error. For decisions about independent technical and contextual review, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-16] [REF-20]

Equity is part of independent technical and contextual review. In interpreting independent technical and contextual review, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about independent technical and contextual review, group levels and population shares should remain visible beside any school average. For independent technical and contextual review, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining independent technical and contextual review, if part of schools near thresholds or facing high stakes is absent, the likely effect on verification and contextual evidence and classification should be reported.[REF-09] [REF-20]

Interpretation of verification and contextual evidence should separate observation, explanation and attribution. For decisions about independent technical and contextual review, for independent technical and contextual review, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For independent technical and contextual review, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining independent technical and contextual review, consequences should be proportionate to evidence strength. Within evidence on independent technical and contextual review, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-10] [REF-16] [REF-20]

Accountability completes independent technical and contextual review. For independent technical and contextual review, a material finding about the classification review should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining independent technical and contextual review, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on independent technical and contextual review, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-09] [REF-10] [REF-16] [REF-20]

34

Public reporting without league-table overclaim

Public reporting without league-table overclaim defines the classification question for the public school report. For public reporting without league-table overclaim, the relevant population or units are families, learners and communities using performance evidence, and the direct evidence concerns finding, uncertainty and explanation. In examining public reporting without league-table overclaim, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on public reporting without league-table overclaim, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of public reporting without league-table overclaim, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-09] [REF-16] [REF-18] [REF-19]

The principal misclassification risk in public reporting without league-table overclaim is that ordinal ranks imply precise differences and one-year causation. In examining public reporting without league-table overclaim, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on public reporting without league-table overclaim, authorities should identify how learners, events and institutions enter or leave the public school report, whether missingness is concentrated, and which assumptions drive the result. For classification of public reporting without league-table overclaim, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-09] [REF-16]

The minimum measurement response for public reporting without league-table overclaim is to lead with levels and distributions and explain uncertainty, context and limits. Within evidence on public reporting without league-table overclaim, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of public reporting without league-table overclaim, it should also state the decision rule, consequence and evidence required for review. In interpreting public reporting without league-table overclaim, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-16] [REF-18]

Source fitness for the public school report depends on coverage. For classification of public reporting without league-table overclaim, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting public reporting without league-table overclaim, for public reporting without league-table overclaim, each source should retain its principal error. For decisions about public reporting without league-table overclaim, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-18] [REF-19]

Equity is part of public reporting without league-table overclaim. In interpreting public reporting without league-table overclaim, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about public reporting without league-table overclaim, group levels and population shares should remain visible beside any school average. For public reporting without league-table overclaim, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining public reporting without league-table overclaim, if part of families, learners and communities using performance evidence is absent, the likely effect on finding, uncertainty and explanation and classification should be reported.[REF-09] [REF-19]

Interpretation of finding, uncertainty and explanation should separate observation, explanation and attribution. For decisions about public reporting without league-table overclaim, for public reporting without league-table overclaim, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For public reporting without league-table overclaim, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining public reporting without league-table overclaim, consequences should be proportionate to evidence strength. Within evidence on public reporting without league-table overclaim, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-16] [REF-18] [REF-19]

Accountability completes public reporting without league-table overclaim. For public reporting without league-table overclaim, a material finding about the public school report should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining public reporting without league-table overclaim, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on public reporting without league-table overclaim, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-09] [REF-16] [REF-18] [REF-19]

35

Appeal, correction and continuity of support

Appeal, correction and continuity of support defines the classification question for the classification remedy. For appeal, correction and continuity of support, the relevant population or units are schools and learners affected by error or sanction, and the direct evidence concerns challenge and corrected decision. In examining appeal, correction and continuity of support, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on appeal, correction and continuity of support, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of appeal, correction and continuity of support, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-10] [REF-13] [REF-16] [REF-20]

The principal misclassification risk in appeal, correction and continuity of support is that error correction is slow or support is withdrawn during dispute. In examining appeal, correction and continuity of support, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on appeal, correction and continuity of support, authorities should identify how learners, events and institutions enter or leave the classification remedy, whether missingness is concentrated, and which assumptions drive the result. For classification of appeal, correction and continuity of support, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-10] [REF-13]

The minimum measurement response for appeal, correction and continuity of support is to provide time-bound appeal, data correction and continuing learner protection. Within evidence on appeal, correction and continuity of support, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of appeal, correction and continuity of support, it should also state the decision rule, consequence and evidence required for review. In interpreting appeal, correction and continuity of support, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-13] [REF-16]

Source fitness for the classification remedy depends on coverage. For classification of appeal, correction and continuity of support, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting appeal, correction and continuity of support, for appeal, correction and continuity of support, each source should retain its principal error. For decisions about appeal, correction and continuity of support, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-16] [REF-20]

Equity is part of appeal, correction and continuity of support. In interpreting appeal, correction and continuity of support, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about appeal, correction and continuity of support, group levels and population shares should remain visible beside any school average. For appeal, correction and continuity of support, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining appeal, correction and continuity of support, if part of schools and learners affected by error or sanction is absent, the likely effect on challenge and corrected decision and classification should be reported.[REF-10] [REF-20]

Interpretation of challenge and corrected decision should separate observation, explanation and attribution. For decisions about appeal, correction and continuity of support, for appeal, correction and continuity of support, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For appeal, correction and continuity of support, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining appeal, correction and continuity of support, consequences should be proportionate to evidence strength. Within evidence on appeal, correction and continuity of support, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-13] [REF-16] [REF-20]

Accountability completes appeal, correction and continuity of support. For appeal, correction and continuity of support, a material finding about the classification remedy should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining appeal, correction and continuity of support, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on appeal, correction and continuity of support, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-10] [REF-13] [REF-16] [REF-20]

36

A proportionate improvement classification

A proportionate improvement classification defines the classification question for the improvement framework. For a proportionate improvement classification, the relevant population or units are schools and system bodies sharing responsibility, and the direct evidence concerns need, support, milestone and escalation. In examining a proportionate improvement classification, a school result should be interpreted only after the decision purpose, unit, learner coverage and reference period are fixed. Within evidence on a proportionate improvement classification, the same evidence can be useful for internal enquiry yet insufficient for a public category or sanction. For classification of a proportionate improvement classification, the report should state what underlying condition is being inferred and what other dimensions of quality remain outside the measure.[REF-01] [REF-10] [REF-16] [REF-24]

The principal misclassification risk in a proportionate improvement classification is that a permanent label replaces a dated diagnosis and improvement plan. In examining a proportionate improvement classification, this can place schools into different categories even though the relevant underlying condition is similar, or conceal a material need behind a favourable summary. Within evidence on a proportionate improvement classification, authorities should identify how learners, events and institutions enter or leave the improvement framework, whether missingness is concentrated, and which assumptions drive the result. For classification of a proportionate improvement classification, an absent value should not become a successful zero, and a precise point estimate should not erase uncertainty.[REF-01] [REF-10]

The minimum measurement response for a proportionate improvement classification is to classify the material need, assign owners and resources, retest and retire the label. Within evidence on a proportionate improvement classification, the specification should include numerator, denominator, cohort or population, source, period, coverage, reliability, uncertainty and group detail. For classification of a proportionate improvement classification, it should also state the decision rule, consequence and evidence required for review. In interpreting a proportionate improvement classification, where adjustment or aggregation occurs, the conceptual reason and influence of each component should remain visible so users can reproduce the category.[REF-10] [REF-16]

Source fitness for the improvement framework depends on coverage. For classification of a proportionate improvement classification, administrative records can describe enrolled learners and events, assessments can describe achievement for a stated target population, surveys can describe climate or circumstance, and observation can examine teaching and conditions. In interpreting a proportionate improvement classification, for a proportionate improvement classification, each source should retain its principal error. For decisions about a proportionate improvement classification, divergence should be investigated before values are merged, because it may reveal different populations, time periods or constructs rather than poor data alone.[REF-16] [REF-24]

Equity is part of a proportionate improvement classification. In interpreting a proportionate improvement classification, the distribution may differ by prior attainment, poverty, location, disability, language, migration or another material characteristic. For decisions about a proportionate improvement classification, group levels and population shares should remain visible beside any school average. For a proportionate improvement classification, context should guide diagnosis and responsibility, not make disadvantaged learners disappear from the expected standard. In examining a proportionate improvement classification, if part of schools and system bodies sharing responsibility is absent, the likely effect on need, support, milestone and escalation and classification should be reported.[REF-01] [REF-24]

Interpretation of need, support, milestone and escalation should separate observation, explanation and attribution. For decisions about a proportionate improvement classification, for a proportionate improvement classification, an outcome difference does not by itself establish school contribution, and adjustment does not prove a causal effect. For a proportionate improvement classification, reports should test plausible alternatives, model sensitivity and the stability of any category. In examining a proportionate improvement classification, consequences should be proportionate to evidence strength. Within evidence on a proportionate improvement classification, all claims should be dated so later methods or results are not read backwards into the 1 April 2017 decision.[REF-10] [REF-16] [REF-24]

Accountability completes a proportionate improvement classification. For a proportionate improvement classification, a material finding about the improvement framework should lead to a bounded diagnostic review, an assigned school or system owner, resources, a milestone and a later learner-facing test. In examining a proportionate improvement classification, schools and communities need access to data correction and appeal, while learners need continuing support during dispute. Within evidence on a proportionate improvement classification, the category should be revised or retired when evidence changes rather than become a permanent reputational label.[REF-01] [REF-10] [REF-16] [REF-24]

A stability record for defining the school as the reporting unit should preserve the definition, population, source coverage, model, threshold and uncertainty used for the school unit. Within evidence on defining the school as the reporting unit, it should show how the category changes under plausible alternatives and whether unit boundary is persistent. For classification of defining the school as the reporting unit, this record prevents a revised method from being presented as changed school quality.[REF-01] [REF-05] [REF-06] [REF-16]

A consequence test for defining the school as the reporting unit should ask how the classification affects admission, curriculum, staffing, resources and community behaviour. Within evidence on defining the school as the reporting unit, if the label creates incentives to exclude institutions, campuses, programmes and shifts grouped as one school or narrows educational provision, the authority should reduce stakes and strengthen corroborating evidence. For classification of defining the school as the reporting unit, the test should protect learner entitlement while error is reviewed.[REF-01] [REF-05]

A stability record for separating description, diagnosis and judgement should preserve the definition, population, source coverage, model, threshold and uncertainty used for the measure purpose. Within evidence on separating description, diagnosis and judgement, it should show how the category changes under plausible alternatives and whether descriptive, diagnostic and evaluative use is persistent. For classification of separating description, diagnosis and judgement, this record prevents a revised method from being presented as changed school quality.[REF-01] [REF-09] [REF-10] [REF-16]

A consequence test for separating description, diagnosis and judgement should ask how the classification affects admission, curriculum, staffing, resources and community behaviour. Within evidence on separating description, diagnosis and judgement, if the label creates incentives to exclude schools and users receiving different forms of evidence or narrows educational provision, the authority should reduce stakes and strengthen corroborating evidence. For classification of separating description, diagnosis and judgement, the test should protect learner entitlement while error is reviewed.[REF-01] [REF-09]

A stability record for defining performance beyond a single outcome should preserve the definition, population, source coverage, model, threshold and uncertainty used for the performance construct. Within evidence on defining performance beyond a single outcome, it should show how the category changes under plausible alternatives and whether access, learning, wellbeing and progression is persistent. For classification of defining performance beyond a single outcome, this record prevents a revised method from being presented as changed school quality.[REF-01] [REF-02] [REF-12] [REF-13]

A consequence test for defining performance beyond a single outcome should ask how the classification affects admission, curriculum, staffing, resources and community behaviour. Within evidence on defining performance beyond a single outcome, if the label creates incentives to exclude learners and schools with multiple education duties or narrows educational provision, the authority should reduce stakes and strengthen corroborating evidence. For classification of defining performance beyond a single outcome, the test should protect learner entitlement while error is reviewed.[REF-01] [REF-02]

A stability record for reference period and cohort alignment should preserve the definition, population, source coverage, model, threshold and uncertainty used for the performance period. Within evidence on reference period and cohort alignment, it should show how the category changes under plausible alternatives and whether annual or cohort result is persistent. For classification of reference period and cohort alignment, this record prevents a revised method from being presented as changed school quality.[REF-05] [REF-06] [REF-17] [REF-19]

A consequence test for reference period and cohort alignment should ask how the classification affects admission, curriculum, staffing, resources and community behaviour. Within evidence on reference period and cohort alignment, if the label creates incentives to exclude learners observed across school years and stages or narrows educational provision, the authority should reduce stakes and strengthen corroborating evidence. For classification of reference period and cohort alignment, the test should protect learner entitlement while error is reviewed.[REF-05] [REF-06]

A stability record for minimum standards and relative ranks should preserve the definition, population, source coverage, model, threshold and uncertainty used for the classification threshold. Within evidence on minimum standards and relative ranks, it should show how the category changes under plausible alternatives and whether absolute level and relative position is persistent. For classification of minimum standards and relative ranks, this record prevents a revised method from being presented as changed school quality.[REF-02] [REF-10] [REF-18] [REF-21]

A consequence test for minimum standards and relative ranks should ask how the classification affects admission, curriculum, staffing, resources and community behaviour. Within evidence on minimum standards and relative ranks, if the label creates incentives to exclude schools compared with standards and one another or narrows educational provision, the authority should reduce stakes and strengthen corroborating evidence. For classification of minimum standards and relative ranks, the test should protect learner entitlement while error is reviewed.[REF-02] [REF-10]

A stability record for high-stakes use and proportional evidence should preserve the definition, population, source coverage, model, threshold and uncertainty used for the decision proportionality. Within evidence on high-stakes use and proportional evidence, it should show how the category changes under plausible alternatives and whether evidence strength and consequence is persistent. For classification of high-stakes use and proportional evidence, this record prevents a revised method from being presented as changed school quality.[REF-09] [REF-10] [REF-16] [REF-20]

A consequence test for high-stakes use and proportional evidence should ask how the classification affects admission, curriculum, staffing, resources and community behaviour. Within evidence on high-stakes use and proportional evidence, if the label creates incentives to exclude schools, staff and learners affected by classification or narrows educational provision, the authority should reduce stakes and strengthen corroborating evidence. For classification of high-stakes use and proportional evidence, the test should protect learner entitlement while error is reviewed.[REF-09] [REF-10]

A stability record for achievement levels and score distributions should preserve the definition, population, source coverage, model, threshold and uncertainty used for the achievement measure. Within evidence on achievement levels and score distributions, it should show how the category changes under plausible alternatives and whether score level, distribution and threshold is persistent. For classification of achievement levels and score distributions, this record prevents a revised method from being presented as changed school quality.[REF-02] [REF-05] [REF-17] [REF-18]

A consequence test for achievement levels and score distributions should ask how the classification affects admission, curriculum, staffing, resources and community behaviour. Within evidence on achievement levels and score distributions, if the label creates incentives to exclude eligible assessed learners and those excluded or absent or narrows educational provision, the authority should reduce stakes and strengthen corroborating evidence. For classification of achievement levels and score distributions, the test should protect learner entitlement while error is reviewed.[REF-02] [REF-05]

A stability record for learner progress and value-added claims should preserve the definition, population, source coverage, model, threshold and uncertainty used for the progress measure. Within evidence on learner progress and value-added claims, it should show how the category changes under plausible alternatives and whether change conditional on prior attainment is persistent. For classification of learner progress and value-added claims, this record prevents a revised method from being presented as changed school quality.[REF-01] [REF-08] [REF-17] [REF-21]

A consequence test for learner progress and value-added claims should ask how the classification affects admission, curriculum, staffing, resources and community behaviour. Within evidence on learner progress and value-added claims, if the label creates incentives to exclude cohorts with comparable prior and later evidence or narrows educational provision, the authority should reduce stakes and strengthen corroborating evidence. For classification of learner progress and value-added claims, the test should protect learner entitlement while error is reviewed.[REF-01] [REF-08]

A stability record for attendance, exclusion and participation should preserve the definition, population, source coverage, model, threshold and uncertainty used for the participation measure. Within evidence on attendance, exclusion and participation, it should show how the category changes under plausible alternatives and whether attendance intensity and exclusion is persistent. For classification of attendance, exclusion and participation, this record prevents a revised method from being presented as changed school quality.[REF-04] [REF-11] [REF-14] [REF-18]

A consequence test for attendance, exclusion and participation should ask how the classification affects admission, curriculum, staffing, resources and community behaviour. Within evidence on attendance, exclusion and participation, if the label creates incentives to exclude enrolled and eligible learners using the school or narrows educational provision, the authority should reduce stakes and strengthen corroborating evidence. For classification of attendance, exclusion and participation, the test should protect learner entitlement while error is reviewed.[REF-04] [REF-11]

A stability record for completion, retention and transition should preserve the definition, population, source coverage, model, threshold and uncertainty used for the pathway measure. Within evidence on completion, retention and transition, it should show how the category changes under plausible alternatives and whether retention, completion and destination is persistent. For classification of completion, retention and transition, this record prevents a revised method from being presented as changed school quality.[REF-05] [REF-06] [REF-18] [REF-24]

A consequence test for completion, retention and transition should ask how the classification affects admission, curriculum, staffing, resources and community behaviour. Within evidence on completion, retention and transition, if the label creates incentives to exclude cohorts approaching programme completion or narrows educational provision, the authority should reduce stakes and strengthen corroborating evidence. For classification of completion, retention and transition, the test should protect learner entitlement while error is reviewed.[REF-05] [REF-06]

A stability record for school climate, safety and learner voice should preserve the definition, population, source coverage, model, threshold and uncertainty used for the climate measure. Within evidence on school climate, safety and learner voice, it should show how the category changes under plausible alternatives and whether safety, belonging and participation is persistent. For classification of school climate, safety and learner voice, this record prevents a revised method from being presented as changed school quality.[REF-10] [REF-12] [REF-14] [REF-15]

A consequence test for school climate, safety and learner voice should ask how the classification affects admission, curriculum, staffing, resources and community behaviour. Within evidence on school climate, safety and learner voice, if the label creates incentives to exclude learners and staff experiencing school conditions or narrows educational provision, the authority should reduce stakes and strengthen corroborating evidence. For classification of school climate, safety and learner voice, the test should protect learner entitlement while error is reviewed.[REF-10] [REF-12]

A stability record for composite indices and hidden trade-offs should preserve the definition, population, source coverage, model, threshold and uncertainty used for the composite performance measure. Within evidence on composite indices and hidden trade-offs, it should show how the category changes under plausible alternatives and whether weighted aggregate and components is persistent. For classification of composite indices and hidden trade-offs, this record prevents a revised method from being presented as changed school quality.[REF-01] [REF-19] [REF-20] [REF-21]

A consequence test for composite indices and hidden trade-offs should ask how the classification affects admission, curriculum, staffing, resources and community behaviour. Within evidence on composite indices and hidden trade-offs, if the label creates incentives to exclude schools summarised across component indicators or narrows educational provision, the authority should reduce stakes and strengthen corroborating evidence. For classification of composite indices and hidden trade-offs, the test should protect learner entitlement while error is reviewed.[REF-01] [REF-19]

A stability record for prior attainment and starting points should preserve the definition, population, source coverage, model, threshold and uncertainty used for the intake baseline. Within evidence on prior attainment and starting points, it should show how the category changes under plausible alternatives and whether prior attainment and later result is persistent. For classification of prior attainment and starting points, this record prevents a revised method from being presented as changed school quality.[REF-02] [REF-04] [REF-17] [REF-21]

A consequence test for prior attainment and starting points should ask how the classification affects admission, curriculum, staffing, resources and community behaviour. Within evidence on prior attainment and starting points, if the label creates incentives to exclude learners entering with varied prior opportunity or narrows educational provision, the authority should reduce stakes and strengthen corroborating evidence. For classification of prior attainment and starting points, the test should protect learner entitlement while error is reviewed.[REF-02] [REF-04]

A stability record for poverty and household-resource composition should preserve the definition, population, source coverage, model, threshold and uncertainty used for the socioeconomic context. Within evidence on poverty and household-resource composition, it should show how the category changes under plausible alternatives and whether school intake and outcome gradient is persistent. For classification of poverty and household-resource composition, this record prevents a revised method from being presented as changed school quality.[REF-03] [REF-04] [REF-07] [REF-11]

A consequence test for poverty and household-resource composition should ask how the classification affects admission, curriculum, staffing, resources and community behaviour. Within evidence on poverty and household-resource composition, if the label creates incentives to exclude learners across household resources and material barriers or narrows educational provision, the authority should reduce stakes and strengthen corroborating evidence. For classification of poverty and household-resource composition, the test should protect learner entitlement while error is reviewed.[REF-03] [REF-04]

References

  1. REF-01

    United Nations Educational, Scientific and Cultural Organization. General Education Quality Analysis and Diagnosis Framework. 2012.

    Systemic analysis of education quality, inputs, teaching, learning and outcomes.

    https://unesdoc.unesco.org/ark:/48223/pf0000217520
  2. REF-02

    Education for All Global Monitoring Report Team. Teaching and Learning: Achieving Quality for All — EFA Global Monitoring Report 2013/4. 2014.

    Evidence on teaching, learning, inequality and education quality.

    https://unesdoc.unesco.org/ark:/48223/pf0000225660
  3. REF-03

    Education for All Global Monitoring Report Team. Overcoming Inequality: Why Governance Matters — EFA Global Monitoring Report 2009. 2008.

    Evidence on governance, inequality, finance and public accountability.

    https://unesdoc.unesco.org/ark:/48223/pf0000177683
  4. REF-04

    Education for All Global Monitoring Report Team. Reaching the Marginalized — EFA Global Monitoring Report 2010. 2010.

    Evidence on intersecting disadvantage and educational marginalisation.

    https://unesdoc.unesco.org/ark:/48223/pf0000186606
  5. REF-05

    UNESCO Institute for Statistics. Education Indicators: Technical Guidelines. 2009.

    Definitions, numerators, denominators and limitations for education indicators.

    https://uis.unesco.org/sites/default/files/documents/education-indicators-technical-guidelines-en_0.pdf
  6. REF-06

    United Nations Educational, Scientific and Cultural Organization. International Standard Classification of Education: ISCED 2011. 2012.

    Common definitions for education programmes and attainment.

    https://uis.unesco.org/sites/default/files/documents/international-standard-classification-of-education-isced-2011-en.pdf
  7. REF-07

    UNESCO Institute for Statistics. Guide to the Analysis and Use of Household Survey and Census Education Data. 2004.

    Methods and limits for household and census education indicators.

    https://uis.unesco.org/sites/default/files/documents/guide-to-the-analysis-and-use-of-household-survey-and-census-education-data-en_0.pdf
  8. REF-08

    United Nations Statistics Division. Household Sample Surveys in Developing and Transition Countries. 2005.

    Guidance on sampling, response, weighting and statistical error.

    https://unstats.un.org/unsd/hhsurveys/sectiona_new.htm
  9. REF-09

    United Nations General Assembly. Fundamental Principles of Official Statistics. 2014.

    Relevance, professional methods, transparency, correction and confidentiality.

    https://undocs.org/A/RES/68/261
  10. REF-10

    Office of the United Nations High Commissioner for Human Rights. Human Rights Indicators: A Guide to Measurement and Implementation. 2012.

    Rights-sensitive measurement, disaggregation and interpretation.

    https://www.ohchr.org/sites/default/files/Documents/Publications/Human_rights_indicators_en.pdf
  11. REF-11

    United Nations Children’s Fund. The State of the World’s Children 2014 in Numbers: Every Child Counts — Revealing Disparities, Advancing Children’s Rights. 2014.

    Evidence on disaggregation, unequal outcomes and statistical visibility.

    https://www.unicef.org/reports/state-worlds-children-2014
  12. REF-12

    United Nations Children’s Fund. Child Friendly Schools Manual. 2009.

    Guidance on inclusive, effective, protective and participatory schools.

    https://www.unicef.org/reports/child-friendly-schools-manual
  13. REF-13

    United Nations Educational, Scientific and Cultural Organization and United Nations Children’s Fund. A Human Rights-Based Approach to Education for All. 2007.

    Rights-based public duties for access, quality, participation and accountability.

    https://unesdoc.unesco.org/ark:/48223/pf0000154861
  14. REF-14

    United Nations General Assembly. Convention on the Rights of the Child. 1989.

    Education, non-discrimination, development, participation and protection obligations.

    https://www.ohchr.org/en/instruments-mechanisms/instruments/convention-rights-child
  15. REF-15

    United Nations General Assembly. Convention on the Rights of Persons with Disabilities. 2006.

    Inclusive education, accessibility and reasonable accommodation.

    https://www.ohchr.org/en/instruments-mechanisms/instruments/convention-rights-persons-disabilities
  16. REF-16

    European Commission/EACEA/Eurydice. Assuring Quality in Education: Policies and Approaches to School Evaluation in Europe. 2015.

    Comparative European evidence on external and internal school evaluation.

    https://op.europa.eu/en/publication-detail/-/publication/4a244ff8-7bac-11e5-9fae-01aa75ed71a1
  17. REF-17

    European Commission/EACEA/Eurydice. National Testing of Pupils in Europe: Objectives, Organisation and Use of Results. 2009.

    European evidence on test purposes, coverage and uses.

    https://op.europa.eu/en/publication-detail/-/publication/df628df4-4e5b-4014-adbd-2ed54a274fd9
  18. REF-18

    European Commission. Education and Training Monitor 2016. 2016.

    European evidence on attainment, early leaving, inequality and education conditions.

    https://op.europa.eu/en/publication-detail/-/publication/d7fd37b9-b130-11e6-871e-01aa75ed71a1
  19. REF-19

    European Statistical System Committee. European Statistics Code of Practice. 2011.

    Institutional and statistical principles for trustworthy public evidence.

    https://ec.europa.eu/eurostat/web/quality/european-quality-standards/european-statistics-code-of-practice
  20. REF-20

    European Parliament and Council of the European Union. Regulation (EC) No 223/2009 on European Statistics. 2009.

    European requirements for independence, quality, confidentiality and dissemination.

    https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32009R0223
  21. REF-21

    Organisation for Economic Co-operation and Development and European Commission Joint Research Centre. Handbook on Constructing Composite Indicators: Methodology and User Guide. 2008.

    Guidance on conceptual frameworks, normalisation, weighting, aggregation, sensitivity and presentation.

    https://www.oecd.org/sdd/42495745.pdf
  22. REF-22

    United Nations High Commissioner for Refugees. Missing Out: Refugee Education in Crisis. 2016.

    Evidence on refugee participation and barriers relevant to school intake and coverage.

    https://www.unhcr.org/media/missing-out-refugee-education-crisis
  23. REF-23

    United Nations General Assembly. New York Declaration for Refugees and Migrants. 2016.

    Commitments concerning refugee and migrant access to public services and education.

    https://undocs.org/A/RES/71/1
  24. REF-24

    Global Education Monitoring Report Team. Education for People and Planet: Creating Sustainable Futures for All — Global Education Monitoring Report 2016. 2016.

    Global monitoring evidence and accountability context in the Education 2030 era.

    https://unesdoc.unesco.org/ark:/48223/pf0000245752