Thematic Research Report

ICEQC-R-2016-07 — National Baselines for Sustainable Development Goal 4: Availability and Comparability

A global comparative indicator study of baseline fitness, source coverage, equity and provisional comparability

Publication date
Research category
Data and Indicator Research
Report archetype
Comparative Indicator Study
Geographic scope
Global
Evidence cut-off date
Responsible body
ICEQC Research and Policy Directorate
International Council for Education Quality Certification

ICEQC-R-2016-07

National Baselines for Sustainable Development Goal 4: Availability and Comparability

A global comparative indicator study of baseline fitness, source coverage, equity and provisional comparability

Publication date
Evidence cut-off date
Publication type
Thematic Research Report
Authoritative language
EN

Publication record

This is the controlled English edition. Evidence and institutional status are stated as at the evidence cut-off date.

Executive summary

National baselines for Sustainable Development Goal 4 must establish where countries and population groups stand at the beginning of implementation. The task is more demanding than locating the nearest national value. A defensible baseline is tied to a target concept, population, reference date, definition, source coverage and stated uncertainty. It preserves the actual observation year and discloses any estimation, correspondence or gap. It is designed for future comparison but does not claim comparability where programme structures, age bands, thresholds or source coverage differ materially.

As at 27 June 2016, indicator arrangements remained at an early stage. The Inter-Agency and Expert Group had presented a proposed global framework, and the United Nations Statistical Commission had agreed it as a practical starting point subject to future refinement. UNESCO’s thematic proposal supplied broader education concepts. This report uses those materials according to their contemporaneous status. It does not treat later 2016 settlement, metadata revision or results as already available.

Goal 4 creates distinct baseline demands. Completion and learning require cohort, stage, threshold and assessment-population evidence. Early childhood requires age, setting, participation and development domains. Post-school participation requires consistent programme boundaries. Skills and literacy require a clear distinction among direct assessment, qualification and self-report. Global-citizenship learning requires evidence beyond curricular keywords. Facilities and teachers require functionality, actual presence, qualification and learner exposure. Equity requires component group levels, not parity indices alone.

No single source covers these demands. Administrative records provide frequent detail on recognised institutions and registered learners but omit people outside the system. Household surveys reveal participation and disadvantage but bring sampling, response and residence limitations. Assessments observe only a defined target population that participates under stated rules. Censuses and civil registration strengthen denominators but may be incomplete or dated. Finance and teacher systems need to show resources received and personnel actually serving learners. Reconciliation begins with a definition map, not an average of conflicting values.

Conflict and displacement expose the consequences of incomplete coverage. The humanitarian summit period and launch of Education Cannot Wait heightened attention to education in crisis, yet mobile, refugee, displaced and host-community learners often fall between national records, registration sources and temporary providers. A baseline should state whether these populations are included, what evidence exists and how missingness affects the national claim. The

The recommended output is a public national baseline statement. For every target component it records the selected value, actual date, definition, source, coverage, disaggregation, uncertainty, status and owner. It reports unavailable, partial and estimated values honestly; explains differences between national and international figures; costs recurrent improvement; and maintains a public revision history. This creates an accountable starting point without inventing precision or importing later indicator hindsight.

Key findings

  • A baseline is a dated observation linked to a defined target concept and population, not simply the nearest available value.
  • The March 2016 indicator framework was a practical starting point subject to refinement at the cutoff.
  • Goal 4 targets require different source populations, units and definitions; enrolment cannot serve as a universal proxy.
  • National averages require distributional evidence by sex, resources, location, disability and other relevant characteristics.
  • Administrative, household, assessment, population, finance and workforce sources should be reconciled through definition mapping.
  • Comparability must be judged for a stated use across concepts, coverage, precision, classifications and time.
  • Crisis-affected and displaced populations must be explicitly included or identified as a material baseline gap.
  • A public baseline statement should preserve source lineage, provisional status, revisions, responsible institutions and recurrent finance.

Scope and method

This global comparative indicator study addresses education ministries, national statistical offices, international custodians and regional organisations. It evaluates baseline availability and comparability for the adopted Goal 4 targets and the indicator proposals officially available by the cutoff. It does not supply new national estimates or assert a later settled global framework.

Evidence is restricted to official United Nations, UNESCO and European institutional materials available by 27 June 2016. Later 2016 indicator revisions, data digests, monitoring reports, metadata and results are excluded. The report interprets official statistical principles together with education target texts, thematic proposals and contemporaneous humanitarian evidence.

Part I

What constitutes a national baseline

1

Baseline purpose and policy claim

Baseline purpose and policy claim establishes the baseline question for the baseline purpose. For baseline purpose and policy claim, the relevant population or responsible units are national authorities and populations entering Goal 4 monitoring, and the direct evidence concerns starting level and distribution. In examining baseline purpose and policy claim, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on baseline purpose and policy claim, the first number found is not necessarily the best starting point. For comparison of baseline purpose and policy claim, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-01] [REF-03] [REF-04] [REF-06]

The principal error risk in baseline purpose and policy claim is that a convenient historical number becomes the baseline without stating the target concept it represents. In examining baseline purpose and policy claim, the error changes the apparent starting position and can distort every later comparison. Within evidence on baseline purpose and policy claim, authorities should list routes by which people, institutions or events enter and leave the baseline purpose, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of baseline purpose and policy claim, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-01] [REF-03]

The minimum baseline requirement for baseline purpose and policy claim is to write the policy claim, population, measure, date and intended comparison before selecting a value. Within evidence on baseline purpose and policy claim, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of baseline purpose and policy claim, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting baseline purpose and policy claim, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-03] [REF-04]

Source fitness for the baseline purpose depends on who and what can be observed. For comparison of baseline purpose and policy claim, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting baseline purpose and policy claim, for baseline purpose and policy claim, each source should retain its actual date and principal error. For decisions about baseline purpose and policy claim, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-04] [REF-06]

Equity is a baseline condition within baseline purpose and policy claim. In interpreting baseline purpose and policy claim, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about baseline purpose and policy claim, each group level and population share should remain visible beside any gap or parity measure. For baseline purpose and policy claim, intersections require sufficient precision and safe disclosure. In examining baseline purpose and policy claim, if part of national authorities and populations entering Goal 4 monitoring cannot be observed, the missing group and its likely effect on starting level and distribution should be stated rather than absorbed into a favourable average.[REF-01] [REF-06]

Comparability of starting level and distribution should be tested across definition, coverage, classification, precision and time. For decisions about baseline purpose and policy claim, for baseline purpose and policy claim, a common label does not cure different stage structures, thresholds, questions or observation years. For baseline purpose and policy claim, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining baseline purpose and policy claim, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-03] [REF-04] [REF-06]

Responsibility completes baseline purpose and policy claim. For baseline purpose and policy claim, a material gap in the baseline purpose should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining baseline purpose and policy claim, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on baseline purpose and policy claim, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-01] [REF-03] [REF-04] [REF-06]

2

Reference date and observation window

Reference date and observation window establishes the baseline question for the baseline timing. For reference date and observation window, the relevant population or responsible units are learners and services observed around the national starting point, and the direct evidence concerns date, period and lag. In examining reference date and observation window, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on reference date and observation window, the first number found is not necessarily the best starting point. For comparison of reference date and observation window, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-04] [REF-05] [REF-07] [REF-12]

The principal error risk in reference date and observation window is that values from unlike years are presented as one 2015 baseline or later observations are projected backwards silently. In examining reference date and observation window, the error changes the apparent starting position and can distort every later comparison. Within evidence on reference date and observation window, authorities should list routes by which people, institutions or events enter and leave the baseline timing, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of reference date and observation window, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-04] [REF-05]

The minimum baseline requirement for reference date and observation window is to retain actual observation dates and label ranges, interpolation and unavailable status. Within evidence on reference date and observation window, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of reference date and observation window, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting reference date and observation window, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-05] [REF-07]

Source fitness for the baseline timing depends on who and what can be observed. For comparison of reference date and observation window, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting reference date and observation window, for reference date and observation window, each source should retain its actual date and principal error. For decisions about reference date and observation window, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-07] [REF-12]

Equity is a baseline condition within reference date and observation window. In interpreting reference date and observation window, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about reference date and observation window, each group level and population share should remain visible beside any gap or parity measure. For reference date and observation window, intersections require sufficient precision and safe disclosure. In examining reference date and observation window, if part of learners and services observed around the national starting point cannot be observed, the missing group and its likely effect on date, period and lag should be stated rather than absorbed into a favourable average.[REF-04] [REF-12]

Comparability of date, period and lag should be tested across definition, coverage, classification, precision and time. For decisions about reference date and observation window, for reference date and observation window, a common label does not cure different stage structures, thresholds, questions or observation years. For reference date and observation window, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining reference date and observation window, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-05] [REF-07] [REF-12]

Responsibility completes reference date and observation window. For reference date and observation window, a material gap in the baseline timing should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining reference date and observation window, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on reference date and observation window, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-04] [REF-05] [REF-07] [REF-12]

3

Indicator status at June 2016

Indicator status at June 2016 establishes the baseline question for the indicator status. For indicator status at june 2016, the relevant population or responsible units are users interpreting global and thematic proposals, and the direct evidence concerns official availability and provisionality. In examining indicator status at june 2016, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on indicator status at june 2016, the first number found is not necessarily the best starting point. For comparison of indicator status at june 2016, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-04] [REF-05] [REF-06] [REF-12]

The principal error risk in indicator status at june 2016 is that the March framework is described as permanently settled or later revisions are read backwards. In examining indicator status at june 2016, the error changes the apparent starting position and can distort every later comparison. Within evidence on indicator status at june 2016, authorities should list routes by which people, institutions or events enter and leave the indicator status, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of indicator status at june 2016, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-04] [REF-05]

The minimum baseline requirement for indicator status at june 2016 is to state that the Statistical Commission accepted a practical starting point subject to refinement. Within evidence on indicator status at june 2016, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of indicator status at june 2016, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting indicator status at june 2016, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-05] [REF-06]

Source fitness for the indicator status depends on who and what can be observed. For comparison of indicator status at june 2016, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting indicator status at june 2016, for indicator status at june 2016, each source should retain its actual date and principal error. For decisions about indicator status at june 2016, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-06] [REF-12]

Equity is a baseline condition within indicator status at june 2016. In interpreting indicator status at june 2016, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about indicator status at june 2016, each group level and population share should remain visible beside any gap or parity measure. For indicator status at june 2016, intersections require sufficient precision and safe disclosure. In examining indicator status at june 2016, if part of users interpreting global and thematic proposals cannot be observed, the missing group and its likely effect on official availability and provisionality should be stated rather than absorbed into a favourable average.[REF-04] [REF-12]

Comparability of official availability and provisionality should be tested across definition, coverage, classification, precision and time. For decisions about indicator status at june 2016, for indicator status at june 2016, a common label does not cure different stage structures, thresholds, questions or observation years. For indicator status at june 2016, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining indicator status at june 2016, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-05] [REF-06] [REF-12]

Responsibility completes indicator status at june 2016. For indicator status at june 2016, a material gap in the indicator status should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining indicator status at june 2016, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on indicator status at june 2016, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-04] [REF-05] [REF-06] [REF-12]

4

Availability, fitness and comparability

Availability, fitness and comparability establishes the baseline question for the baseline readiness. For availability, fitness and comparability, the relevant population or responsible units are institutions holding candidate values, and the direct evidence concerns data existence, conceptual fit and comparative use. In examining availability, fitness and comparability, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on availability, fitness and comparability, the first number found is not necessarily the best starting point. For comparison of availability, fitness and comparability, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-05] [REF-06] [REF-12] [REF-19]

The principal error risk in availability, fitness and comparability is that availability is confused with fitness and a national figure is assumed comparable because labels match. In examining availability, fitness and comparability, the error changes the apparent starting position and can distort every later comparison. Within evidence on availability, fitness and comparability, authorities should list routes by which people, institutions or events enter and leave the baseline readiness, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of availability, fitness and comparability, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-05] [REF-06]

The minimum baseline requirement for availability, fitness and comparability is to assess existence, definition, coverage, quality, date, disaggregation and correspondence separately. Within evidence on availability, fitness and comparability, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of availability, fitness and comparability, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting availability, fitness and comparability, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-06] [REF-12]

Source fitness for the baseline readiness depends on who and what can be observed. For comparison of availability, fitness and comparability, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting availability, fitness and comparability, for availability, fitness and comparability, each source should retain its actual date and principal error. For decisions about availability, fitness and comparability, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-12] [REF-19]

Equity is a baseline condition within availability, fitness and comparability. In interpreting availability, fitness and comparability, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about availability, fitness and comparability, each group level and population share should remain visible beside any gap or parity measure. For availability, fitness and comparability, intersections require sufficient precision and safe disclosure. In examining availability, fitness and comparability, if part of institutions holding candidate values cannot be observed, the missing group and its likely effect on data existence, conceptual fit and comparative use should be stated rather than absorbed into a favourable average.[REF-05] [REF-19]

Comparability of data existence, conceptual fit and comparative use should be tested across definition, coverage, classification, precision and time. For decisions about availability, fitness and comparability, for availability, fitness and comparability, a common label does not cure different stage structures, thresholds, questions or observation years. For availability, fitness and comparability, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining availability, fitness and comparability, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-06] [REF-12] [REF-19]

Responsibility completes availability, fitness and comparability. For availability, fitness and comparability, a material gap in the baseline readiness should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining availability, fitness and comparability, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on availability, fitness and comparability, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-05] [REF-06] [REF-12] [REF-19]

5

National ownership and international estimation

National ownership and international estimation establishes the baseline question for the baseline ownership. For national ownership and international estimation, the relevant population or responsible units are national statistical authorities and international custodians, and the direct evidence concerns reported and estimated values. In examining national ownership and international estimation, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on national ownership and international estimation, the first number found is not necessarily the best starting point. For comparison of national ownership and international estimation, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-03] [REF-05] [REF-12] [REF-19]

The principal error risk in national ownership and international estimation is that an international estimate overrides national evidence without transparent reconciliation. In examining national ownership and international estimation, the error changes the apparent starting position and can distort every later comparison. Within evidence on national ownership and international estimation, authorities should list routes by which people, institutions or events enter and leave the baseline ownership, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of national ownership and international estimation, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-03] [REF-05]

The minimum baseline requirement for national ownership and international estimation is to retain national authority, publish estimation methods and resolve differences openly. Within evidence on national ownership and international estimation, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of national ownership and international estimation, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting national ownership and international estimation, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-05] [REF-12]

Source fitness for the baseline ownership depends on who and what can be observed. For comparison of national ownership and international estimation, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting national ownership and international estimation, for national ownership and international estimation, each source should retain its actual date and principal error. For decisions about national ownership and international estimation, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-12] [REF-19]

Equity is a baseline condition within national ownership and international estimation. In interpreting national ownership and international estimation, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about national ownership and international estimation, each group level and population share should remain visible beside any gap or parity measure. For national ownership and international estimation, intersections require sufficient precision and safe disclosure. In examining national ownership and international estimation, if part of national statistical authorities and international custodians cannot be observed, the missing group and its likely effect on reported and estimated values should be stated rather than absorbed into a favourable average.[REF-03] [REF-19]

Comparability of reported and estimated values should be tested across definition, coverage, classification, precision and time. For decisions about national ownership and international estimation, for national ownership and international estimation, a common label does not cure different stage structures, thresholds, questions or observation years. For national ownership and international estimation, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining national ownership and international estimation, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-05] [REF-12] [REF-19]

Responsibility completes national ownership and international estimation. For national ownership and international estimation, a material gap in the baseline ownership should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining national ownership and international estimation, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on national ownership and international estimation, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-03] [REF-05] [REF-12] [REF-19]

6

Baseline uncertainty and revision

Baseline uncertainty and revision establishes the baseline question for the revision discipline. For baseline uncertainty and revision, the relevant population or responsible units are users comparing future progress with the starting point, and the direct evidence concerns uncertainty and version. In examining baseline uncertainty and revision, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on baseline uncertainty and revision, the first number found is not necessarily the best starting point. For comparison of baseline uncertainty and revision, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-05] [REF-11] [REF-12] [REF-19]

The principal error risk in baseline uncertainty and revision is that later improved evidence silently changes the baseline and rewrites apparent progress. In examining baseline uncertainty and revision, the error changes the apparent starting position and can distort every later comparison. Within evidence on baseline uncertainty and revision, authorities should list routes by which people, institutions or events enter and leave the revision discipline, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of baseline uncertainty and revision, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-05] [REF-11]

The minimum baseline requirement for baseline uncertainty and revision is to publish uncertainty, version, reason, effect and preserved prior release. Within evidence on baseline uncertainty and revision, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of baseline uncertainty and revision, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting baseline uncertainty and revision, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-11] [REF-12]

Source fitness for the revision discipline depends on who and what can be observed. For comparison of baseline uncertainty and revision, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting baseline uncertainty and revision, for baseline uncertainty and revision, each source should retain its actual date and principal error. For decisions about baseline uncertainty and revision, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-12] [REF-19]

Equity is a baseline condition within baseline uncertainty and revision. In interpreting baseline uncertainty and revision, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about baseline uncertainty and revision, each group level and population share should remain visible beside any gap or parity measure. For baseline uncertainty and revision, intersections require sufficient precision and safe disclosure. In examining baseline uncertainty and revision, if part of users comparing future progress with the starting point cannot be observed, the missing group and its likely effect on uncertainty and version should be stated rather than absorbed into a favourable average.[REF-05] [REF-19]

Comparability of uncertainty and version should be tested across definition, coverage, classification, precision and time. For decisions about baseline uncertainty and revision, for baseline uncertainty and revision, a common label does not cure different stage structures, thresholds, questions or observation years. For baseline uncertainty and revision, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining baseline uncertainty and revision, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-11] [REF-12] [REF-19]

Responsibility completes baseline uncertainty and revision. For baseline uncertainty and revision, a material gap in the revision discipline should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining baseline uncertainty and revision, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on baseline uncertainty and revision, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-05] [REF-11] [REF-12] [REF-19]

Part II

Baseline demands across Goal 4 targets

7

Completion and learning in primary and secondary education

Completion and learning in primary and secondary education establishes the baseline question for the target 4.1 baseline. For completion and learning in primary and secondary education, the relevant population or responsible units are children and adolescents approaching school stages, and the direct evidence concerns completion and minimum learning. In examining completion and learning in primary and secondary education, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on completion and learning in primary and secondary education, the first number found is not necessarily the best starting point. For comparison of completion and learning in primary and secondary education, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-01] [REF-03] [REF-04] [REF-08]

The principal error risk in completion and learning in primary and secondary education is that administrative completion is combined with assessment values from different cohorts and years. In examining completion and learning in primary and secondary education, the error changes the apparent starting position and can distort every later comparison. Within evidence on completion and learning in primary and secondary education, authorities should list routes by which people, institutions or events enter and leave the target 4.1 baseline, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of completion and learning in primary and secondary education, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-01] [REF-03]

The minimum baseline requirement for completion and learning in primary and secondary education is to state cohort, stage, completion rule, assessment population, threshold and date separately. Within evidence on completion and learning in primary and secondary education, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of completion and learning in primary and secondary education, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting completion and learning in primary and secondary education, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-03] [REF-04]

Source fitness for the target 4.1 baseline depends on who and what can be observed. For comparison of completion and learning in primary and secondary education, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting completion and learning in primary and secondary education, for completion and learning in primary and secondary education, each source should retain its actual date and principal error. For decisions about completion and learning in primary and secondary education, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-04] [REF-08]

Equity is a baseline condition within completion and learning in primary and secondary education. In interpreting completion and learning in primary and secondary education, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about completion and learning in primary and secondary education, each group level and population share should remain visible beside any gap or parity measure. For completion and learning in primary and secondary education, intersections require sufficient precision and safe disclosure. In examining completion and learning in primary and secondary education, if part of children and adolescents approaching school stages cannot be observed, the missing group and its likely effect on completion and minimum learning should be stated rather than absorbed into a favourable average.[REF-01] [REF-08]

Comparability of completion and minimum learning should be tested across definition, coverage, classification, precision and time. For decisions about completion and learning in primary and secondary education, for completion and learning in primary and secondary education, a common label does not cure different stage structures, thresholds, questions or observation years. For completion and learning in primary and secondary education, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining completion and learning in primary and secondary education, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-03] [REF-04] [REF-08]

Responsibility completes completion and learning in primary and secondary education. For completion and learning in primary and secondary education, a material gap in the target 4.1 baseline should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining completion and learning in primary and secondary education, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on completion and learning in primary and secondary education, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-01] [REF-03] [REF-04] [REF-08]

8

Early childhood development and pre-primary education

Early childhood development and pre-primary education establishes the baseline question for the target 4.2 baseline. For early childhood development and pre-primary education, the relevant population or responsible units are young children and children one year before primary entry, and the direct evidence concerns development and organised participation. In examining early childhood development and pre-primary education, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on early childhood development and pre-primary education, the first number found is not necessarily the best starting point. For comparison of early childhood development and pre-primary education, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-01] [REF-03] [REF-06] [REF-16]

The principal error risk in early childhood development and pre-primary education is that pre-primary records represent all early learning or one development measure stands for the whole child. In examining early childhood development and pre-primary education, the error changes the apparent starting position and can distort every later comparison. Within evidence on early childhood development and pre-primary education, authorities should list routes by which people, institutions or events enter and leave the target 4.2 baseline, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of early childhood development and pre-primary education, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-01] [REF-03]

The minimum baseline requirement for early childhood development and pre-primary education is to separate age, setting, programme participation, developmental domain and household coverage. Within evidence on early childhood development and pre-primary education, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of early childhood development and pre-primary education, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting early childhood development and pre-primary education, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-03] [REF-06]

Source fitness for the target 4.2 baseline depends on who and what can be observed. For comparison of early childhood development and pre-primary education, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting early childhood development and pre-primary education, for early childhood development and pre-primary education, each source should retain its actual date and principal error. For decisions about early childhood development and pre-primary education, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-06] [REF-16]

Equity is a baseline condition within early childhood development and pre-primary education. In interpreting early childhood development and pre-primary education, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about early childhood development and pre-primary education, each group level and population share should remain visible beside any gap or parity measure. For early childhood development and pre-primary education, intersections require sufficient precision and safe disclosure. In examining early childhood development and pre-primary education, if part of young children and children one year before primary entry cannot be observed, the missing group and its likely effect on development and organised participation should be stated rather than absorbed into a favourable average.[REF-01] [REF-16]

Comparability of development and organised participation should be tested across definition, coverage, classification, precision and time. For decisions about early childhood development and pre-primary education, for early childhood development and pre-primary education, a common label does not cure different stage structures, thresholds, questions or observation years. For early childhood development and pre-primary education, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining early childhood development and pre-primary education, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-03] [REF-06] [REF-16]

Responsibility completes early childhood development and pre-primary education. For early childhood development and pre-primary education, a material gap in the target 4.2 baseline should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining early childhood development and pre-primary education, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on early childhood development and pre-primary education, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-01] [REF-03] [REF-06] [REF-16]

9

Technical, vocational and tertiary participation

Technical, vocational and tertiary participation establishes the baseline question for the target 4.3 baseline. For technical, vocational and tertiary participation, the relevant population or responsible units are youth and adults entering formal and non-formal post-school learning, and the direct evidence concerns participation by route. In examining technical, vocational and tertiary participation, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on technical, vocational and tertiary participation, the first number found is not necessarily the best starting point. For comparison of technical, vocational and tertiary participation, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-01] [REF-03] [REF-06] [REF-09]

The principal error risk in technical, vocational and tertiary participation is that programme classifications and short-course boundaries differ across countries and sources. In examining technical, vocational and tertiary participation, the error changes the apparent starting position and can distort every later comparison. Within evidence on technical, vocational and tertiary participation, authorities should list routes by which people, institutions or events enter and leave the target 4.3 baseline, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of technical, vocational and tertiary participation, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-01] [REF-03]

The minimum baseline requirement for technical, vocational and tertiary participation is to map national programmes to common levels and distinguish person, activity and enrolment counts. Within evidence on technical, vocational and tertiary participation, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of technical, vocational and tertiary participation, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting technical, vocational and tertiary participation, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-03] [REF-06]

Source fitness for the target 4.3 baseline depends on who and what can be observed. For comparison of technical, vocational and tertiary participation, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting technical, vocational and tertiary participation, for technical, vocational and tertiary participation, each source should retain its actual date and principal error. For decisions about technical, vocational and tertiary participation, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-06] [REF-09]

Equity is a baseline condition within technical, vocational and tertiary participation. In interpreting technical, vocational and tertiary participation, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about technical, vocational and tertiary participation, each group level and population share should remain visible beside any gap or parity measure. For technical, vocational and tertiary participation, intersections require sufficient precision and safe disclosure. In examining technical, vocational and tertiary participation, if part of youth and adults entering formal and non-formal post-school learning cannot be observed, the missing group and its likely effect on participation by route should be stated rather than absorbed into a favourable average.[REF-01] [REF-09]

Comparability of participation by route should be tested across definition, coverage, classification, precision and time. For decisions about technical, vocational and tertiary participation, for technical, vocational and tertiary participation, a common label does not cure different stage structures, thresholds, questions or observation years. For technical, vocational and tertiary participation, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining technical, vocational and tertiary participation, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-03] [REF-06] [REF-09]

Responsibility completes technical, vocational and tertiary participation. For technical, vocational and tertiary participation, a material gap in the target 4.3 baseline should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining technical, vocational and tertiary participation, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on technical, vocational and tertiary participation, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-01] [REF-03] [REF-06] [REF-09]

10

Skills for employment, decent work and entrepreneurship

Skills for employment, decent work and entrepreneurship establishes the baseline question for the target 4.4 baseline. For skills for employment, decent work and entrepreneurship, the relevant population or responsible units are youth and adults with differing education and labour status, and the direct evidence concerns skill possession and participation. In examining skills for employment, decent work and entrepreneurship, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on skills for employment, decent work and entrepreneurship, the first number found is not necessarily the best starting point. For comparison of skills for employment, decent work and entrepreneurship, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-01] [REF-03] [REF-06] [REF-18]

The principal error risk in skills for employment, decent work and entrepreneurship is that qualification, self-report and demonstrated skill are treated as equivalent. In examining skills for employment, decent work and entrepreneurship, the error changes the apparent starting position and can distort every later comparison. Within evidence on skills for employment, decent work and entrepreneurship, authorities should list routes by which people, institutions or events enter and leave the target 4.4 baseline, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of skills for employment, decent work and entrepreneurship, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-01] [REF-03]

The minimum baseline requirement for skills for employment, decent work and entrepreneurship is to state skill construct, assessment or proxy, age, work status and inference limit. Within evidence on skills for employment, decent work and entrepreneurship, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of skills for employment, decent work and entrepreneurship, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting skills for employment, decent work and entrepreneurship, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-03] [REF-06]

Source fitness for the target 4.4 baseline depends on who and what can be observed. For comparison of skills for employment, decent work and entrepreneurship, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting skills for employment, decent work and entrepreneurship, for skills for employment, decent work and entrepreneurship, each source should retain its actual date and principal error. For decisions about skills for employment, decent work and entrepreneurship, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-06] [REF-18]

Equity is a baseline condition within skills for employment, decent work and entrepreneurship. In interpreting skills for employment, decent work and entrepreneurship, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about skills for employment, decent work and entrepreneurship, each group level and population share should remain visible beside any gap or parity measure. For skills for employment, decent work and entrepreneurship, intersections require sufficient precision and safe disclosure. In examining skills for employment, decent work and entrepreneurship, if part of youth and adults with differing education and labour status cannot be observed, the missing group and its likely effect on skill possession and participation should be stated rather than absorbed into a favourable average.[REF-01] [REF-18]

Comparability of skill possession and participation should be tested across definition, coverage, classification, precision and time. For decisions about skills for employment, decent work and entrepreneurship, for skills for employment, decent work and entrepreneurship, a common label does not cure different stage structures, thresholds, questions or observation years. For skills for employment, decent work and entrepreneurship, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining skills for employment, decent work and entrepreneurship, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-03] [REF-06] [REF-18]

Responsibility completes skills for employment, decent work and entrepreneurship. For skills for employment, decent work and entrepreneurship, a material gap in the target 4.4 baseline should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining skills for employment, decent work and entrepreneurship, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on skills for employment, decent work and entrepreneurship, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-01] [REF-03] [REF-06] [REF-18]

11

Youth and adult literacy and numeracy

Youth and adult literacy and numeracy establishes the baseline question for the target 4.6 baseline. For youth and adult literacy and numeracy, the relevant population or responsible units are youth and adults across age and prior schooling, and the direct evidence concerns proficiency and attainment. In examining youth and adult literacy and numeracy, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on youth and adult literacy and numeracy, the first number found is not necessarily the best starting point. For comparison of youth and adult literacy and numeracy, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-01] [REF-03] [REF-06] [REF-10]

The principal error risk in youth and adult literacy and numeracy is that self-declared literacy or years of schooling replace assessed competence without qualification. In examining youth and adult literacy and numeracy, the error changes the apparent starting position and can distort every later comparison. Within evidence on youth and adult literacy and numeracy, authorities should list routes by which people, institutions or events enter and leave the target 4.6 baseline, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of youth and adult literacy and numeracy, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-01] [REF-03]

The minimum baseline requirement for youth and adult literacy and numeracy is to distinguish direct assessment, self-report, attainment and programme participation. Within evidence on youth and adult literacy and numeracy, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of youth and adult literacy and numeracy, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting youth and adult literacy and numeracy, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-03] [REF-06]

Source fitness for the target 4.6 baseline depends on who and what can be observed. For comparison of youth and adult literacy and numeracy, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting youth and adult literacy and numeracy, for youth and adult literacy and numeracy, each source should retain its actual date and principal error. For decisions about youth and adult literacy and numeracy, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-06] [REF-10]

Equity is a baseline condition within youth and adult literacy and numeracy. In interpreting youth and adult literacy and numeracy, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about youth and adult literacy and numeracy, each group level and population share should remain visible beside any gap or parity measure. For youth and adult literacy and numeracy, intersections require sufficient precision and safe disclosure. In examining youth and adult literacy and numeracy, if part of youth and adults across age and prior schooling cannot be observed, the missing group and its likely effect on proficiency and attainment should be stated rather than absorbed into a favourable average.[REF-01] [REF-10]

Comparability of proficiency and attainment should be tested across definition, coverage, classification, precision and time. For decisions about youth and adult literacy and numeracy, for youth and adult literacy and numeracy, a common label does not cure different stage structures, thresholds, questions or observation years. For youth and adult literacy and numeracy, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining youth and adult literacy and numeracy, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-03] [REF-06] [REF-10]

Responsibility completes youth and adult literacy and numeracy. For youth and adult literacy and numeracy, a material gap in the target 4.6 baseline should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining youth and adult literacy and numeracy, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on youth and adult literacy and numeracy, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-01] [REF-03] [REF-06] [REF-10]

12

Global citizenship and sustainable-development learning

Global citizenship and sustainable-development learning establishes the baseline question for the target 4.7 baseline. For global citizenship and sustainable-development learning, the relevant population or responsible units are learners and institutions at different stages, and the direct evidence concerns curriculum, teacher preparation and learner evidence. In examining global citizenship and sustainable-development learning, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on global citizenship and sustainable-development learning, the first number found is not necessarily the best starting point. For comparison of global citizenship and sustainable-development learning, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-01] [REF-03] [REF-06] [REF-17]

The principal error risk in global citizenship and sustainable-development learning is that keyword presence is treated as learning or a later indicator is imposed on early evidence. In examining global citizenship and sustainable-development learning, the error changes the apparent starting position and can distort every later comparison. Within evidence on global citizenship and sustainable-development learning, authorities should list routes by which people, institutions or events enter and leave the target 4.7 baseline, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of global citizenship and sustainable-development learning, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-01] [REF-03]

The minimum baseline requirement for global citizenship and sustainable-development learning is to record intended curriculum, opportunity, enacted teaching and available learner evidence by date. Within evidence on global citizenship and sustainable-development learning, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of global citizenship and sustainable-development learning, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting global citizenship and sustainable-development learning, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-03] [REF-06]

Source fitness for the target 4.7 baseline depends on who and what can be observed. For comparison of global citizenship and sustainable-development learning, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting global citizenship and sustainable-development learning, for global citizenship and sustainable-development learning, each source should retain its actual date and principal error. For decisions about global citizenship and sustainable-development learning, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-06] [REF-17]

Equity is a baseline condition within global citizenship and sustainable-development learning. In interpreting global citizenship and sustainable-development learning, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about global citizenship and sustainable-development learning, each group level and population share should remain visible beside any gap or parity measure. For global citizenship and sustainable-development learning, intersections require sufficient precision and safe disclosure. In examining global citizenship and sustainable-development learning, if part of learners and institutions at different stages cannot be observed, the missing group and its likely effect on curriculum, teacher preparation and learner evidence should be stated rather than absorbed into a favourable average.[REF-01] [REF-17]

Comparability of curriculum, teacher preparation and learner evidence should be tested across definition, coverage, classification, precision and time. For decisions about global citizenship and sustainable-development learning, for global citizenship and sustainable-development learning, a common label does not cure different stage structures, thresholds, questions or observation years. For global citizenship and sustainable-development learning, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining global citizenship and sustainable-development learning, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-03] [REF-06] [REF-17]

Responsibility completes global citizenship and sustainable-development learning. For global citizenship and sustainable-development learning, a material gap in the target 4.7 baseline should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining global citizenship and sustainable-development learning, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on global citizenship and sustainable-development learning, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-01] [REF-03] [REF-06] [REF-17]

Part III

Equity, enabling conditions and crisis visibility

13

Parity and distribution under target 4.5

Parity and distribution under target 4.5 establishes the baseline question for the equity baseline. For parity and distribution under target 4.5, the relevant population or responsible units are learners by sex, wealth, location and other relevant characteristics, and the direct evidence concerns group levels, gaps and parity. In examining parity and distribution under target 4.5, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on parity and distribution under target 4.5, the first number found is not necessarily the best starting point. For comparison of parity and distribution under target 4.5, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-01] [REF-03] [REF-06] [REF-14]

The principal error risk in parity and distribution under target 4.5 is that a parity index conceals low levels for both groups and missing populations. In examining parity and distribution under target 4.5, the error changes the apparent starting position and can distort every later comparison. Within evidence on parity and distribution under target 4.5, authorities should list routes by which people, institutions or events enter and leave the equity baseline, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of parity and distribution under target 4.5, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-01] [REF-03]

The minimum baseline requirement for parity and distribution under target 4.5 is to publish component levels, population shares, absolute and relative gaps and uncertainty. Within evidence on parity and distribution under target 4.5, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of parity and distribution under target 4.5, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting parity and distribution under target 4.5, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-03] [REF-06]

Source fitness for the equity baseline depends on who and what can be observed. For comparison of parity and distribution under target 4.5, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting parity and distribution under target 4.5, for parity and distribution under target 4.5, each source should retain its actual date and principal error. For decisions about parity and distribution under target 4.5, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-06] [REF-14]

Equity is a baseline condition within parity and distribution under target 4.5. In interpreting parity and distribution under target 4.5, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about parity and distribution under target 4.5, each group level and population share should remain visible beside any gap or parity measure. For parity and distribution under target 4.5, intersections require sufficient precision and safe disclosure. In examining parity and distribution under target 4.5, if part of learners by sex, wealth, location and other relevant characteristics cannot be observed, the missing group and its likely effect on group levels, gaps and parity should be stated rather than absorbed into a favourable average.[REF-01] [REF-14]

Comparability of group levels, gaps and parity should be tested across definition, coverage, classification, precision and time. For decisions about parity and distribution under target 4.5, for parity and distribution under target 4.5, a common label does not cure different stage structures, thresholds, questions or observation years. For parity and distribution under target 4.5, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining parity and distribution under target 4.5, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-03] [REF-06] [REF-14]

Responsibility completes parity and distribution under target 4.5. For parity and distribution under target 4.5, a material gap in the equity baseline should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining parity and distribution under target 4.5, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on parity and distribution under target 4.5, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-01] [REF-03] [REF-06] [REF-14]

14

Disability and functional difficulty

Disability and functional difficulty establishes the baseline question for the disability baseline. For disability and functional difficulty, the relevant population or responsible units are learners with different functional and accommodation requirements, and the direct evidence concerns participation, learning and accessibility. In examining disability and functional difficulty, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on disability and functional difficulty, the first number found is not necessarily the best starting point. For comparison of disability and functional difficulty, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-06] [REF-13] [REF-15] [REF-16]

The principal error risk in disability and functional difficulty is that diagnosis-based school records omit never-enrolled learners and incomparable categories are merged. In examining disability and functional difficulty, the error changes the apparent starting position and can distort every later comparison. Within evidence on disability and functional difficulty, authorities should list routes by which people, institutions or events enter and leave the disability baseline, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of disability and functional difficulty, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-06] [REF-13]

The minimum baseline requirement for disability and functional difficulty is to use respectful functional evidence and document accessibility, accommodation and coverage. Within evidence on disability and functional difficulty, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of disability and functional difficulty, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting disability and functional difficulty, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-13] [REF-15]

Source fitness for the disability baseline depends on who and what can be observed. For comparison of disability and functional difficulty, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting disability and functional difficulty, for disability and functional difficulty, each source should retain its actual date and principal error. For decisions about disability and functional difficulty, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-15] [REF-16]

Equity is a baseline condition within disability and functional difficulty. In interpreting disability and functional difficulty, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about disability and functional difficulty, each group level and population share should remain visible beside any gap or parity measure. For disability and functional difficulty, intersections require sufficient precision and safe disclosure. In examining disability and functional difficulty, if part of learners with different functional and accommodation requirements cannot be observed, the missing group and its likely effect on participation, learning and accessibility should be stated rather than absorbed into a favourable average.[REF-06] [REF-16]

Comparability of participation, learning and accessibility should be tested across definition, coverage, classification, precision and time. For decisions about disability and functional difficulty, for disability and functional difficulty, a common label does not cure different stage structures, thresholds, questions or observation years. For disability and functional difficulty, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining disability and functional difficulty, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-13] [REF-15] [REF-16]

Responsibility completes disability and functional difficulty. For disability and functional difficulty, a material gap in the disability baseline should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining disability and functional difficulty, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on disability and functional difficulty, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-06] [REF-13] [REF-15] [REF-16]

15

Facilities, safety and inclusive environments

Facilities, safety and inclusive environments establishes the baseline question for the target 4.a baseline. For facilities, safety and inclusive environments, the relevant population or responsible units are learners and institutions using different facilities and routes, and the direct evidence concerns usable water, sanitation, access, safety and learning conditions. In examining facilities, safety and inclusive environments, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on facilities, safety and inclusive environments, the first number found is not necessarily the best starting point. For comparison of facilities, safety and inclusive environments, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-01] [REF-03] [REF-06] [REF-23]

The principal error risk in facilities, safety and inclusive environments is that facility presence is counted although unavailable, inaccessible or unsafe in regular use. In examining facilities, safety and inclusive environments, the error changes the apparent starting position and can distort every later comparison. Within evidence on facilities, safety and inclusive environments, authorities should list routes by which people, institutions or events enter and leave the target 4.a baseline, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of facilities, safety and inclusive environments, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-01] [REF-03]

The minimum baseline requirement for facilities, safety and inclusive environments is to define functionality and learner use and verify distribution across institutions and groups. Within evidence on facilities, safety and inclusive environments, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of facilities, safety and inclusive environments, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting facilities, safety and inclusive environments, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-03] [REF-06]

Source fitness for the target 4.a baseline depends on who and what can be observed. For comparison of facilities, safety and inclusive environments, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting facilities, safety and inclusive environments, for facilities, safety and inclusive environments, each source should retain its actual date and principal error. For decisions about facilities, safety and inclusive environments, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-06] [REF-23]

Equity is a baseline condition within facilities, safety and inclusive environments. In interpreting facilities, safety and inclusive environments, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about facilities, safety and inclusive environments, each group level and population share should remain visible beside any gap or parity measure. For facilities, safety and inclusive environments, intersections require sufficient precision and safe disclosure. In examining facilities, safety and inclusive environments, if part of learners and institutions using different facilities and routes cannot be observed, the missing group and its likely effect on usable water, sanitation, access, safety and learning conditions should be stated rather than absorbed into a favourable average.[REF-01] [REF-23]

Comparability of usable water, sanitation, access, safety and learning conditions should be tested across definition, coverage, classification, precision and time. For decisions about facilities, safety and inclusive environments, for facilities, safety and inclusive environments, a common label does not cure different stage structures, thresholds, questions or observation years. For facilities, safety and inclusive environments, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining facilities, safety and inclusive environments, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-03] [REF-06] [REF-23]

Responsibility completes facilities, safety and inclusive environments. For facilities, safety and inclusive environments, a material gap in the target 4.a baseline should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining facilities, safety and inclusive environments, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on facilities, safety and inclusive environments, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-01] [REF-03] [REF-06] [REF-23]

16

Teachers, qualifications and distribution

Teachers, qualifications and distribution establishes the baseline question for the target 4.c baseline. For teachers, qualifications and distribution, the relevant population or responsible units are teachers and learners in differently resourced settings, and the direct evidence concerns qualification, training and learner exposure. In examining teachers, qualifications and distribution, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on teachers, qualifications and distribution, the first number found is not necessarily the best starting point. For comparison of teachers, qualifications and distribution, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-01] [REF-03] [REF-06] [REF-17]

The principal error risk in teachers, qualifications and distribution is that national teacher shares conceal subject gaps, double counting and unequal assignment. In examining teachers, qualifications and distribution, the error changes the apparent starting position and can distort every later comparison. Within evidence on teachers, qualifications and distribution, authorities should list routes by which people, institutions or events enter and leave the target 4.c baseline, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of teachers, qualifications and distribution, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-01] [REF-03]

The minimum baseline requirement for teachers, qualifications and distribution is to define teacher unit and qualification and report learner exposure by location and level. Within evidence on teachers, qualifications and distribution, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of teachers, qualifications and distribution, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting teachers, qualifications and distribution, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-03] [REF-06]

Source fitness for the target 4.c baseline depends on who and what can be observed. For comparison of teachers, qualifications and distribution, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting teachers, qualifications and distribution, for teachers, qualifications and distribution, each source should retain its actual date and principal error. For decisions about teachers, qualifications and distribution, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-06] [REF-17]

Equity is a baseline condition within teachers, qualifications and distribution. In interpreting teachers, qualifications and distribution, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about teachers, qualifications and distribution, each group level and population share should remain visible beside any gap or parity measure. For teachers, qualifications and distribution, intersections require sufficient precision and safe disclosure. In examining teachers, qualifications and distribution, if part of teachers and learners in differently resourced settings cannot be observed, the missing group and its likely effect on qualification, training and learner exposure should be stated rather than absorbed into a favourable average.[REF-01] [REF-17]

Comparability of qualification, training and learner exposure should be tested across definition, coverage, classification, precision and time. For decisions about teachers, qualifications and distribution, for teachers, qualifications and distribution, a common label does not cure different stage structures, thresholds, questions or observation years. For teachers, qualifications and distribution, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining teachers, qualifications and distribution, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-03] [REF-06] [REF-17]

Responsibility completes teachers, qualifications and distribution. For teachers, qualifications and distribution, a material gap in the target 4.c baseline should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining teachers, qualifications and distribution, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on teachers, qualifications and distribution, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-01] [REF-03] [REF-06] [REF-17]

17

Scholarships and cross-border evidence

Scholarships and cross-border evidence establishes the baseline question for the target 4.b baseline. For scholarships and cross-border evidence, the relevant population or responsible units are scholarship recipients and eligible populations, and the direct evidence concerns award, field, destination and completion. In examining scholarships and cross-border evidence, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on scholarships and cross-border evidence, the first number found is not necessarily the best starting point. For comparison of scholarships and cross-border evidence, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-01] [REF-03] [REF-04] [REF-06]

The principal error risk in scholarships and cross-border evidence is that announced funds or awards replace recipient and learning evidence and duplicate sources count the same person. In examining scholarships and cross-border evidence, the error changes the apparent starting position and can distort every later comparison. Within evidence on scholarships and cross-border evidence, authorities should list routes by which people, institutions or events enter and leave the target 4.b baseline, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of scholarships and cross-border evidence, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-01] [REF-03]

The minimum baseline requirement for scholarships and cross-border evidence is to separate commitments, awards, uptake, study status and result with safe records. Within evidence on scholarships and cross-border evidence, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of scholarships and cross-border evidence, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting scholarships and cross-border evidence, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-03] [REF-04]

Source fitness for the target 4.b baseline depends on who and what can be observed. For comparison of scholarships and cross-border evidence, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting scholarships and cross-border evidence, for scholarships and cross-border evidence, each source should retain its actual date and principal error. For decisions about scholarships and cross-border evidence, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-04] [REF-06]

Equity is a baseline condition within scholarships and cross-border evidence. In interpreting scholarships and cross-border evidence, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about scholarships and cross-border evidence, each group level and population share should remain visible beside any gap or parity measure. For scholarships and cross-border evidence, intersections require sufficient precision and safe disclosure. In examining scholarships and cross-border evidence, if part of scholarship recipients and eligible populations cannot be observed, the missing group and its likely effect on award, field, destination and completion should be stated rather than absorbed into a favourable average.[REF-01] [REF-06]

Comparability of award, field, destination and completion should be tested across definition, coverage, classification, precision and time. For decisions about scholarships and cross-border evidence, for scholarships and cross-border evidence, a common label does not cure different stage structures, thresholds, questions or observation years. For scholarships and cross-border evidence, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining scholarships and cross-border evidence, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-03] [REF-04] [REF-06]

Responsibility completes scholarships and cross-border evidence. For scholarships and cross-border evidence, a material gap in the target 4.b baseline should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining scholarships and cross-border evidence, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on scholarships and cross-border evidence, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-01] [REF-03] [REF-04] [REF-06]

18

Conflict, displacement and humanitarian gaps

Conflict, displacement and humanitarian gaps establishes the baseline question for the crisis baseline. For conflict, displacement and humanitarian gaps, the relevant population or responsible units are refugee, displaced, crisis-affected and host-community learners, and the direct evidence concerns population, participation and continuity. In examining conflict, displacement and humanitarian gaps, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on conflict, displacement and humanitarian gaps, the first number found is not necessarily the best starting point. For comparison of conflict, displacement and humanitarian gaps, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-21] [REF-22] [REF-23] [REF-24]

The principal error risk in conflict, displacement and humanitarian gaps is that national systems omit mobile populations or temporary provision and therefore overstate coverage. In examining conflict, displacement and humanitarian gaps, the error changes the apparent starting position and can distort every later comparison. Within evidence on conflict, displacement and humanitarian gaps, authorities should list routes by which people, institutions or events enter and leave the crisis baseline, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of conflict, displacement and humanitarian gaps, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-21] [REF-22]

The minimum baseline requirement for conflict, displacement and humanitarian gaps is to use interoperable bounded records, household or registration evidence and explicit missingness. Within evidence on conflict, displacement and humanitarian gaps, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of conflict, displacement and humanitarian gaps, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting conflict, displacement and humanitarian gaps, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-22] [REF-23]

Source fitness for the crisis baseline depends on who and what can be observed. For comparison of conflict, displacement and humanitarian gaps, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting conflict, displacement and humanitarian gaps, for conflict, displacement and humanitarian gaps, each source should retain its actual date and principal error. For decisions about conflict, displacement and humanitarian gaps, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-23] [REF-24]

Equity is a baseline condition within conflict, displacement and humanitarian gaps. In interpreting conflict, displacement and humanitarian gaps, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about conflict, displacement and humanitarian gaps, each group level and population share should remain visible beside any gap or parity measure. For conflict, displacement and humanitarian gaps, intersections require sufficient precision and safe disclosure. In examining conflict, displacement and humanitarian gaps, if part of refugee, displaced, crisis-affected and host-community learners cannot be observed, the missing group and its likely effect on population, participation and continuity should be stated rather than absorbed into a favourable average.[REF-21] [REF-24]

Comparability of population, participation and continuity should be tested across definition, coverage, classification, precision and time. For decisions about conflict, displacement and humanitarian gaps, for conflict, displacement and humanitarian gaps, a common label does not cure different stage structures, thresholds, questions or observation years. For conflict, displacement and humanitarian gaps, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining conflict, displacement and humanitarian gaps, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-22] [REF-23] [REF-24]

Responsibility completes conflict, displacement and humanitarian gaps. For conflict, displacement and humanitarian gaps, a material gap in the crisis baseline should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining conflict, displacement and humanitarian gaps, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on conflict, displacement and humanitarian gaps, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-21] [REF-22] [REF-23] [REF-24]

Part IV

Source architecture and reconciliation

19

Administrative education records

Administrative education records establishes the baseline question for the administrative baseline source. For administrative education records, the relevant population or responsible units are recognised institutions, staff and registered learners, and the direct evidence concerns service, enrolment and flow. In examining administrative education records, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on administrative education records, the first number found is not necessarily the best starting point. For comparison of administrative education records, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-05] [REF-06] [REF-08] [REF-12]

The principal error risk in administrative education records is that never-enrolled learners and non-reporting institutions are absent while duplicate records inflate totals. In examining administrative education records, the error changes the apparent starting position and can distort every later comparison. Within evidence on administrative education records, authorities should list routes by which people, institutions or events enter and leave the administrative baseline source, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of administrative education records, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-05] [REF-06]

The minimum baseline requirement for administrative education records is to publish institution coverage, reporting completeness, unique-person rules and event definitions. Within evidence on administrative education records, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of administrative education records, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting administrative education records, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-06] [REF-08]

Source fitness for the administrative baseline source depends on who and what can be observed. For comparison of administrative education records, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting administrative education records, for administrative education records, each source should retain its actual date and principal error. For decisions about administrative education records, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-08] [REF-12]

Equity is a baseline condition within administrative education records. In interpreting administrative education records, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about administrative education records, each group level and population share should remain visible beside any gap or parity measure. For administrative education records, intersections require sufficient precision and safe disclosure. In examining administrative education records, if part of recognised institutions, staff and registered learners cannot be observed, the missing group and its likely effect on service, enrolment and flow should be stated rather than absorbed into a favourable average.[REF-05] [REF-12]

Comparability of service, enrolment and flow should be tested across definition, coverage, classification, precision and time. For decisions about administrative education records, for administrative education records, a common label does not cure different stage structures, thresholds, questions or observation years. For administrative education records, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining administrative education records, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-06] [REF-08] [REF-12]

Responsibility completes administrative education records. For administrative education records, a material gap in the administrative baseline source should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining administrative education records, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on administrative education records, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-05] [REF-06] [REF-08] [REF-12]

20

Household surveys

Household surveys establishes the baseline question for the household baseline source. For household surveys, the relevant population or responsible units are persons living in sampled households, and the direct evidence concerns participation, attainment and disadvantage. In examining household surveys, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on household surveys, the first number found is not necessarily the best starting point. For comparison of household surveys, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-10] [REF-11] [REF-14] [REF-19]

The principal error risk in household surveys is that sampling, proxy response and exclusion of institutional or mobile populations distort group results. In examining household surveys, the error changes the apparent starting position and can distort every later comparison. Within evidence on household surveys, authorities should list routes by which people, institutions or events enter and leave the household baseline source, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of household surveys, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-10] [REF-11]

The minimum baseline requirement for household surveys is to state frame, questions, response, weights, intervals and excluded populations. Within evidence on household surveys, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of household surveys, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting household surveys, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-11] [REF-14]

Source fitness for the household baseline source depends on who and what can be observed. For comparison of household surveys, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting household surveys, for household surveys, each source should retain its actual date and principal error. For decisions about household surveys, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-14] [REF-19]

Equity is a baseline condition within household surveys. In interpreting household surveys, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about household surveys, each group level and population share should remain visible beside any gap or parity measure. For household surveys, intersections require sufficient precision and safe disclosure. In examining household surveys, if part of persons living in sampled households cannot be observed, the missing group and its likely effect on participation, attainment and disadvantage should be stated rather than absorbed into a favourable average.[REF-10] [REF-19]

Comparability of participation, attainment and disadvantage should be tested across definition, coverage, classification, precision and time. For decisions about household surveys, for household surveys, a common label does not cure different stage structures, thresholds, questions or observation years. For household surveys, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining household surveys, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-11] [REF-14] [REF-19]

Responsibility completes household surveys. For household surveys, a material gap in the household baseline source should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining household surveys, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on household surveys, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-10] [REF-11] [REF-14] [REF-19]

21

Learning assessments

Learning assessments establishes the baseline question for the assessment baseline source. For learning assessments, the relevant population or responsible units are eligible learners reached under stated conditions, and the direct evidence concerns achievement distribution. In examining learning assessments, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on learning assessments, the first number found is not necessarily the best starting point. For comparison of learning assessments, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-04] [REF-06] [REF-15] [REF-17]

The principal error risk in learning assessments is that absence, language, disability and school exclusion remove learners at greatest risk. In examining learning assessments, the error changes the apparent starting position and can distort every later comparison. Within evidence on learning assessments, authorities should list routes by which people, institutions or events enter and leave the assessment baseline source, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of learning assessments, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-04] [REF-06]

The minimum baseline requirement for learning assessments is to report target population, exclusions, participation, accommodations, thresholds and uncertainty. Within evidence on learning assessments, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of learning assessments, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting learning assessments, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-06] [REF-15]

Source fitness for the assessment baseline source depends on who and what can be observed. For comparison of learning assessments, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting learning assessments, for learning assessments, each source should retain its actual date and principal error. For decisions about learning assessments, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-15] [REF-17]

Equity is a baseline condition within learning assessments. In interpreting learning assessments, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about learning assessments, each group level and population share should remain visible beside any gap or parity measure. For learning assessments, intersections require sufficient precision and safe disclosure. In examining learning assessments, if part of eligible learners reached under stated conditions cannot be observed, the missing group and its likely effect on achievement distribution should be stated rather than absorbed into a favourable average.[REF-04] [REF-17]

Comparability of achievement distribution should be tested across definition, coverage, classification, precision and time. For decisions about learning assessments, for learning assessments, a common label does not cure different stage structures, thresholds, questions or observation years. For learning assessments, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining learning assessments, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-06] [REF-15] [REF-17]

Responsibility completes learning assessments. For learning assessments, a material gap in the assessment baseline source should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining learning assessments, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on learning assessments, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-04] [REF-06] [REF-15] [REF-17]

22

Population censuses and civil registration

Population censuses and civil registration establishes the baseline question for the population baseline source. For population censuses and civil registration, the relevant population or responsible units are residents and birth cohorts within stated rules, and the direct evidence concerns denominator and cohort size. In examining population censuses and civil registration, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on population censuses and civil registration, the first number found is not necessarily the best starting point. For comparison of population censuses and civil registration, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-05] [REF-10] [REF-12] [REF-16]

The principal error risk in population censuses and civil registration is that undercount, outdated projections and weak registration distort age-specific rates. In examining population censuses and civil registration, the error changes the apparent starting position and can distort every later comparison. Within evidence on population censuses and civil registration, authorities should list routes by which people, institutions or events enter and leave the population baseline source, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of population censuses and civil registration, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-05] [REF-10]

The minimum baseline requirement for population censuses and civil registration is to state reference date, residence, undercount, adjustment and revision. Within evidence on population censuses and civil registration, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of population censuses and civil registration, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting population censuses and civil registration, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-10] [REF-12]

Source fitness for the population baseline source depends on who and what can be observed. For comparison of population censuses and civil registration, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting population censuses and civil registration, for population censuses and civil registration, each source should retain its actual date and principal error. For decisions about population censuses and civil registration, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-12] [REF-16]

Equity is a baseline condition within population censuses and civil registration. In interpreting population censuses and civil registration, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about population censuses and civil registration, each group level and population share should remain visible beside any gap or parity measure. For population censuses and civil registration, intersections require sufficient precision and safe disclosure. In examining population censuses and civil registration, if part of residents and birth cohorts within stated rules cannot be observed, the missing group and its likely effect on denominator and cohort size should be stated rather than absorbed into a favourable average.[REF-05] [REF-16]

Comparability of denominator and cohort size should be tested across definition, coverage, classification, precision and time. For decisions about population censuses and civil registration, for population censuses and civil registration, a common label does not cure different stage structures, thresholds, questions or observation years. For population censuses and civil registration, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining population censuses and civil registration, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-10] [REF-12] [REF-16]

Responsibility completes population censuses and civil registration. For population censuses and civil registration, a material gap in the population baseline source should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining population censuses and civil registration, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on population censuses and civil registration, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-05] [REF-10] [REF-12] [REF-16]

23

Finance and workforce systems

Finance and workforce systems establishes the baseline question for the resource baseline source. For finance and workforce systems, the relevant population or responsible units are budgets, institutions, teachers and learners receiving resources, and the direct evidence concerns expenditure and staffing. In examining finance and workforce systems, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on finance and workforce systems, the first number found is not necessarily the best starting point. For comparison of finance and workforce systems, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-03] [REF-05] [REF-17] [REF-18]

The principal error risk in finance and workforce systems is that allocations or authorised posts replace expenditure received and staff actually teaching. In examining finance and workforce systems, the error changes the apparent starting position and can distort every later comparison. Within evidence on finance and workforce systems, authorities should list routes by which people, institutions or events enter and leave the resource baseline source, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of finance and workforce systems, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-03] [REF-05]

The minimum baseline requirement for finance and workforce systems is to trace release, receipt, personnel presence, qualification and learner exposure. Within evidence on finance and workforce systems, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of finance and workforce systems, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting finance and workforce systems, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-05] [REF-17]

Source fitness for the resource baseline source depends on who and what can be observed. For comparison of finance and workforce systems, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting finance and workforce systems, for finance and workforce systems, each source should retain its actual date and principal error. For decisions about finance and workforce systems, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-17] [REF-18]

Equity is a baseline condition within finance and workforce systems. In interpreting finance and workforce systems, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about finance and workforce systems, each group level and population share should remain visible beside any gap or parity measure. For finance and workforce systems, intersections require sufficient precision and safe disclosure. In examining finance and workforce systems, if part of budgets, institutions, teachers and learners receiving resources cannot be observed, the missing group and its likely effect on expenditure and staffing should be stated rather than absorbed into a favourable average.[REF-03] [REF-18]

Comparability of expenditure and staffing should be tested across definition, coverage, classification, precision and time. For decisions about finance and workforce systems, for finance and workforce systems, a common label does not cure different stage structures, thresholds, questions or observation years. For finance and workforce systems, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining finance and workforce systems, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-05] [REF-17] [REF-18]

Responsibility completes finance and workforce systems. For finance and workforce systems, a material gap in the resource baseline source should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining finance and workforce systems, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on finance and workforce systems, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-03] [REF-05] [REF-17] [REF-18]

24

Cross-source reconciliation

Cross-source reconciliation establishes the baseline question for the source correspondence. For cross-source reconciliation, the relevant population or responsible units are values carrying similar indicator labels, and the direct evidence concerns conceptual and numerical consistency. In examining cross-source reconciliation, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on cross-source reconciliation, the first number found is not necessarily the best starting point. For comparison of cross-source reconciliation, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-05] [REF-06] [REF-12] [REF-19]

The principal error risk in cross-source reconciliation is that sources are averaged or linked despite different populations, dates and units. In examining cross-source reconciliation, the error changes the apparent starting position and can distort every later comparison. Within evidence on cross-source reconciliation, authorities should list routes by which people, institutions or events enter and leave the source correspondence, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of cross-source reconciliation, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-05] [REF-06]

The minimum baseline requirement for cross-source reconciliation is to map every definition, investigate divergence and link only with lawful minimum data. Within evidence on cross-source reconciliation, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of cross-source reconciliation, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting cross-source reconciliation, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-06] [REF-12]

Source fitness for the source correspondence depends on who and what can be observed. For comparison of cross-source reconciliation, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting cross-source reconciliation, for cross-source reconciliation, each source should retain its actual date and principal error. For decisions about cross-source reconciliation, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-12] [REF-19]

Equity is a baseline condition within cross-source reconciliation. In interpreting cross-source reconciliation, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about cross-source reconciliation, each group level and population share should remain visible beside any gap or parity measure. For cross-source reconciliation, intersections require sufficient precision and safe disclosure. In examining cross-source reconciliation, if part of values carrying similar indicator labels cannot be observed, the missing group and its likely effect on conceptual and numerical consistency should be stated rather than absorbed into a favourable average.[REF-05] [REF-19]

Comparability of conceptual and numerical consistency should be tested across definition, coverage, classification, precision and time. For decisions about cross-source reconciliation, for cross-source reconciliation, a common label does not cure different stage structures, thresholds, questions or observation years. For cross-source reconciliation, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining cross-source reconciliation, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-06] [REF-12] [REF-19]

Responsibility completes cross-source reconciliation. For cross-source reconciliation, a material gap in the source correspondence should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining cross-source reconciliation, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on cross-source reconciliation, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-05] [REF-06] [REF-12] [REF-19]

Part V

Tests of availability and comparability

25

Concept and definition correspondence

Concept and definition correspondence establishes the baseline question for the conceptual comparability. For concept and definition correspondence, the relevant population or responsible units are countries and sources reporting a common education concept, and the direct evidence concerns definition match. In examining concept and definition correspondence, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on concept and definition correspondence, the first number found is not necessarily the best starting point. For comparison of concept and definition correspondence, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-04] [REF-06] [REF-09] [REF-19]

The principal error risk in concept and definition correspondence is that the same label covers different ages, stages, events or programme boundaries. In examining concept and definition correspondence, the error changes the apparent starting position and can distort every later comparison. Within evidence on concept and definition correspondence, authorities should list routes by which people, institutions or events enter and leave the conceptual comparability, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of concept and definition correspondence, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-04] [REF-06]

The minimum baseline requirement for concept and definition correspondence is to publish numerator, denominator, reference period, classification and correspondence note. Within evidence on concept and definition correspondence, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of concept and definition correspondence, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting concept and definition correspondence, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-06] [REF-09]

Source fitness for the conceptual comparability depends on who and what can be observed. For comparison of concept and definition correspondence, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting concept and definition correspondence, for concept and definition correspondence, each source should retain its actual date and principal error. For decisions about concept and definition correspondence, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-09] [REF-19]

Equity is a baseline condition within concept and definition correspondence. In interpreting concept and definition correspondence, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about concept and definition correspondence, each group level and population share should remain visible beside any gap or parity measure. For concept and definition correspondence, intersections require sufficient precision and safe disclosure. In examining concept and definition correspondence, if part of countries and sources reporting a common education concept cannot be observed, the missing group and its likely effect on definition match should be stated rather than absorbed into a favourable average.[REF-04] [REF-19]

Comparability of definition match should be tested across definition, coverage, classification, precision and time. For decisions about concept and definition correspondence, for concept and definition correspondence, a common label does not cure different stage structures, thresholds, questions or observation years. For concept and definition correspondence, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining concept and definition correspondence, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-06] [REF-09] [REF-19]

Responsibility completes concept and definition correspondence. For concept and definition correspondence, a material gap in the conceptual comparability should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining concept and definition correspondence, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on concept and definition correspondence, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-04] [REF-06] [REF-09] [REF-19]

26

Coverage and missingness

Coverage and missingness establishes the baseline question for the coverage comparability. For coverage and missingness, the relevant population or responsible units are entitled populations and reporting units, and the direct evidence concerns observed share and omission. In examining coverage and missingness, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on coverage and missingness, the first number found is not necessarily the best starting point. For comparison of coverage and missingness, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-05] [REF-11] [REF-12] [REF-14]

The principal error risk in coverage and missingness is that partial institutional or geographic coverage is described as national and missingness is coded as zero. In examining coverage and missingness, the error changes the apparent starting position and can distort every later comparison. Within evidence on coverage and missingness, authorities should list routes by which people, institutions or events enter and leave the coverage comparability, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of coverage and missingness, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-05] [REF-11]

The minimum baseline requirement for coverage and missingness is to quantify unit and item completeness and profile concentrated omission. Within evidence on coverage and missingness, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of coverage and missingness, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting coverage and missingness, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-11] [REF-12]

Source fitness for the coverage comparability depends on who and what can be observed. For comparison of coverage and missingness, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting coverage and missingness, for coverage and missingness, each source should retain its actual date and principal error. For decisions about coverage and missingness, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-12] [REF-14]

Equity is a baseline condition within coverage and missingness. In interpreting coverage and missingness, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about coverage and missingness, each group level and population share should remain visible beside any gap or parity measure. For coverage and missingness, intersections require sufficient precision and safe disclosure. In examining coverage and missingness, if part of entitled populations and reporting units cannot be observed, the missing group and its likely effect on observed share and omission should be stated rather than absorbed into a favourable average.[REF-05] [REF-14]

Comparability of observed share and omission should be tested across definition, coverage, classification, precision and time. For decisions about coverage and missingness, for coverage and missingness, a common label does not cure different stage structures, thresholds, questions or observation years. For coverage and missingness, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining coverage and missingness, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-11] [REF-12] [REF-14]

Responsibility completes coverage and missingness. For coverage and missingness, a material gap in the coverage comparability should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining coverage and missingness, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on coverage and missingness, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-05] [REF-11] [REF-12] [REF-14]

27

Precision and statistical uncertainty

Precision and statistical uncertainty establishes the baseline question for the precision comparability. For precision and statistical uncertainty, the relevant population or responsible units are sampled groups and modelled estimates, and the direct evidence concerns interval and error. In examining precision and statistical uncertainty, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on precision and statistical uncertainty, the first number found is not necessarily the best starting point. For comparison of precision and statistical uncertainty, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-05] [REF-11] [REF-13] [REF-19]

The principal error risk in precision and statistical uncertainty is that small differences are ranked despite sampling or model uncertainty. In examining precision and statistical uncertainty, the error changes the apparent starting position and can distort every later comparison. Within evidence on precision and statistical uncertainty, authorities should list routes by which people, institutions or events enter and leave the precision comparability, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of precision and statistical uncertainty, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-05] [REF-11]

The minimum baseline requirement for precision and statistical uncertainty is to publish intervals, effective samples, assumptions and limits on ranking. Within evidence on precision and statistical uncertainty, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of precision and statistical uncertainty, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting precision and statistical uncertainty, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-11] [REF-13]

Source fitness for the precision comparability depends on who and what can be observed. For comparison of precision and statistical uncertainty, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting precision and statistical uncertainty, for precision and statistical uncertainty, each source should retain its actual date and principal error. For decisions about precision and statistical uncertainty, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-13] [REF-19]

Equity is a baseline condition within precision and statistical uncertainty. In interpreting precision and statistical uncertainty, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about precision and statistical uncertainty, each group level and population share should remain visible beside any gap or parity measure. For precision and statistical uncertainty, intersections require sufficient precision and safe disclosure. In examining precision and statistical uncertainty, if part of sampled groups and modelled estimates cannot be observed, the missing group and its likely effect on interval and error should be stated rather than absorbed into a favourable average.[REF-05] [REF-19]

Comparability of interval and error should be tested across definition, coverage, classification, precision and time. For decisions about precision and statistical uncertainty, for precision and statistical uncertainty, a common label does not cure different stage structures, thresholds, questions or observation years. For precision and statistical uncertainty, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining precision and statistical uncertainty, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-11] [REF-13] [REF-19]

Responsibility completes precision and statistical uncertainty. For precision and statistical uncertainty, a material gap in the precision comparability should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining precision and statistical uncertainty, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on precision and statistical uncertainty, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-05] [REF-11] [REF-13] [REF-19]

28

Classification and threshold alignment

Classification and threshold alignment establishes the baseline question for the classification comparability. For classification and threshold alignment, the relevant population or responsible units are programmes, skills and proficiency categories, and the direct evidence concerns category correspondence. In examining classification and threshold alignment, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on classification and threshold alignment, the first number found is not necessarily the best starting point. For comparison of classification and threshold alignment, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-04] [REF-06] [REF-09] [REF-17]

The principal error risk in classification and threshold alignment is that national standards are renamed as common thresholds without evidence of equivalence. In examining classification and threshold alignment, the error changes the apparent starting position and can distort every later comparison. Within evidence on classification and threshold alignment, authorities should list routes by which people, institutions or events enter and leave the classification comparability, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of classification and threshold alignment, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-04] [REF-06]

The minimum baseline requirement for classification and threshold alignment is to use documented mapping, empirical linking where justified and visible non-equivalence. Within evidence on classification and threshold alignment, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of classification and threshold alignment, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting classification and threshold alignment, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-06] [REF-09]

Source fitness for the classification comparability depends on who and what can be observed. For comparison of classification and threshold alignment, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting classification and threshold alignment, for classification and threshold alignment, each source should retain its actual date and principal error. For decisions about classification and threshold alignment, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-09] [REF-17]

Equity is a baseline condition within classification and threshold alignment. In interpreting classification and threshold alignment, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about classification and threshold alignment, each group level and population share should remain visible beside any gap or parity measure. For classification and threshold alignment, intersections require sufficient precision and safe disclosure. In examining classification and threshold alignment, if part of programmes, skills and proficiency categories cannot be observed, the missing group and its likely effect on category correspondence should be stated rather than absorbed into a favourable average.[REF-04] [REF-17]

Comparability of category correspondence should be tested across definition, coverage, classification, precision and time. For decisions about classification and threshold alignment, for classification and threshold alignment, a common label does not cure different stage structures, thresholds, questions or observation years. For classification and threshold alignment, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining classification and threshold alignment, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-06] [REF-09] [REF-17]

Responsibility completes classification and threshold alignment. For classification and threshold alignment, a material gap in the classification comparability should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining classification and threshold alignment, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on classification and threshold alignment, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-04] [REF-06] [REF-09] [REF-17]

29

Timeliness and mixed-year baselines

Timeliness and mixed-year baselines establishes the baseline question for the temporal comparability. For timeliness and mixed-year baselines, the relevant population or responsible units are values observed at different dates, and the direct evidence concerns reference year and lag. In examining timeliness and mixed-year baselines, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on timeliness and mixed-year baselines, the first number found is not necessarily the best starting point. For comparison of timeliness and mixed-year baselines, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-05] [REF-06] [REF-12] [REF-18]

The principal error risk in timeliness and mixed-year baselines is that country rankings compare different years or combine target dimensions from unrelated periods. In examining timeliness and mixed-year baselines, the error changes the apparent starting position and can distort every later comparison. Within evidence on timeliness and mixed-year baselines, authorities should list routes by which people, institutions or events enter and leave the temporal comparability, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of timeliness and mixed-year baselines, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-05] [REF-06]

The minimum baseline requirement for timeliness and mixed-year baselines is to show actual dates and avoid a single baseline year label where observation windows differ. Within evidence on timeliness and mixed-year baselines, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of timeliness and mixed-year baselines, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting timeliness and mixed-year baselines, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-06] [REF-12]

Source fitness for the temporal comparability depends on who and what can be observed. For comparison of timeliness and mixed-year baselines, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting timeliness and mixed-year baselines, for timeliness and mixed-year baselines, each source should retain its actual date and principal error. For decisions about timeliness and mixed-year baselines, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-12] [REF-18]

Equity is a baseline condition within timeliness and mixed-year baselines. In interpreting timeliness and mixed-year baselines, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about timeliness and mixed-year baselines, each group level and population share should remain visible beside any gap or parity measure. For timeliness and mixed-year baselines, intersections require sufficient precision and safe disclosure. In examining timeliness and mixed-year baselines, if part of values observed at different dates cannot be observed, the missing group and its likely effect on reference year and lag should be stated rather than absorbed into a favourable average.[REF-05] [REF-18]

Comparability of reference year and lag should be tested across definition, coverage, classification, precision and time. For decisions about timeliness and mixed-year baselines, for timeliness and mixed-year baselines, a common label does not cure different stage structures, thresholds, questions or observation years. For timeliness and mixed-year baselines, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining timeliness and mixed-year baselines, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-06] [REF-12] [REF-18]

Responsibility completes timeliness and mixed-year baselines. For timeliness and mixed-year baselines, a material gap in the temporal comparability should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining timeliness and mixed-year baselines, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on timeliness and mixed-year baselines, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-05] [REF-06] [REF-12] [REF-18]

30

Confidentiality and safe disaggregation

Confidentiality and safe disaggregation establishes the baseline question for the disclosure comparability. For confidentiality and safe disaggregation, the relevant population or responsible units are small and sensitive groups, and the direct evidence concerns safe public detail. In examining confidentiality and safe disaggregation, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on confidentiality and safe disaggregation, the first number found is not necessarily the best starting point. For comparison of confidentiality and safe disaggregation, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-12] [REF-13] [REF-15] [REF-20]

The principal error risk in confidentiality and safe disaggregation is that greater visibility exposes identity or suppression removes disadvantaged populations from policy. In examining confidentiality and safe disaggregation, the error changes the apparent starting position and can distort every later comparison. Within evidence on confidentiality and safe disaggregation, authorities should list routes by which people, institutions or events enter and leave the disclosure comparability, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of confidentiality and safe disaggregation, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-12] [REF-13]

The minimum baseline requirement for confidentiality and safe disaggregation is to apply necessity, minimum detail, controlled access and publish the evidence gap. Within evidence on confidentiality and safe disaggregation, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of confidentiality and safe disaggregation, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting confidentiality and safe disaggregation, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-13] [REF-15]

Source fitness for the disclosure comparability depends on who and what can be observed. For comparison of confidentiality and safe disaggregation, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting confidentiality and safe disaggregation, for confidentiality and safe disaggregation, each source should retain its actual date and principal error. For decisions about confidentiality and safe disaggregation, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-15] [REF-20]

Equity is a baseline condition within confidentiality and safe disaggregation. In interpreting confidentiality and safe disaggregation, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about confidentiality and safe disaggregation, each group level and population share should remain visible beside any gap or parity measure. For confidentiality and safe disaggregation, intersections require sufficient precision and safe disclosure. In examining confidentiality and safe disaggregation, if part of small and sensitive groups cannot be observed, the missing group and its likely effect on safe public detail should be stated rather than absorbed into a favourable average.[REF-12] [REF-20]

Comparability of safe public detail should be tested across definition, coverage, classification, precision and time. For decisions about confidentiality and safe disaggregation, for confidentiality and safe disaggregation, a common label does not cure different stage structures, thresholds, questions or observation years. For confidentiality and safe disaggregation, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining confidentiality and safe disaggregation, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-13] [REF-15] [REF-20]

Responsibility completes confidentiality and safe disaggregation. For confidentiality and safe disaggregation, a material gap in the disclosure comparability should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining confidentiality and safe disaggregation, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on confidentiality and safe disaggregation, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-12] [REF-13] [REF-15] [REF-20]

Part VI

A public national baseline statement

31

National target-to-source inventory

National target-to-source inventory establishes the baseline question for the baseline inventory. For national target-to-source inventory, the relevant population or responsible units are institutions holding candidate evidence, and the direct evidence concerns measure, source and owner map. In examining national target-to-source inventory, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on national target-to-source inventory, the first number found is not necessarily the best starting point. For comparison of national target-to-source inventory, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-03] [REF-05] [REF-06] [REF-12]

The principal error risk in national target-to-source inventory is that a list of datasets omits concept, coverage, date, quality and recurring cost. In examining national target-to-source inventory, the error changes the apparent starting position and can distort every later comparison. Within evidence on national target-to-source inventory, authorities should list routes by which people, institutions or events enter and leave the baseline inventory, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of national target-to-source inventory, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-03] [REF-05]

The minimum baseline requirement for national target-to-source inventory is to record every target component with definition, source, groups, quality, owner and finance. Within evidence on national target-to-source inventory, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of national target-to-source inventory, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting national target-to-source inventory, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-05] [REF-06]

Source fitness for the baseline inventory depends on who and what can be observed. For comparison of national target-to-source inventory, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting national target-to-source inventory, for national target-to-source inventory, each source should retain its actual date and principal error. For decisions about national target-to-source inventory, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-06] [REF-12]

Equity is a baseline condition within national target-to-source inventory. In interpreting national target-to-source inventory, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about national target-to-source inventory, each group level and population share should remain visible beside any gap or parity measure. For national target-to-source inventory, intersections require sufficient precision and safe disclosure. In examining national target-to-source inventory, if part of institutions holding candidate evidence cannot be observed, the missing group and its likely effect on measure, source and owner map should be stated rather than absorbed into a favourable average.[REF-03] [REF-12]

Comparability of measure, source and owner map should be tested across definition, coverage, classification, precision and time. For decisions about national target-to-source inventory, for national target-to-source inventory, a common label does not cure different stage structures, thresholds, questions or observation years. For national target-to-source inventory, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining national target-to-source inventory, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-05] [REF-06] [REF-12]

Responsibility completes national target-to-source inventory. For national target-to-source inventory, a material gap in the baseline inventory should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining national target-to-source inventory, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on national target-to-source inventory, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-03] [REF-05] [REF-06] [REF-12]

32

Selecting the best available value

Selecting the best available value establishes the baseline question for the value-selection rule. For selecting the best available value, the relevant population or responsible units are candidate national and international observations, and the direct evidence concerns fitness and decision use. In examining selecting the best available value, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on selecting the best available value, the first number found is not necessarily the best starting point. For comparison of selecting the best available value, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-05] [REF-06] [REF-12] [REF-19]

The principal error risk in selecting the best available value is that the newest or most favourable value is selected without a pre-stated criterion. In examining selecting the best available value, the error changes the apparent starting position and can distort every later comparison. Within evidence on selecting the best available value, authorities should list routes by which people, institutions or events enter and leave the value-selection rule, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of selecting the best available value, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-05] [REF-06]

The minimum baseline requirement for selecting the best available value is to rank conceptual fit, coverage, quality, date and disaggregation and record rejection reasons. Within evidence on selecting the best available value, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of selecting the best available value, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting selecting the best available value, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-06] [REF-12]

Source fitness for the value-selection rule depends on who and what can be observed. For comparison of selecting the best available value, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting selecting the best available value, for selecting the best available value, each source should retain its actual date and principal error. For decisions about selecting the best available value, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-12] [REF-19]

Equity is a baseline condition within selecting the best available value. In interpreting selecting the best available value, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about selecting the best available value, each group level and population share should remain visible beside any gap or parity measure. For selecting the best available value, intersections require sufficient precision and safe disclosure. In examining selecting the best available value, if part of candidate national and international observations cannot be observed, the missing group and its likely effect on fitness and decision use should be stated rather than absorbed into a favourable average.[REF-05] [REF-19]

Comparability of fitness and decision use should be tested across definition, coverage, classification, precision and time. For decisions about selecting the best available value, for selecting the best available value, a common label does not cure different stage structures, thresholds, questions or observation years. For selecting the best available value, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining selecting the best available value, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-06] [REF-12] [REF-19]

Responsibility completes selecting the best available value. For selecting the best available value, a material gap in the value-selection rule should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining selecting the best available value, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on selecting the best available value, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-05] [REF-06] [REF-12] [REF-19]

33

Reporting unavailable and partial baselines

Reporting unavailable and partial baselines establishes the baseline question for the gap statement. For reporting unavailable and partial baselines, the relevant population or responsible units are targets and groups without fit evidence, and the direct evidence concerns unavailable, partial or estimated status. In examining reporting unavailable and partial baselines, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on reporting unavailable and partial baselines, the first number found is not necessarily the best starting point. For comparison of reporting unavailable and partial baselines, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-03] [REF-05] [REF-13] [REF-14]

The principal error risk in reporting unavailable and partial baselines is that pressure for completeness produces false precision or hides omitted populations. In examining reporting unavailable and partial baselines, the error changes the apparent starting position and can distort every later comparison. Within evidence on reporting unavailable and partial baselines, authorities should list routes by which people, institutions or events enter and leave the gap statement, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of reporting unavailable and partial baselines, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-03] [REF-05]

The minimum baseline requirement for reporting unavailable and partial baselines is to use explicit status, reason, affected population, interim evidence and funded next step. Within evidence on reporting unavailable and partial baselines, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of reporting unavailable and partial baselines, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting reporting unavailable and partial baselines, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-05] [REF-13]

Source fitness for the gap statement depends on who and what can be observed. For comparison of reporting unavailable and partial baselines, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting reporting unavailable and partial baselines, for reporting unavailable and partial baselines, each source should retain its actual date and principal error. For decisions about reporting unavailable and partial baselines, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-13] [REF-14]

Equity is a baseline condition within reporting unavailable and partial baselines. In interpreting reporting unavailable and partial baselines, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about reporting unavailable and partial baselines, each group level and population share should remain visible beside any gap or parity measure. For reporting unavailable and partial baselines, intersections require sufficient precision and safe disclosure. In examining reporting unavailable and partial baselines, if part of targets and groups without fit evidence cannot be observed, the missing group and its likely effect on unavailable, partial or estimated status should be stated rather than absorbed into a favourable average.[REF-03] [REF-14]

Comparability of unavailable, partial or estimated status should be tested across definition, coverage, classification, precision and time. For decisions about reporting unavailable and partial baselines, for reporting unavailable and partial baselines, a common label does not cure different stage structures, thresholds, questions or observation years. For reporting unavailable and partial baselines, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining reporting unavailable and partial baselines, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-05] [REF-13] [REF-14]

Responsibility completes reporting unavailable and partial baselines. For reporting unavailable and partial baselines, a material gap in the gap statement should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining reporting unavailable and partial baselines, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on reporting unavailable and partial baselines, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-03] [REF-05] [REF-13] [REF-14]

34

National and international reconciliation

National and international reconciliation establishes the baseline question for the reconciliation statement. For national and international reconciliation, the relevant population or responsible units are national authorities and international custodians, and the direct evidence concerns reported, adjusted and estimated value. In examining national and international reconciliation, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on national and international reconciliation, the first number found is not necessarily the best starting point. For comparison of national and international reconciliation, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-04] [REF-05] [REF-12] [REF-19]

The principal error risk in national and international reconciliation is that users cannot tell why national and international values differ. In examining national and international reconciliation, the error changes the apparent starting position and can distort every later comparison. Within evidence on national and international reconciliation, authorities should list routes by which people, institutions or events enter and leave the reconciliation statement, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of national and international reconciliation, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-04] [REF-05]

The minimum baseline requirement for national and international reconciliation is to publish source lineage, adjustments, uncertainty and a route to resolve material differences. Within evidence on national and international reconciliation, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of national and international reconciliation, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting national and international reconciliation, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-05] [REF-12]

Source fitness for the reconciliation statement depends on who and what can be observed. For comparison of national and international reconciliation, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting national and international reconciliation, for national and international reconciliation, each source should retain its actual date and principal error. For decisions about national and international reconciliation, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-12] [REF-19]

Equity is a baseline condition within national and international reconciliation. In interpreting national and international reconciliation, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about national and international reconciliation, each group level and population share should remain visible beside any gap or parity measure. For national and international reconciliation, intersections require sufficient precision and safe disclosure. In examining national and international reconciliation, if part of national authorities and international custodians cannot be observed, the missing group and its likely effect on reported, adjusted and estimated value should be stated rather than absorbed into a favourable average.[REF-04] [REF-19]

Comparability of reported, adjusted and estimated value should be tested across definition, coverage, classification, precision and time. For decisions about national and international reconciliation, for national and international reconciliation, a common label does not cure different stage structures, thresholds, questions or observation years. For national and international reconciliation, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining national and international reconciliation, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-05] [REF-12] [REF-19]

Responsibility completes national and international reconciliation. For national and international reconciliation, a material gap in the reconciliation statement should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining national and international reconciliation, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on national and international reconciliation, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-04] [REF-05] [REF-12] [REF-19]

35

Financing recurrent baseline improvement

Financing recurrent baseline improvement establishes the baseline question for the statistical capability plan. For financing recurrent baseline improvement, the relevant population or responsible units are agencies maintaining records, surveys and assessments, and the direct evidence concerns recurrent cost and milestone. In examining financing recurrent baseline improvement, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on financing recurrent baseline improvement, the first number found is not necessarily the best starting point. For comparison of financing recurrent baseline improvement, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-03] [REF-05] [REF-21] [REF-22]

The principal error risk in financing recurrent baseline improvement is that one-time support creates a value that cannot be updated or disaggregated. In examining financing recurrent baseline improvement, the error changes the apparent starting position and can distort every later comparison. Within evidence on financing recurrent baseline improvement, authorities should list routes by which people, institutions or events enter and leave the statistical capability plan, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of financing recurrent baseline improvement, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-03] [REF-05]

The minimum baseline requirement for financing recurrent baseline improvement is to cost staff, fieldwork, systems, quality, release, protection and later rounds. Within evidence on financing recurrent baseline improvement, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of financing recurrent baseline improvement, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting financing recurrent baseline improvement, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-05] [REF-21]

Source fitness for the statistical capability plan depends on who and what can be observed. For comparison of financing recurrent baseline improvement, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting financing recurrent baseline improvement, for financing recurrent baseline improvement, each source should retain its actual date and principal error. For decisions about financing recurrent baseline improvement, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-21] [REF-22]

Equity is a baseline condition within financing recurrent baseline improvement. In interpreting financing recurrent baseline improvement, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about financing recurrent baseline improvement, each group level and population share should remain visible beside any gap or parity measure. For financing recurrent baseline improvement, intersections require sufficient precision and safe disclosure. In examining financing recurrent baseline improvement, if part of agencies maintaining records, surveys and assessments cannot be observed, the missing group and its likely effect on recurrent cost and milestone should be stated rather than absorbed into a favourable average.[REF-03] [REF-22]

Comparability of recurrent cost and milestone should be tested across definition, coverage, classification, precision and time. For decisions about financing recurrent baseline improvement, for financing recurrent baseline improvement, a common label does not cure different stage structures, thresholds, questions or observation years. For financing recurrent baseline improvement, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining financing recurrent baseline improvement, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-05] [REF-21] [REF-22]

Responsibility completes financing recurrent baseline improvement. For financing recurrent baseline improvement, a material gap in the statistical capability plan should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining financing recurrent baseline improvement, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on financing recurrent baseline improvement, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-03] [REF-05] [REF-21] [REF-22]

36

Release, revision and public correction

Release, revision and public correction establishes the baseline question for the public baseline governance. For release, revision and public correction, the relevant population or responsible units are people represented in and using Goal 4 evidence, and the direct evidence concerns publication, version and remedy. In examining release, revision and public correction, baseline selection should begin with the policy claim and observation unit, then identify a value that fits them. Within evidence on release, revision and public correction, the first number found is not necessarily the best starting point. For comparison of release, revision and public correction, a national statement should make clear what the measure represents, which target component it can inform and which conclusions remain beyond its scope at 27 June 2016.[REF-05] [REF-12] [REF-13] [REF-20]

The principal error risk in release, revision and public correction is that baseline methods or values change silently and misuse remains uncorrected. In examining release, revision and public correction, the error changes the apparent starting position and can distort every later comparison. Within evidence on release, revision and public correction, authorities should list routes by which people, institutions or events enter and leave the public baseline governance, determine whether omission is concentrated by group or place, and separate missing evidence from a true zero. For comparison of release, revision and public correction, a national label should not be used where the source has only partial institutional, territorial or population coverage.[REF-05] [REF-12]

The minimum baseline requirement for release, revision and public correction is to publish calendar, metadata, revisions, correction, complaint and responsible authority. Within evidence on release, revision and public correction, the specification should include numerator, denominator, age or stage, reference date, classification, source, coverage, quality and selected disaggregation. For comparison of release, revision and public correction, it should name the national owner and indicate whether the value is reported, adjusted, estimated, partial or unavailable. In interpreting release, revision and public correction, if the measure corresponds only approximately to an internationally proposed indicator, the difference should remain visible instead of being erased by a common title.[REF-12] [REF-13]

Source fitness for the public baseline governance depends on who and what can be observed. For comparison of release, revision and public correction, administrative records can describe registered learners and services, household evidence can include people outside institutions, assessments can describe achievement for a stated target population, and population sources can strengthen denominators. In interpreting release, revision and public correction, for release, revision and public correction, each source should retain its actual date and principal error. For decisions about release, revision and public correction, divergence should be investigated as evidence of different concepts, coverage or quality before any combination is attempted.[REF-13] [REF-20]

Equity is a baseline condition within release, revision and public correction. In interpreting release, revision and public correction, the relevant groups may include sex, household resources, location, disability, language, migration, displacement or another nationally material characteristic. For decisions about release, revision and public correction, each group level and population share should remain visible beside any gap or parity measure. For release, revision and public correction, intersections require sufficient precision and safe disclosure. In examining release, revision and public correction, if part of people represented in and using Goal 4 evidence cannot be observed, the missing group and its likely effect on publication, version and remedy should be stated rather than absorbed into a favourable average.[REF-05] [REF-20]

Comparability of publication, version and remedy should be tested across definition, coverage, classification, precision and time. For decisions about release, revision and public correction, for release, revision and public correction, a common label does not cure different stage structures, thresholds, questions or observation years. For release, revision and public correction, metadata should support the precise comparison being made, and uncertainty should limit ranking and causal language. In examining release, revision and public correction, because arrangements were still being refined at the cutoff, reports should preserve provisional status and avoid reading later indicator decisions or results backwards into the baseline.[REF-12] [REF-13] [REF-20]

Responsibility completes release, revision and public correction. For release, revision and public correction, a material gap in the public baseline governance should lead to an owned improvement in records, surveys, assessments, population evidence, disaggregation or public documentation, with finance and a date. In examining release, revision and public correction, if no immediate collection is justified, the reason and available interim evidence should be published. Within evidence on release, revision and public correction, the baseline should finish as a transparent chain of definition, source, value, uncertainty, owner and revision rather than an unexplained figure that future users cannot reproduce.[REF-05] [REF-12] [REF-13] [REF-20]

References

  1. REF-01

    United Nations General Assembly. Transforming Our World: The 2030 Agenda for Sustainable Development. 2015.

    Adopted Goal 4, targets, universality and the commitment to leave no one behind.

    https://undocs.org/A/RES/70/1
  2. REF-02

    World Education Forum 2015. Incheon Declaration: Education 2030 — Towards Inclusive and Equitable Quality Education and Lifelong Learning for All. 2015.

    Education 2030 vision for inclusion, equity, quality and lifelong learning.

    https://unesdoc.unesco.org/ark:/48223/pf0000233137
  3. REF-03

    World Education Forum 2015 and United Nations Educational, Scientific and Cultural Organization. Education 2030 Framework for Action. 2015.

    Implementation framework adopted in November 2015, including thematic monitoring directions.

    https://unesdoc.unesco.org/ark:/48223/pf0000245656
  4. REF-04

    Inter-Agency and Expert Group on Sustainable Development Goal Indicators. Report of the Inter-Agency and Expert Group on Sustainable Development Goal Indicators. 2016.

    Proposed global indicator framework presented to the Statistical Commission as an initial basis.

    https://undocs.org/E/CN.3/2016/2/Rev.1
  5. REF-05

    United Nations Statistical Commission. Report on the Forty-seventh Session. 2016.

    Commission agreement to the proposed global indicator framework as a practical starting point subject to refinement.

    https://undocs.org/E/2016/24
  6. REF-06

    UNESCO Institute for Statistics Technical Advisory Group. Proposal for Thematic Indicators to Monitor the Education 2030 Agenda. 2015.

    Pre-cutoff proposal for broader thematic education indicators and definitions.

    https://uis.unesco.org/sites/default/files/documents/proposal-for-thematic-indicators-to-monitor-the-education-2030-agenda-2015-en.pdf
  7. REF-07

    Education for All Global Monitoring Report Team. Education for All 2000–2015: Achievements and Challenges — EFA Global Monitoring Report 2015. 2015.

    Latest pre-cutoff global assessment of education levels, inequalities and data gaps.

    https://unesdoc.unesco.org/ark:/48223/pf0000232205
  8. REF-08

    UNESCO Institute for Statistics. Global Education Digest 2012: Opportunities Lost — The Impact of Grade Repetition and Early School Leaving. 2012.

    Comparable evidence on progression, repetition and early leaving.

    https://uis.unesco.org/sites/default/files/documents/global-education-digest-2012-opportunities-lost-the-impact-of-grade-repetition-and-early-school-leaving-en_0.pdf
  9. REF-09

    United Nations Educational, Scientific and Cultural Organization. International Standard Classification of Education: ISCED 2011. 2012.

    Common definitions for education programmes and attainment.

    https://uis.unesco.org/sites/default/files/documents/international-standard-classification-of-education-isced-2011-en.pdf
  10. REF-10

    UNESCO Institute for Statistics. Guide to the Analysis and Use of Household Survey and Census Education Data. 2004.

    Methods for education indicators from household and census sources.

    https://uis.unesco.org/sites/default/files/documents/guide-to-the-analysis-and-use-of-household-survey-and-census-education-data-en_0.pdf
  11. REF-11

    United Nations Statistics Division. Household Sample Surveys in Developing and Transition Countries. 2005.

    Guidance on frames, questions, sampling error, response, weighting and analysis.

    https://unstats.un.org/unsd/hhsurveys/sectiona_new.htm
  12. REF-12

    United Nations General Assembly. Fundamental Principles of Official Statistics. 2014.

    Global principles of relevance, professional methods, transparency, correction and confidentiality.

    https://undocs.org/A/RES/68/261
  13. REF-13

    Office of the United Nations High Commissioner for Human Rights. Human Rights Indicators: A Guide to Measurement and Implementation. 2012.

    Rights-sensitive indicator design, disaggregation and interpretation.

    https://www.ohchr.org/sites/default/files/Documents/Publications/Human_rights_indicators_en.pdf
  14. REF-14

    United Nations Children’s Fund. The State of the World’s Children 2014 in Numbers: Every Child Counts — Revealing Disparities, Advancing Children’s Rights. 2014.

    Evidence on disaggregation, unequal outcomes and statistical visibility.

    https://www.unicef.org/reports/state-worlds-children-2014
  15. REF-15

    United Nations General Assembly. Convention on the Rights of Persons with Disabilities. 2006.

    Inclusive education, equality, accessibility and disability-sensitive information.

    https://www.ohchr.org/en/instruments-mechanisms/instruments/convention-rights-persons-disabilities
  16. REF-16

    United Nations General Assembly. Convention on the Rights of the Child. 1989.

    Education, non-discrimination, development and participation obligations.

    https://www.ohchr.org/en/instruments-mechanisms/instruments/convention-rights-child
  17. REF-17

    Education for All Global Monitoring Report Team. Teaching and Learning: Achieving Quality for All — EFA Global Monitoring Report 2013/4. 2014.

    Evidence on unequal teaching, learning and education quality.

    https://unesdoc.unesco.org/ark:/48223/pf0000225660
  18. REF-18

    European Commission. Education and Training Monitor 2015. 2015.

    Contemporaneous European evidence on attainment, early leaving, inequality and investment.

    https://op.europa.eu/en/publication-detail/-/publication/818a1177-a61b-11e5-b528-01aa75ed71a1
  19. REF-19

    European Statistical System Committee. European Statistics Code of Practice. 2011.

    Institutional and statistical principles for trustworthy official evidence.

    https://ec.europa.eu/eurostat/web/quality/european-quality-standards/european-statistics-code-of-practice
  20. REF-20

    European Parliament and Council of the European Union. Regulation (EC) No 223/2009 on European Statistics. 2009.

    European requirements for independence, quality, confidentiality and dissemination.

    https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32009R0223
  21. REF-21

    United Nations Secretary-General. One Humanity: Shared Responsibility. 2016.

    Pre-summit analysis of humanitarian need, displacement and responsibility relevant to education baselines.

    https://undocs.org/A/70/709
  22. REF-22

    Education Cannot Wait. Education Cannot Wait: A Fund for Education in Emergencies. 2016.

    Launch material on education finance and evidence in emergencies.

    https://www.educationcannotwait.org/
  23. REF-23

    United Nations General Assembly. The Right to Education in Emergency Situations. 2010.

    Continuity, protection, inclusion and quality of education during emergencies.

    https://undocs.org/A/RES/64/290
  24. REF-24

    United Nations High Commissioner for Refugees. Global Trends 2014: World at War. 2015.

    Pre-cutoff displacement evidence illustrating populations often absent from national education sources.

    https://www.unhcr.org/statistics/country/556725e69/unhcr-global-trends-2014.html