Data and research analysis

Public performance reporting: definitions and comparability

Data Research

Considers what the available data can establish about public performance reporting and identifies the limitations that should accompany any public conclusion.

The 2017 accountability agenda provides the immediate context for public performance reporting. For the evidence under review, the available evidence should be interpreted with close attention to definitions, population coverage, collection methods and the limits of comparison. A reliable review extends beyond the central process to material variation across programmes, sites, delivery arrangements and learner groups. Central policy alone does not establish consistent operation across the declared scope.

The governing expectation for the evidence under review should be capable of consistent application. A decision concerning the evidence under review should recognise that reported averages should be accompanied by sufficient distributional information to identify material differences between learner groups, locations and forms of provision. Terms governing eligibility, support, assessment, reporting or review should prevent materially different treatment without recorded justification.

The present position

The stated reference is the 2017 accountability agenda. Application to public performance reporting depends on evidence from the relevant jurisdiction or institution. Authorities and providers should distinguish established fact, policy expectation and matters left to institutional judgement. That distinction should remain visible in the decision record, public reporting and later review.

The quality significance of the evidence under review follows from a basic distinction between availability and effective provision. In reviewing the analytical question, education indicators should support decisions by describing outcomes and variation with definitions and limitations that permit responsible interpretation. Oversight should examine implementation throughout the learner journey, not only at entry or through one reported outcome.

  • Define the decision the indicator will inform within a defined period and review the result.
  • Analyse missing information within a defined period and review the result.
  • Report uncertainty and revisions before using it to determine a learner or provider outcome.
  • Avoid causal claims unsupported by the design and retain evidence sufficient for independent review.
  • Test comparability, including material exceptions and unequal effects.

The substantive quality question

The technical issue within public performance reporting concerns the basis on which a conclusion is reached. In reviewing the evidence under review, comparison requires more than the use of a common label. Definitions, reference periods, population coverage, institutional boundaries and collection practices must be sufficiently aligned for the observed difference to have a stable meaning. The decision record should distinguish the scope supported by evidence from any scope that remains unresolved.

The principal risks in relation to the analytical question are averages concealing distribution, incomplete coverage, small differences overstated, and data revisions not carried through to published conclusions. A weakness in one part of the control environment may obscure a related failure elsewhere. Review should follow the sequence of decisions and records rather than assess documents in isolation.

Readily available material should not define the enquiry if it cannot answer the relevant decision question. For the evidence under review, the most relevant material is likely to include disaggregated results, revision and comparability records, triangulation with administrative and qualitative evidence, and coverage and missingness analysis. No source should carry more weight than its coverage and reliability permit, and unresolved uncertainty should remain visible.

  • Do the reference periods align?
  • Is the remaining difference educationally material?
  • Has a classification changed?
  • Are the populations defined on the same basis?
  • Are exclusions and missing records comparable?

Basis for a reliable conclusion

For operational review of public performance reporting, authorities and providers should proceed in a defined sequence. A competent review of the comparison should prepare a comparability table before analysing results. Record common elements, material differences, breaks in series and the direction in which each limitation may affect the conclusion; do not rank systems where those limitations remain material. A finding must identify its evidential basis, reach and required response, without giving informal observations a status they do not have.

Publication of findings on the analytical question should distinguish observed values, estimates and interpretation. Revisions, breaks in series and changes in classification should be visible. Where disaggregation creates small or unstable groups, confidentiality and uncertainty should be managed without concealing a material disparity that requires further investigation.

Interpretation of the reported measure should avoid two errors: treating a formal commitment as proof of effect, and treating one adverse case as proof that every part of the system has failed. Oversight of the analytical question should reflect the principle that measurement can reveal where outcomes differ; it does not by itself establish why they differ or which intervention will work. The analysis of the reported measure proceeds on the basis that missing or delayed information may be patterned rather than random. Conclusions should account for the possibility that excluded learners or providers differ from those observed.

The assurance record for the comparison should retain the date of the evidence, the source responsible for it, the scope examined and the version of any instrument or definition applied. Traceable source and version information allow genuine improvement to be distinguished from administrative revision. Revision should not remove an earlier conclusion from the record where reliance has occurred.

Public reporting on the comparison should distinguish established fact, analytical judgement and planned action. A material change should not remove the earlier position from the evidential trail. If definitions, coverage or evidence alter an earlier conclusion, the reason should be stated so that revision is not mistaken for changed performance.

No individual measure is sufficient to establish effective operation of the evidence under review across the affected scope. Assurance should be based on the combined legal or policy basis, operating evidence and learner effect, not on one element alone.