Data and research analysis

Using administrative data to examine science learning outcomes

Data Research

Provides a disciplined basis for interpreting evidence on science learning outcomes, including material variation, missing information and revision risk.

The present attention to administrative data to examine science learning outcomes follows the PISA 2006 results released in December 2007 and requires a careful distinction between public commitment, institutional practice and demonstrated result. For the reported measure, the principal analytical task is to separate an observed difference from a conclusion about its cause. The scope should include every materially affected setting, with differences in location, programme, delivery mode and learner population kept visible. Central policy alone does not establish consistent operation across the declared scope.

The PISA 2006 cycle gives particular emphasis to science while also assessing reading and mathematics among 15-year-old students. The results provide a comparative account of performance and its distribution across participating systems. Interpretation should take account of the sampled population, uncertainty and contextual information; a system-level association does not establish the cause of an individual learner’s result or the effectiveness of a particular provider.

A proper review of the comparison should establish the intended outcome before selecting controls or indicators. The analysis of the evidence under review proceeds on the basis that where an indicator is used as a proxy, the relationship between the proxy and the underlying educational outcome should be stated and tested. The record should explain why the approach suits the affected context, how material departures are authorised and when review will occur.

The present position

The stated reference is PISA 2006 results released in December 2007. Interpretation should preserve the unit and population represented in the data collection. A national or international pattern may justify closer review of administrative data to examine science learning outcomes, but provider-level action requires evidence relating to the affected provision. Variation in population coverage, reference period or classification should accompany the reported comparison.

The quality significance of the evidence under review follows from a basic distinction between availability and effective provision. The analysis of the evidence under review proceeds on the basis that education indicators should support decisions by describing outcomes and variation with definitions and limitations that permit responsible interpretation. Review should cover the stages at which learners receive information, provision, assessment, support and remedy.

  • Analyse missing information, including material exceptions and unequal effects.
  • Disaggregate material results within a defined period and review the result.
  • Test comparability and retain evidence sufficient for independent review.
  • Define the decision the indicator will inform within a defined period and review the result.
  • Document numerator and denominator within a defined period and review the result.

Implications for education statistics and performance measurement

The technical issue within administrative data to examine science learning outcomes concerns the basis on which a conclusion is reached. Oversight of the analytical question should reflect the principle that a proxy is useful only where its relationship with the intended outcome is sufficiently understood. Participation, activity and expenditure may support learning, but none constitutes direct evidence of learning without an explicit and tested connection. Any condition preventing complete assurance should appear with the evidence on which the judgement relies.

A narrow control over the analytical question may create false assurance. In the present context, data revisions not carried through to published conclusions, changes in definition presented as changes in performance and proxy measures treated as direct outcomes may produce acceptable aggregate reporting while individual learners remain exposed to material disadvantage. Testing should include exceptions and adverse cases, not only routine or successful operation.

Relevant evidence for the matter examined will normally include triangulation with administrative and qualitative evidence, revision and comparability records, population and sampling information, uncertainty estimates where relevant, and coverage and missingness analysis. Evidence should be current for the reference period, attributable and representative of the conclusion's stated scope. Contradictory evidence should be investigated and resolved, not omitted from the record.

  • For which groups may it fail?
  • Is direct evidence available?
  • Would the decision change if the proxy were inaccurate?
  • What outcome does the proxy represent?
  • What evidence supports the relationship?

What should be examined

The review method for administrative data to examine science learning outcomes should be reproducible. For the comparison, the reviewer should state the construct to be measured, explain why the proxy is expected to represent it, test that relationship against direct evidence and identify circumstances in which the proxy may fail. Do not allow convenience to determine the measure. A competent reviewer should be able to follow the record from source selection to conclusion and exception handling.

The analytical record for the reported measure should state the research question, data source, unit of analysis, reference period, coverage, exclusions, treatment of missing values and principal limitations. Results should be reproducible from the retained data and method. Any causal explanation should be identified separately from descriptive findings and supported by an appropriate design.

Interpretation of the analytical question should avoid two errors: treating a formal commitment as proof of effect, and treating one adverse case as proof that every part of the system has failed. In reviewing the matter examined, measurement can reveal where outcomes differ; it does not by itself establish why they differ or which intervention will work. For the reported measure, association should not be presented as causation, and statistical significance should not be treated as evidence of educational importance without further analysis.

Records relating to the matter examined should preserve both the conclusion and its limits. New evidence should trigger a traceable correction and review of decisions materially affected by the earlier conclusion. The correction process should identify prior users and decisions where published information has had material effect.

Accountability for the analytical question should follow decision-making authority. Evidence of material risk should be placed before the body with authority to act, together with a traceable decision. Operational tasks may be delegated, but accountability for material effects on learners must remain identifiable.

No individual measure is sufficient to establish effective operation of the analytical question across the affected scope. A reasoned conclusion should reconcile the governing requirement, evidence of operation, learner outcomes and residual risk, and remain open to better evidence.