Thematic Research Report

ICEQC-R-2006-10 — Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data

A global comparative indicator study of concepts, survey methods, direct assessment, coverage, linguistic context and responsible interpretation

Publication date
Research category
Data and Indicator Research
Report archetype
Comparative Indicator Study
Geographic scope
Global
Evidence cut-off date
Responsible body
ICEQC Research and Policy Directorate
International Council for Education Quality Certification

ICEQC-R-2006-10

Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data

A global comparative indicator study of concepts, survey methods, direct assessment, coverage, linguistic context and responsible interpretation

Publication date
Evidence cut-off date
Publication type
Thematic Research Report
Authoritative language
EN

Publication record

This is the controlled English edition. Evidence and institutional status are stated as at the evidence cut-off date.

Executive summary

Adult literacy statistics are indispensable to education policy and unusually difficult to interpret. A national rate may be drawn from a census question answered by the individual, a statement made by another household member, an assumption based on schooling, a short reading task or a broader assessment of performance across several kinds of text. These methods do not necessarily measure the same construct, cover the same population or produce comparable values.

The distinction is material. A self-reported response can describe how a person identifies or functions in a familiar setting, but it is affected by question wording, social expectations, language, privacy, proxy response and the threshold understood by interviewer and respondent. A direct assessment observes performance under specified conditions, but it is affected by task selection, translation, administration, sampling, non-response and the relationship between the assessment setting and adults’ everyday literacy practices.

Literacy is not adequately represented as a permanent binary possession. Contemporary international analysis describes a continuum of capabilities used for different purposes and in different social, linguistic and economic settings. Direct international assessments have demonstrated substantial distributions of prose, document, quantitative and numeracy proficiency within populations and within levels of educational attainment.

This does not render census and household-survey indicators without value. Censuses can provide broad population coverage and small-area information that a specialised assessment cannot ordinarily match. Household surveys can connect literacy responses to poverty, work, health, language, migration and participation. Their value depends on a clear account of what was asked, who answered, which languages were available, how schooling was treated and who was omitted.

The report establishes four classes of literacy evidence: self-declaration by the person concerned; proxy declaration by another household member; indirect classification through educational participation or attainment; and direct performance assessment. It also distinguishes a short verification task from a multi-domain assessment. No class is declared universally superior. Each answers a narrower set of questions and carries characteristic error.

Comparability requires more than applying the same label. Concept, population, age range, language, reference period, question, response categories, mode, respondent, task demands, scoring, sampling, weighting and non-response must be examined together. A time series can break when one of these changes. A cross-national table can contain accurately reported national values that remain unsuitable for direct ranking.

Participation is both an object of measurement and a condition of measurement quality. Adults with weaker literacy, unfamiliarity with the survey language, disability, insecure residence, long working hours, institutional residence or fear of official contact may be less likely to be reached or to complete an assessment. If their exclusion is not measured, the published distribution may overstate population proficiency.

The report proposes an indicator-use protocol rather than one universal literacy rate. Every published value should carry a method class, population, age range, languages, respondent rule, reference date, coverage statement and explicit interpretation boundary. Comparisons should be classified as strong, qualified, descriptive only or not supportable.

Where countries require both extensive coverage and stronger evidence of skill, a layered design is appropriate: a concise census or household module for breadth; direct assessment in a probability sample for proficiency; and a background questionnaire for language, education, literacy practices and participation. The contemporary development of the Literacy Assessment and Monitoring Programme reflects the need for more relevant and reliable evidence while building national statistical capacity.

The public interest is not served by replacing one uncertain estimate with an opaque assessment score. Responsible measurement must preserve the definition, sampling, uncertainty, accessibility, confidentiality and policy meaning of the result. Literacy statistics should identify barriers and inform provision; they should not stigmatise adults, languages or communities.

Key findings

  • “Literacy rate” is not a method-neutral term. Its meaning depends on the concept, question or task, population, respondent, language and threshold used.
  • Self-declaration, proxy declaration, schooling-based inference, short verification and multi-domain direct assessment measure overlapping but non-identical conditions.
  • A binary response may support broad monitoring but cannot describe the distribution of literacy practices and proficiency within the population.
  • Educational attainment is related to literacy but is not an adequate substitute for observed proficiency. Adults with the same formal level may have substantially different skills.
  • Direct assessment reduces reliance on perception and proxy reporting but introduces its own construct, task, translation, sampling, administration and participation requirements.
  • Proxy reporting should be separately identified. A household member may not know another adult’s reading and writing practices, particularly where ability is concealed or used outside the home.
  • Question wording can change the threshold. “Can read and write,” “can read a simple statement,” and “reads without difficulty” do not create one interchangeable indicator.
  • Language is part of measurement validity. Failure in an unavailable or unfamiliar language cannot be interpreted automatically as absence of literacy in all languages.
  • Survey participation is socially distributed. Coverage and non-response analysis should examine adults least likely to be reached or assessed, not only the achieved sample total.
  • Cross-national comparison requires method metadata beside the value. A common title or denominator is insufficient.
  • Trend analysis should identify breaks caused by changes in question, respondent rule, language, frame, age range, assumption based on schooling or direct-assessment method.
  • National aggregates should be accompanied by sex, age, location and other policy-relevant distributions where sample and confidentiality permit. The choice should respond to participation and provision questions.
  • Thresholds and proficiency levels are reporting devices. They should not conceal score distributions, uncertainty or variation near a cut point.
  • Census breadth, household-survey context and specialised assessment depth are complementary when linked through transparent definitions and appropriate samples.
  • Published literacy evidence should support educational provision and equal participation, not rank communities through measures that are not comparable.

Scope and method

The report examines adult literacy measurement as at 7 October 2006. Its principal population is persons aged 15 years and over, consistent with common international adult-literacy reporting, while recognising that national systems may use other age ranges. It addresses censuses, general household surveys, specialised literacy surveys and international comparative assessments.

The evidence base comprises the Dakar Framework and international Literacy Decade action; the 2006 global monitoring report on literacy; UNESCO analysis of literacy plurality and assessment; the UIS Literacy Assessment and Monitoring Programme; the International Adult Literacy Survey and Adult Literacy and Life Skills Survey; United Nations census and household-survey guidance; principles of official statistics; international education classification; and contemporary global monitoring.

The report compares methods, not countries. Numerical cases are constructed to demonstrate denominator, non-response, standard error, classification and trend decisions. They do not estimate any identified national population.

“Self-report” means an adult’s answer about that adult’s own literacy. “Proxy report” means an answer supplied by another person. “Indirect classification” means assignment from another characteristic, such as completed schooling. “Direct assessment” means performance on one or more specified tasks. “Comparability” means sufficient equivalence of concept, population and measurement to support the stated comparison; it does not require complete identity in every operational detail.

Part I

What an Adult Literacy Indicator Claims

1

Indicator purpose

feasibility and require, different, measures in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 1. Indicator purpose to connect evidence with delivery, unequal effect and correction. Broad population monitoring, local programme planning, identification of service barriers and evaluation of a learning programme may require different measures.

The method should be selected after the question. A readily available rate should not determine the purpose retrospectively.

2

Literacy as capability and practice

capacity and language, setting, opportunity in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 2. Literacy as capability and practice to connect evidence with delivery, unequal effect and correction. It is shaped by purpose, language, text, setting and opportunity.[REF-02] [REF-04]

A measure that observes one task should state the domain it represents. It should not claim every social and communicative dimension of literacy.

3

The binary convention

equity and levels, domains, categories in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 3. The binary convention to connect evidence with delivery, unequal effect and correction. The convention can provide a concise count but compresses different levels, domains and uses into two categories.

Movement near the operational threshold can alter the rate without representing a sharp division in adults’ capabilities.

4

A continuum of proficiency

timing and analysis, distribution, demand in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 4. A continuum of proficiency to connect evidence with delivery, unequal effect and correction. This allows analysis of distribution and task demand.[REF-06] [REF-07]

coverage and treated, substantively, discontinuous in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 4. A continuum of proficiency to connect evidence with delivery, unequal effect and correction. Adults within a level are not identical, and performance near adjacent cut points should not be treated as substantively discontinuous.

5

Reading and writing

uncertainty and observed, inference, justified in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 5. Reading and writing to connect evidence with delivery, unequal effect and correction. A reading task cannot establish writing performance unless writing is separately observed or the inference is justified.

Published metadata should identify whether reading, writing or both entered the classification.

6

Numeracy

comparability and distinct, assessed, domain in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 6. Numeracy to connect evidence with delivery, unequal effect and correction. The Adult Literacy and Life Skills Survey treats numeracy as a distinct assessed domain.[REF-07]

A combined label should not conceal different constructs or thresholds. Policy responses may differ where the principal difficulty concerns text, quantity or their interaction.

7

Documents and prose

authority and locating, integrating, information in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 7. Documents and prose to connect evidence with delivery, unequal effect and correction. Continuous text, forms, schedules, maps and tables may place different demands on locating, integrating and using information.[REF-06]

A short sentence-reading item provides little evidence about the use of complex documents, even where it is useful for a minimum verification purpose.

8

Functional interpretation

coverage and educational, cultural, settings in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 8. Functional interpretation to connect evidence with delivery, unequal effect and correction. The term requires specification because functions and demands differ across work, family, civic, educational and cultural settings.

It should not be used as an undefined higher threshold added to a basic rate.

9

Social and linguistic plurality

Adults may read and write in one language or script but not another. Literacy practices may be strongest in domains recognised locally but absent from the survey.[REF-04]

Measurement should identify the language and script permitted. A classification in one language is not automatically a classification across all languages.

10

Skills and opportunities for use

Observed proficiency reflects learning and opportunities to use skills. Adults may gain, maintain or lose fluency as demands and practices change.[REF-06] [REF-07]

Policy should therefore examine the environments that support literacy, not treat the measured result solely as an individual attribute.

11

Population specification

distribution and population, insufficient, metadata in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 11. Population specification to connect evidence with delivery, unequal effect and correction. “Adult population” is insufficient metadata.

timing and educational, opportunity, cohort in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 11. Population specification to connect evidence with delivery, unequal effect and correction. Values from different upper-age rules can diverge where proficiency and educational opportunity vary by cohort.

12

Reference period

feasibility and migration, household, change in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 12. Reference period to connect evidence with delivery, unequal effect and correction. Fieldwork extending across months may require a defined reference date and treatment of birthdays, migration and household change.

Modelled estimates and projections should be identified separately from observed results.

13

Unit of observation

capacity and access, replace, record in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 13. Unit of observation to connect evidence with delivery, unequal effect and correction. Household-level conditions may explain practice and access, but they do not replace the adult record.

Weights and denominators should correspond to persons, not interviewed households, when the reported result is an adult literacy rate.

14

Table 1: claim specification

Table 1. Minimum specification of an adult literacy indicator claim
Claim fieldRequired statementQuestion protectedMisinterpretation prevented
purposemonitoring, planning, diagnosis or evaluationwhy the value is producedavailable statistic determines policy question
constructreading, writing, numeracy, domain and usewhat capability is representedbroad literacy inferred from one task
classificationbinary, ordered category or scalehow performance becomes a resultthreshold treated as natural divide
populationage, residence and coverageto whom the result appliesunlike populations compared
language and scriptpermitted and administered formsin what language performance was observedlanguage mismatch treated as no literacy
respondentself, proxy or assessed personwho supplied the evidenceproxy answer treated as self-report
methodquestion, indirect rule or direct taskhow evidence was obtainedmethod-neutral “rate” assumed
periodfieldwork and reference datewhen the condition was observedpublication year treated as measurement year
uncertaintysampling and material non-sampling limitshow precisely result is knownpoint estimate treated as exact
use boundaryclaims not supportedwhere interpretation must stopindicator used for unsupported ranking

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

15

Public meaning

An indicator becomes public evidence when users can understand its scope and limitations. Technical complexity does not justify publication of an unexplained headline rate.

The public meaning should be accurate enough to protect policy and respectful enough to avoid presenting adults below a threshold as incapable of learning or participation.

Part II

Self-Reported and Proxy-Reported Literacy

16

Nature of self-report

remedy and questions, designed, purposes in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 16. Nature of self-report to connect evidence with delivery, unequal effect and correction. It can be rapid, inexpensive and suitable for large population instruments. It can also capture perceived difficulty or actual practice when questions are designed for those purposes.

It does not directly observe performance. Its validity depends on the relationship between the reported judgement and the construct used in policy.

17

Question wording

continuity and difficulty, provides, differentiation in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 17. Question wording to connect evidence with delivery, unequal effect and correction. ” leaves the language, material, understanding, ease and threshold to interpretation. A question about reading a simple message is narrower; a graded question about difficulty provides more differentiation.

Wording changes can alter responses even when underlying ability is unchanged. Exact question text should accompany comparisons.

18

Combined questions

distribution and understanding, expected, category in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 18. Combined questions to connect evidence with delivery, unequal effect and correction. Respondents may answer according to the stronger skill, the weaker skill or their understanding of the expected category.

Separate questions improve diagnostic value, though they do not remove self-evaluation error.

19

Response categories

feasibility and anchors, consistent, administration in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 19. Response categories to connect evidence with delivery, unequal effect and correction. Ordered categories such as easily, with difficulty and not at all can retain useful variation but require clear anchors and consistent administration.

An additional “unknown” or “not stated” category should not be combined with non-literacy. Its distribution may reveal respondent or fieldwork problems.

20

Social desirability

capacity and standard, produce, understatement in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 20. Social desirability to connect evidence with delivery, unequal effect and correction. In other settings, modesty or a demanding personal standard may produce understatement.

Privacy, interviewer conduct and the perceived consequence of an answer affect this error. It cannot be assumed constant between groups or countries.

21

Reference standard

equity and standard, formal, writing in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 21. Reference standard to connect evidence with delivery, unequal effect and correction. A person who manages familiar work documents may answer yes; another with similar performance may answer no because the understood standard is formal writing.

Comparability is weakened when respondents supply both the evidence and an unstated threshold.

22

Literacy practices as self-report

timing and practice, opportunity, maintenance in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 22. Literacy practices as self-report to connect evidence with delivery, unequal effect and correction.[REF-06] [REF-07]

Practice questions should state the period and purpose and allow for locally relevant materials. Frequency is not a substitute for quality or understanding.

23

Proxy reporting

uncertainty and uncertainty, knowledge, judgement in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 23. Proxy reporting to connect evidence with delivery, unequal effect and correction. Proxy reporting reduces field cost and permits completion when members are absent. It also introduces uncertainty about knowledge and judgement.

The dataset should identify whether the answer was self or proxy wherever operationally possible.

24

Knowledge within the household

comparability and stigma, household, relations in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 24. Knowledge within the household to connect evidence with delivery, unequal effect and correction. Ability may also be concealed because of stigma or household relations.

Proxy accuracy should not be presumed from relationship alone. Survey evaluation can compare self and proxy answers in a subsample.

25

Household hierarchy

authority and systematically, favour, particular in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 25. Household hierarchy to connect evidence with delivery, unequal effect and correction. Interview scheduling and respondent rules can systematically favour particular ages or sexes.

Field protocols should identify the most knowledgeable eligible respondent and record who answered each literacy item.

26

Inference from schooling

coverage and demonstrate, current, performance in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 26. Inference from schooling to connect evidence with delivery, unequal effect and correction. Educational attainment is associated with literacy but does not demonstrate current performance.[REF-06]

feasibility and convert, literacy, observations in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 26. Inference from schooling to connect evidence with delivery, unequal effect and correction. International classification helps describe education levels but does not convert them into literacy observations.[REF-10]

27

Assumed literacy and assumed illiteracy

remedy and illiteracy, misclassify, adults in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 27. Assumed literacy and assumed illiteracy to connect evidence with delivery, unequal effect and correction. Both can misclassify adults.

Every assumed classification should be separately identified in metadata and sensitivity analysis.

28

Non-response

continuity and estimates, biased, upward in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 28. Non-response to connect evidence with delivery, unequal effect and correction. If adults with weaker literacy are less likely to respond personally, complete-case estimates may be biased upward.

Substitution by proxy may improve coverage while changing the measurement method. The trade-off should be documented rather than hidden in an overall response rate.

29

Interviewer effects

distribution and procedures, treatment, uncertainty in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 29. Interviewer effects to connect evidence with delivery, unequal effect and correction. Training should address neutrality, exact wording, language procedures and treatment of uncertainty.[REF-11] [REF-12]

Unusually high or low literacy rates by interviewer may indicate assignment differences or field error and require controlled review.

30

Mode effects

feasibility and without, changing, meaning in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 30. Mode effects to connect evidence with delivery, unequal effect and correction. A self-completed literacy question can exclude the very respondent whose evidence is sought unless accessible assistance is provided without changing meaning.

Mode changes in a time series require evaluation and, where material, a break in comparability.

31

Language of interview

capacity and rather, literal, wording in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 31. Language of interview to connect evidence with delivery, unequal effect and correction. Translation should preserve the construct and response thresholds rather than literal wording alone.

The language actually used should be recorded, together with interpreter use and any cases in which an eligible adult could not be interviewed in an adequate language.

32

Table 2: self- and proxy-report risk register

Table 2. Self-reported and proxy-reported literacy risk register
RiskMechanismEvidence neededReporting response
unstated thresholdrespondent defines “literate”cognitive testing and question textrestrict interpretation to declared status
social desirabilitystigma or perceived benefit affects answerprivacy review and validation subsampledisclose probable direction and uncertainty
combined skillsreading and writing collapsedseparate-item comparisondo not infer domain-specific ability
proxy knowledgerespondent does not observe another adult’s practiceself–proxy re-interviewidentify proxy share and disagreement
schooling assumptionattainment substitutes for present skilldirect or self-report comparisonpublish assumed component separately
missing languageeligible adult cannot use interview languagelanguage non-interview countqualify coverage and improve provision
mode changeresponse process differs over timebridge study or parallel administrationclassify trend as qualified or broken
interviewer variationprobing or coding differsinterviewer-level quality reviewretrain, verify and correct where supported
non-responseparticipation related to literacyresponse by group and follow-up evidenceweight cautiously and retain residual bias

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

33

Appropriate uses

timing and relevant, service, access in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 33. Appropriate uses to connect evidence with delivery, unequal effect and correction. Graded questions can identify perceived difficulty relevant to service access.

It is weaker for estimating proficiency distributions, evaluating instruction or ranking populations with different languages and response conventions.

34

Interpretation boundary

A self-reported literacy rate is evidence of classification under the survey question and field conditions. It should not be described as an observed demonstration of literacy.

This boundary is not a dismissal. It is the condition for using the measure honestly and for deciding when complementary assessment is necessary.

Part III

Direct Assessment of Adult Literacy

35

What direct assessment adds

continuity and education, literacy, practices in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 35. What direct assessment adds to connect evidence with delivery, unequal effect and correction. It can distinguish levels and domains, examine distributions and relate performance to education, work and literacy practices.[REF-06] [REF-07]

The result remains conditional on the construct, tasks, language, administration and population reached. “Direct” does not mean complete or error-free.

36

Construct definition

distribution and rather, available, collection in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 36. Construct definition to connect evidence with delivery, unequal effect and correction. Content should follow the construct rather than an available collection of items.

A broad policy term such as functional literacy requires a narrower assessment statement before results can be interpreted.

37

Domain coverage

Prose, document, writing, numeracy and problem solving are related but distinct. An assessment may cover one or several.[REF-06] [REF-07]

The report title should not imply domains omitted by design. Composite reporting should retain domain-level evidence where policy conclusions differ.

38

Task authenticity

capacity and assessment, intends, measure in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 38. Task authenticity to connect evidence with delivery, unequal effect and correction. Familiar appearance does not alone establish validity; a task must require the capability the assessment intends to measure.

Everyday context can improve relevance while creating cultural or occupational differences that require review.

39

Task difficulty

equity and defensible, increases, demand in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 39. Task difficulty to connect evidence with delivery, unequal effect and correction. A progression should reflect defensible increases in demand.

Empirical performance assists calibration, but statistical difficulty does not by itself explain what capability a task requires.

40

Short verification tasks

timing and adults, simple, sentence in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 40. Short verification tasks to connect evidence with delivery, unequal effect and correction. It is feasible within a broader household survey and may identify adults able to read all, part or none of a simple sentence.

It does not measure writing, extended comprehension, documents or the full continuum of proficiency. Its economy depends on a strict interpretation boundary.

41

Multi-item assessment

uncertainty and capacity, respondent, cooperation in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 41. Multi-item assessment to connect evidence with delivery, unequal effect and correction. They can support scaled scores and proficiency distributions but require more time, technical capacity and respondent cooperation.

The burden should be considered in sample design and field protocol. Longer assessment is not automatically more valid if non-completion becomes selective.

42

Screening and routing

A screening stage may route adults to tasks suited to their demonstrated entry performance. Routing can reduce frustration and improve information at lower proficiency.

The rule, measurement consequence and treatment of non-attempts must be explicit. A screening failure should not create an unobserved group outside the reported distribution.

43

Item sampling

authority and scores, become, precise in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 43. Item sampling to connect evidence with delivery, unequal effect and correction. This expands content coverage while each person completes only part of the pool. Population proficiency can be estimated, but individual scores become less precise.

Published results should distinguish population inference from diagnostic claims about a named adult.

44

Translation and adaptation

coverage and convention, changes, difficulty in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 44. Translation and adaptation to connect evidence with delivery, unequal effect and correction. Literal equivalence may fail where grammar, script, word length or document convention changes difficulty.

Adaptation decisions require bilingual, subject and measurement review, field testing and documentation.

45

Multiple languages

remedy and comparisons, versions, differ in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 45. Multiple languages to connect evidence with delivery, unequal effect and correction. Each approach affects the population meaning. Choice can show capability in a preferred language but complicate comparisons where versions differ.

Results should record assessment language and avoid classifying performance in one language as literacy in every language.

46

Script and orthography

continuity and equivalence, psychometric, equivalence in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 46. Script and orthography to connect evidence with delivery, unequal effect and correction. Item design should not treat surface equivalence as psychometric equivalence.

Where oral and written language relationships differ, instructions and practice items require particular attention.

47

Cultural and contextual review

distribution and neither, possible, desirable in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 47. Cultural and contextual review to connect evidence with delivery, unequal effect and correction. Removal of every contextual feature is neither possible nor desirable.

The objective is to minimise irrelevant difficulty and document the context retained.

48

Administration conditions

feasibility and preserve, intended, comparison in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 48. Administration conditions to connect evidence with delivery, unequal effect and correction. Standardisation should be practical and sufficient to preserve the intended comparison.

Deviations and interrupted sessions should be recorded rather than coded automatically as low proficiency.

49

Interviewer and assessor competence

capacity and language, without, authorisation in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 49. Interviewer and assessor competence to connect evidence with delivery, unequal effect and correction. They should not teach the task, signal correctness or change language without authorisation.

Certification of field competence should be based on observed administration and correction, not attendance alone.

50

Accessibility

equity and changing, capability, measured in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 50. Accessibility to connect evidence with delivery, unequal effect and correction. An accommodation should remove irrelevant barriers without changing the capability being measured.

Exclusion from assessment should be reported by reason. An assessment that omits a population cannot support a whole-population claim without qualification.

51

Anxiety and stigma

timing and absence, individual, sanction in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 51. Anxiety and stigma to connect evidence with delivery, unequal effect and correction. Information should explain purpose, confidentiality, voluntary participation where applicable and absence of individual sanction.

Respectful administration is an ethical requirement and a data-quality condition.

52

Scoring

Scoring rules should define full, partial and incorrect responses, omissions and invalid administrations. Constructed responses require scorer training and reliability checks.

Changes made after fieldwork should be documented and applied consistently. Ambiguous items should be reviewed before population estimates are finalised.

53

Scaling

Scaled scores place performance from different item sets on a common continuum under a statistical model. Model fit, item behaviour and uncertainty require technical examination.

A scale is not a physical quantity with an absolute zero. Differences should be interpreted through task demand and sampling error.

54

Proficiency levels

Levels translate a score continuum into descriptions of tasks likely to be completed. They aid communication but create boundaries not present in the underlying scale.

Reports should provide distributions and standard errors and avoid describing all adults below a chosen level as having no literacy.

55

Plausible values and population use

coverage and method, estimates, variance in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 55. Plausible values and population use to connect evidence with delivery, unequal effect and correction. Users should follow the technical method for group estimates and variance.

Simple averaging of an individual-score field may understate uncertainty or produce biased relationships.

56

Quality control

Field monitoring, response checks, item analysis, scoring reliability, data cleaning and weighting should form one documented quality system.[REF-05] [REF-12]

Quality flags should lead to verification and defined treatment. Removal of inconvenient records without a rule can bias the distribution.

57

Table 3: direct-assessment validity record

Table 3. Direct-assessment validity and participation record
DomainRequired evidencePrincipal riskRelease limitation
constructframework and intended interpretationtasks narrower than claimname assessed domain only
contentitem specification and reviewirrelevant knowledge determines responsequalify or remove affected items
languagetranslation, adaptation and administration recordversion difficulty differstest equivalence and report language
samplingframe, selection and weightspopulation not representedrestrict population claim
participationcontact, response and completion by groupweakest adults under-representednon-response analysis required
accessibilityaccommodation and exclusion recorddisability becomes test barrierqualify coverage and redesign access
administrationtraining, observation and deviationsconditions alter performanceflag or exclude under fixed rule
scoringrules and agreement evidencescorer variationmoderate and rescore
scalingmodel, fit and uncertaintyfalse precisionreport standard error and diagnostics
interpretationlevel descriptions and boundariesbelow-level equated with no literacypublish continuum and task meaning

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

58

Appropriate uses

distribution and change, designs, comparable in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 58. Appropriate uses to connect evidence with delivery, unequal effect and correction.

It is not automatically suited to individual certification, diagnosis or high-stakes allocation. The sampling and measurement design governs permissible use.

59

Direct-assessment boundary

feasibility and qualitative, knowledge, barriers in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 59. Direct-assessment boundary to connect evidence with delivery, unequal effect and correction. It does not observe every literacy practice, language or context and should not displace qualitative knowledge about barriers and use.

Its public authority depends on transparent design, participation and uncertainty, not on technical complexity alone.

Part IV

Comparability Across Methods, Populations and Time

60

Comparability as a claim

comparability and inadequate, ranking, estimates in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 60. Comparability as a claim to connect evidence with delivery, unequal effect and correction. It is a judgement about whether they can support a particular contrast. Two values may be adequate for broad description but inadequate for ranking or small trend estimates.

The intended use should therefore accompany every comparability decision.

61

Concept equivalence

authority and identical, comprehension, sentence in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 61. Concept equivalence to connect evidence with delivery, unequal effect and correction. A binary declaration of ability to read and write is not conceptually identical to a scale of prose comprehension or a short sentence task.

They may be examined together as different evidence about literacy, but their numerical results should not be treated as interchangeable.

62

Population equivalence

coverage and cohorts, differ, materially in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 62. Population equivalence to connect evidence with delivery, unequal effect and correction. A survey of ages 16–65 should not be compared directly with a census rate for all persons aged 15 and over where older cohorts differ materially.

Recalculation to a common age range may improve comparability if microdata and weights permit.

63

Time equivalence

Literacy data are collected irregularly. A table labelled with one year may combine observations from different years or modelled values. The observation year should be prominent.

Economic, educational and cohort change can make distant observations unsuitable for a current comparison.

64

Method equivalence

Self, proxy, indirect and direct methods should be coded explicitly. A method change can alter a time series independently of real population change.

Bridge studies using parallel methods on the same sample can estimate the direction and magnitude of the discontinuity.

65

Question equivalence

Small wording differences can change difficulty and threshold. Comparability review should preserve exact wording, response categories, routing and interviewer instructions.

A translated label in an international table cannot demonstrate equivalent national questions.

66

Language equivalence

feasibility and permitting, written, language in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 66. Language equivalence to connect evidence with delivery, unequal effect and correction. A national rate based on one official language differs in meaning from a rate permitting any written language.

The comparison should state whether it concerns literacy in a specified language or literacy demonstrated in any assessed language.

67

Respondent equivalence

A self-response series cannot be assumed comparable with a later household-proxy series. The proxy share may also change with interview timing, migration or employment patterns.

Respondent status should enter the method metadata and, where possible, published quality tables.

68

Educational assumption equivalence

equity and different, learning, opportunities in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 68. Educational assumption equivalence to connect evidence with delivery, unequal effect and correction. Even a common grade label may represent different learning opportunities.

Attainment-based estimates should be separated from reported or assessed literacy and should not be used to fill missing observations without disclosure.

69

Threshold equivalence

timing and thresholds, thresholds, equivalent in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 69. Threshold equivalence to connect evidence with delivery, unequal effect and correction. Similar percentages below two thresholds do not show that the thresholds are equivalent.

Linking requires common items, common persons or another defensible design with uncertainty.

70

Sampling equivalence

Different frames may omit remote areas, collective households, migrants or persons without stable addresses. Sampling stages and clustering affect precision.

Comparable point estimates remain misleading when one design excludes populations central to the policy question.

71

Weighting equivalence

Weights reflect selection probability, non-response adjustment and calibration. Differences can improve representativeness but also signal different assumptions.

Analysts should apply design weights and appropriate variance methods. Unweighted comparisons of complex samples are generally insufficient.

72

Coverage error

authority and examined, demographic, controls in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 72. Coverage error to connect evidence with delivery, unequal effect and correction. Coverage ratios should be examined with demographic controls.

Adjustment cannot fully correct a group that has no frame representation and no reliable external total.

73

Non-response equivalence

Overall response rates can be similar while group patterns differ. Contact failure, refusal and assessment non-completion have different implications and should be separated.

Adjustment models reduce known imbalance but cannot guarantee removal of bias related to unobserved literacy.

74

Standard errors

remedy and presented, definitive, ranking in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 74. Standard errors to connect evidence with delivery, unequal effect and correction. A difference smaller than its uncertainty should not be presented as a definitive ranking.

Non-sampling error remains outside the interval and requires separate discussion.

75

Rounding and rank

continuity and substantively, meaningful, difference in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 75. Rounding and rank to connect evidence with delivery, unequal effect and correction. Published tables should discourage ordinal claims unsupported by statistically and substantively meaningful difference.

Country order is a display choice, not a finding.

76

Trend breaks

distribution and during, bridge, period in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 76. Trend breaks to connect evidence with delivery, unequal effect and correction. The old and new series may overlap during a bridge period.

Silently joining the values can produce a false improvement or decline.

77

Comparability classes

feasibility and unlikely, reverse, conclusion in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 77. Comparability classes to connect evidence with delivery, unequal effect and correction. Qualified comparability permits bounded differences that are documented and unlikely to reverse the broad conclusion.

Descriptive-only comparison shows context without supporting magnitude or rank. Not supportable applies where differences are fundamental or unknown.

78

Table 4: comparability decision matrix

Table 4. Comparability decision matrix
DimensionStrongQualifiedDescriptive onlyNot supportable
constructsame domain and interpretationbounded domain differencerelated literacy conceptunrelated or unspecified
populationharmonised age and coveragesmall documented differencematerial difference retainedpopulation unknown
methodsame or linked designevaluated method variationdifferent known methodmethod unknown
question or taskequivalent and stableminor tested adaptationdifferent known demandwording or task unavailable
languageequivalent access and adaptationbounded language differencedifferent language rule disclosedlanguage condition unknown
periodsame or policy-relevant intervalmodest documented intervaldistant observation used as contextobservation date unknown
participationadequate and similarly distributedresidual bounded differencematerial difference disclosedselective participation unexamined
uncertaintydesign-based and sufficientapproximate but decision-stablepoint estimates onlyprecision cannot be judged
permitted usemagnitude, distribution and trendbroad magnitude with qualificationcontextual juxtapositionno comparative conclusion

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

79

Comparison record

equity and should, policy, relevant in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 79. Comparison record to connect evidence with delivery, unequal effect and correction. The reason for inclusion should be policy-relevant.

Users should be able to identify why two values appear together and which inference is authorised.

80

Comparability conclusion

timing and assembled, unlike, measures in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 80. Comparability conclusion to connect evidence with delivery, unequal effect and correction. A transparent descriptive comparison can be more authoritative than an exact rank assembled from unlike measures.

The absence of full comparability should direct investment in better evidence, not justify treating current values as equivalent.

Part V

Coverage, Sampling and Participation

81

Participation as a statistical condition

remedy and population, represented, result in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 81. Participation as a statistical condition to connect evidence with delivery, unequal effect and correction. Failure at any stage can alter the population represented by the result.

Participation should therefore be analysed as part of measurement, not reported only as an operational total.

82

Target population

continuity and another, defined, population in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 82. Target population to connect evidence with delivery, unequal effect and correction. It should state age, residence, territory and any exclusions. A study may target usual residents, de facto residents or another defined population.

The choice affects migrants, temporary workers, displaced persons and adults who divide time between households.

83

Frame population

distribution and institutional, geographic, omissions in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 83. Frame population to connect evidence with delivery, unequal effect and correction. It may differ from the target population because of age, address, institutional or geographic omissions.

A coverage assessment should quantify known differences and identify groups for which no reliable estimate is possible.

84

Household-based coverage

feasibility and histories, literacy, access in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 84. Household-based coverage to connect evidence with delivery, unequal effect and correction. These groups may have distinctive educational histories and literacy access.

The published population should name exclusions. An estimate should not be called national whole-population evidence where important residents are outside the design.

85

Remote and sparsely populated areas

capacity and exclusion, substantively, important in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 85. Remote and sparsely populated areas to connect evidence with delivery, unequal effect and correction. Language diversity and educational access can make this exclusion substantively important.

Cost decisions should be explicit, and national summaries should not imply equal territorial coverage.

86

Adults in institutions

equity and affect, inclusion, privacy in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 86. Adults in institutions to connect evidence with delivery, unequal effect and correction. Institutional gatekeeping can affect both inclusion and privacy.

Where these adults are excluded, the result applies to the household population. Separate studies may be needed for policy concerning institutional education.

87

Homeless and mobile populations

timing and solely, because, absent in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 87. Homeless and mobile populations to connect evidence with delivery, unequal effect and correction. Mobility also affects repeated contact and weighting. Their omission should not be assumed negligible solely because they are absent from the frame.

Supplementary location-based or service-based approaches may inform planning, but estimates derived from them require their own population definition.

88

Sample size and precision

uncertainty and insufficient, linguistic, regional in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 88. Sample size and precision to connect evidence with delivery, unequal effect and correction. A large national sample may remain insufficient for a small linguistic or regional group.

Precision requirements should follow policy use. Publication of numerous unstable subgroup rates is not improved transparency.

89

Stratification

comparability and weights, reflect, design in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 89. Stratification to connect evidence with delivery, unequal effect and correction. Selection probabilities and analysis weights must reflect the design.

Post hoc division of a sample into many groups does not provide the same assurance as planned representation.

90

Clustering

Area-based surveys often select clusters of households. Adults within clusters may be more similar than adults selected independently, increasing variance for a given sample size.

Variance estimation should use the actual design. Treating clustered observations as a simple random sample produces unjustified precision.

91

Selection within households

coverage and available, confident, person in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 91. Selection within households to connect evidence with delivery, unequal effect and correction. The rule should prevent interviewers or households from choosing the most available or confident person.

Substitution of another adult changes selection probability and can bias the result. It should not be permitted without a controlled design.

92

Contact procedures

remedy and refusal, language, barrier in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 92. Contact procedures to connect evidence with delivery, unequal effect and correction. Repeated visits and varied timing improve inclusion. The contact record should distinguish unavailable address, no eligible adult, temporary absence, refusal and language barrier.

These categories support targeted field correction and later bias analysis.

93

Informed participation

continuity and require, literacy, assessed in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 93. Informed participation to connect evidence with delivery, unequal effect and correction. Information should not require the literacy level being assessed.

Oral explanation and accessible formats may be necessary. Consent obtained through unreadable material is not an adequate participation safeguard.

94

Refusal

Refusal may arise from time, distrust, fear, stigma or assessment burden. Conversion efforts should remain respectful and should not become coercive.

Refusal rates should be examined by area and observable characteristics. A low overall rate can conceal concentration in a group central to literacy policy.

95

Break-off and partial completion

feasibility and language, concern, performance in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 95. Break-off and partial completion to connect evidence with delivery, unequal effect and correction. Reasons include fatigue, difficulty, interruption, disability, language and concern about performance.

Partial completion should be categorised under predetermined rules. Treating every non-attempt as the lowest score may confound proficiency with access and participation.

99

Weighting for unequal selection

Base weights reflect inverse selection probabilities. Further adjustments may address non-response and align the sample with reliable population totals.[REF-12]

Weights can correct known imbalances under assumptions; they cannot recreate information for a wholly absent population or guarantee removal of literacy-related bias.

100

Non-response adjustment classes

comparability and create, unstable, weights in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 100. Non-response adjustment classes to connect evidence with delivery, unequal effect and correction. Excessively broad classes leave bias; excessively narrow classes can create unstable weights.

The procedure and weight distribution should be documented, including trimming and its effect.

101

Calibration

authority and consistently, survey, population in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 101. Calibration to connect evidence with delivery, unequal effect and correction. Controls should be reliable, temporally appropriate and defined consistently with the survey population.

Agreement on calibration variables does not ensure agreement on unobserved literacy. Residual bias remains an interpretation issue.

102

Imputation

coverage and because, principal, outcome in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 102. Imputation to connect evidence with delivery, unequal effect and correction. Imputing literacy proficiency or declared status requires stronger justification because the value is the principal outcome.

Observed and imputed shares should be reported. Imputation should not convert non-participation into apparently direct evidence.

103

Response-rate components

remedy and stages, conceals, mechanism in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 103. Response-rate components to connect evidence with delivery, unequal effect and correction. One percentage combining these stages conceals the mechanism of loss.

Comparable surveys should use consistent disposition rules and denominators.

104

Table 5: participation flow

Table 5. Participation and coverage flow for an adult literacy survey
StageRequired countQuality questionInterpretation if lost
target populationestimated eligible adultswho is intended to be representeddefines inference
frame coverageadults represented by framewhich groups are absent or duplicatedcoverage error
selected sampleselected eligible unitswas probability selection preservedselection integrity
contacted samplehouseholds or adults reachedare contact failures patternedpotential availability bias
cooperating adultseligible adults agreeingare refusals related to trust or burdencooperation bias
background completersusable contextual interviewcan non-assessment be analysedauxiliary evidence available
assessment startersadults beginning tasksdid access or anxiety prevent startparticipation barrier
valid completersscoreable assessment evidenceare break-offs selectiveachieved proficiency sample
weighted populationrepresented adults after adjustmentdo weights rely on defensible controlsbounded population inference

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

105

Participation profile

distribution and assessment, completion, separately in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 105. Participation profile to connect evidence with delivery, unequal effect and correction. It should also examine contact, refusal and assessment completion separately.

Differences guide field improvement and qualify substantive estimates. They should not be used to blame groups for non-participation.

106

Fieldwork correction

feasibility and authorised, language, support in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 106. Fieldwork correction to connect evidence with delivery, unequal effect and correction. Corrective action may include additional visits, reassignment, extended hours or authorised language support.

Correction should preserve probability selection and respondent rights. Replacing difficult cases with convenient ones is not correction.

107

Participation conclusion

capacity and reasoned, account, material in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 107. Participation conclusion to connect evidence with delivery, unequal effect and correction. Reliable adult literacy measurement requires an explicit chain from target population to valid completion and a reasoned account of every material loss.

Where participation is plausibly related to literacy, uncertainty extends beyond the sampling error and must constrain the public claim.

Part VI

Language, Culture and Accessible Measurement

108

Language as part of the construct

Literacy is exercised through particular languages and scripts. Assessment language is therefore not merely a delivery choice; it helps define what capability is observed.

Reports should specify whether the policy question concerns literacy in any language, an official language, a language of schooling or another defined set.

109

Language mapping

Design should begin with evidence on languages read, written and used by the target population. Spoken-language prevalence alone is insufficient because written use and script may differ.

The map should inform instrument languages, sample needs, recruitment and public interpretation.

110

Language choice

Respondent choice can respect capability and improve participation. It may also select among versions with different task characteristics. A controlled choice procedure should record preferred language, administered language and reason where they differ.

Forced use of one language answers a narrower question and should be labelled accordingly.

111

Translation process

Translation requires construct analysis, forward drafting, independent review, reconciliation and field testing. Back translation can identify some differences but cannot alone establish functional equivalence.

Reviewers should examine vocabulary, syntax, text genre, numerical convention and expected familiarity.

112

Adaptation register

Every departure from the source task should be recorded with reason, affected demand, reviewers and empirical evidence. Adaptation may be necessary when an institution, object or layout has no equivalent.

The objective is comparable demand, not identical surface form.

113

Item functioning across languages

Items that show unexpected differences between language groups after relevant proficiency is considered require investigation. Statistical evidence can signal a problem but cannot identify whether the cause is translation, culture, curriculum or genuine variation.

Content review and field evidence should accompany the analysis.

114

Multilingual adults

Adults may distribute reading and writing practices across languages. One language may be used at home, another at work and another for official documents.

A single-language assessment captures part of this repertoire. Background questions should record relevant practices without converting every language into a full assessment requirement.

115

Minority and indigenous languages

Excluding a written minority language can understate capability and obscure demand for literacy materials and education. Inclusion may require script, terminology and sampling expertise not available in a central instrument.

The decision and its limitations should be public, and development should involve competent language communities.

116

Oral languages and emerging written conventions

Some languages may have limited standardised written use or several orthographies. A conventional reading assessment may then measure exposure to one standard as much as general literacy.

Policy should distinguish oral capability, written-language development and literacy in other languages rather than assign a global deficit.

117

Sign languages and communication access

Instructions and consent may require sign-language access, while the assessed construct concerns engagement with written text. Interpretation should not supply answers or alter text demand.

The protocol should separate communication accommodation from substantive task modification.

118

Visual accessibility

Print size, contrast, lighting and layout can create irrelevant difficulty. Large print or other presentation may be provided where it preserves text and response demand.

Braille assessment, oral presentation or assistive technology may represent a related but different mode requiring separate validity analysis.

119

Physical access

Writing or page-handling tasks may disadvantage adults with motor impairments. A response accommodation can preserve the intended literacy construct if motor production is not part of that construct.

The use and effect of accommodation should be recorded without publicly identifying the adult.

120

Cognitive and learning differences

Instructions, pace and task load may affect adults with learning or cognitive impairments. Simplifying assessed text would change difficulty; providing accessible instructions and appropriate time may remove a procedural barrier.

No one adjustment is universally valid. The decision follows the construct and intended population claim.

121

Cultural familiarity

Documents, transactions and topics differ across settings. An unfamiliar form can create difficulty unrelated to the intended information-processing demand, while over-familiar content can advantage a subgroup.

Balanced task pools and expert review should reduce systematic irrelevant differences.

122

Gendered practices

Access to schooling, paid work, official documents and leisure reading may differ by sex and social setting. Tasks drawn mainly from one domain can reflect opportunity as well as capability.

Assessment should not erase these conditions. Background evidence helps interpret them and supports policy on access and use.

123

Rural and urban contexts

Text exposure and common documents may vary between rural and urban settings. Adaptation should preserve comparable cognitive demand without assuming that urban administrative materials define universal functionality.

Sampling and reporting should allow location differences to be examined where precision permits.

124

Respect and non-stigmatisation

Field language should not describe adults as deficient persons. It should describe observed performance under defined conditions and recognise the possibility of learning and varied practice.

Public categories should be tested for unintended stigma, particularly where results are published for small linguistic communities.

125

Table 6: language and accessibility assurance

Table 6. Language, culture and accessibility assurance record
Assurance areaEvidence before fieldworkField recordReporting boundary
language populationspoken and written language mappreferred and administered languagescope limited to offered languages
translationconstruct-based review and testingversion identifierno presumed equivalence without evidence
adaptationreason and demand analysisitem version usedmaterial change disclosed
multilingual practicebackground questionslanguages of useone-language score not whole repertoire
cultural contextexpert and participant reviewdifficulty observationsirrelevant familiarity considered
communication accessaccessible information and instructionssupport providedsupport not treated as item performance
visual accesspresentation and accommodation rulesformat used and exclusionswhole-population claim qualified
physical accessresponse-method protocolauthorised adjustmentmotor barrier not confused with literacy
non-stigmatisationterminology reviewcomplaint or distress recordcategories describe performance, not worth

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

126

Accessible reporting

Literacy findings should themselves be communicated in forms accessible to adults with varied literacy. Oral briefings, clear graphics and community-language explanations can accompany the technical report.

Simplification should preserve uncertainty and method. Accessibility is not permission to turn a qualified result into an absolute statement.

127

Language and accessibility conclusion

A literacy measure cannot be separated from the language, script and mode through which performance is elicited. Comparability requires equivalent construct access, not uniform administration that excludes part of the population.

The report should state who could demonstrate capability, under which conditions, and whose capability remains unobserved.

Part VII

Linking Self-Report and Direct Assessment

128

Purpose of linking

Collecting self-report and direct assessment from the same adults can show disagreement, examine reporting thresholds and support adjustment of broad survey indicators. It can also connect perceived difficulty and practice to observed performance.

The exercise should not assume that one measure contains every truth and the other only error.

129

Validation subsample

A probability subsample of respondents to a census-linked or household survey module may complete direct assessment. Selection probability, non-response and language access require separate weights.

Volunteers are unlikely to represent adults who decline or fear assessment and should not be used to calibrate a national rate without qualification.

130

Cross-classification

Self-reported categories can be cross-tabulated against proficiency levels or task outcomes. The table should show counts, weighted proportions and uncertainty.

Disagreement is expected because constructs and thresholds differ. It should be analysed, not labelled automatically as false reporting.

131

Sensitivity and specificity

If a direct threshold is adopted for a defined purpose, self-report sensitivity describes the share above that threshold reporting literacy; specificity describes the share below reporting non-literacy. Both depend on the chosen threshold and population.

They should not turn a contested literacy boundary into a natural fact.

132

Predictive models

Background variables and self-reported practices may predict assessed proficiency. Models should be developed and validated on separate data where feasible and should report error across groups.

Predicted proficiency is not directly assessed proficiency. Publication should not obscure that distinction.

133

Differential reporting

The relationship between self-report and assessment may differ by age, sex, language, education, location and social expectations. A single national correction factor can therefore misclassify subgroups.

Analysis should test interaction and retain adequate sample size and protection.

134

Proxy validation

Where both adult self-response and household proxy response are available, agreement can be analysed before direct assessment is considered. Three-way comparison can distinguish proxy disagreement from self-assessment disagreement.

Re-interview conditions should minimise learning or disclosure effects.

135

Method transition

A country moving from self-report to direct assessment should conduct an overlap or bridge study. Both methods are applied under controlled conditions during at least one period.

The published series should mark the transition. A difference between methods should not be presented as population change.

136

Composite systems

A layered system can use a short household module for broad coverage, a direct-assessment subsample for proficiency and qualitative enquiry for literacy practices and barriers. Each component retains its own inference.

Integration occurs through common identifiers, definitions and analysis, not by collapsing all evidence into one rate.

137

Table 7: method-linking record

Table 7. Self-report and direct-assessment linking record
Linking elementRequired designValid inferenceProhibited inference
overlap sampleprobability selection and response weightspopulation relationship between methodsvolunteer agreement as national validity
common populationaligned age, residence and language rulesmethod difference within defined groupdifference attributed to time trend
cross-classificationweighted cells and uncertaintypattern of agreement and disagreementself-report labelled dishonest
threshold analysisexplicit policy thresholdconditional sensitivity and specificitythreshold treated as universal literacy divide
subgroup analysisadequate protected samplesdifferential reporting patternsunstable cells ranked
predictionmodel validation and errorbounded estimate for stated usepredicted value called observed skill
bridge periodparallel administrationestimated series discontinuityspliced series without marker
publicationseparate results and interpretationcomplementary evidenceone method silently replaces another

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

138

Ethical linking

Linking requires clear authority, confidentiality and purpose. Adults should not experience a declaration as a commitment that triggers an undisclosed test or personal consequence.

Identifiers should be protected and removed from analytical files where no longer necessary.

139

Linking conclusion

Linked evidence can improve understanding of method effects and support a responsible transition to stronger measurement. It cannot make unlike constructs identical.

The value lies in explaining difference and uncertainty, not manufacturing a single definitive number.

Part VIII

Reporting Literacy Indicators Responsibly

140

Publication purpose

Literacy statistics should inform education, language, labour, social and community policy while respecting the people represented. Publication should permit independent understanding of the measure and its limits.[REF-09]

A headline rate without method and observation year is incomplete public evidence.

141

Core metadata beside the value

Every principal table should state source, observation period, population age, geographic coverage, method class, respondent rule, language condition and whether schooling assumptions were used.

Detailed technical notes may follow, but essential meaning should not be separated from the number.

142

Counts and denominators

Rates should be accompanied by weighted population counts where appropriate and by unweighted sample counts for survey transparency. These serve different purposes and should be labelled.

Small unweighted cells require suppression or caution even where weighted populations appear large.

143

Precision

Sample estimates should show standard errors, confidence intervals or another accepted measure of sampling uncertainty. Tables should indicate where design or sample size makes an estimate unstable.

Decimal places should reflect precision. Additional digits do not create information.

144

Non-sampling limitations

Coverage, non-response, language, question, mode, scoring and model limitations should be stated near the affected result. Sampling intervals do not capture these errors.

A concise direction-of-bias statement is useful where evidence supports it; otherwise uncertainty should not be assigned a convenient sign.

145

Distribution before averages

Direct assessments should report proficiency distributions and relevant percentiles or level shares, not only a mean. A common average can conceal different lower and upper tails.

Policy concerning minimum access requires attention to adults facing the greatest task difficulty without reducing them to a permanent label.

146

Disaggregation

Sex, age, location, language, education, work and socio-economic measures may reveal unequal opportunity and participation. Selection should follow policy relevance and sample capacity.

Disaggregation is not a licence for an unlimited table of unstable comparisons.

147

Time series

Charts should mark observation years and method breaks. Lines should not imply annual measurement where only intermittent observations exist.

Modelled interpolation and projection should be visually and textually distinct from observed values.

148

Cross-national tables

Country values should be grouped or annotated by method comparability. Alphabetical presentation is often preferable to an unsupported rank.

Where method differences are fundamental, separate panels are more honest than a single ordered column.

149

Threshold language

“Below the selected proficiency level” states a measurement result. “Illiterate population” may imply an absolute condition not established by the assessment, especially where one language or domain was tested.

Labels should follow construct and avoid stigma.

150

Associations and causes

Relationships between literacy and income, employment, health or civic participation are policy-relevant. Cross-sectional association does not establish that measured literacy alone caused the difference.[REF-06] [REF-07]

Education, background, opportunity and selection may contribute. Claims should match design.

151

Change and programme effect

Population change between surveys can reflect cohorts, migration, education, practice, economic conditions and measurement. It should not be attributed to one literacy programme without an appropriate evaluation.

Programme participants are also not a random population sample; their change answers a narrower question.

152

Public communication

Briefings should explain what adults were asked or required to do, who was included and what a level means in practical task terms. Visual material should preserve denominators and uncertainty.

Communication should identify actionable barriers and provision needs rather than sensationalise a population count.

153

Confidentiality

Small cells, rare languages and local areas can expose identities or stigmatise communities. Disclosure control should consider direct and inferential risk.

Suppression and aggregation should be applied consistently and should not conceal material inequality from authorised policy review.

154

Corrections

Errors in weights, coding, labels or tables should be corrected publicly with date, affected results and interpretive consequence. The original release remains documented.

Quiet replacement weakens official-statistics accountability.

155

Table 8: publication assurance

Table 8. Adult literacy indicator publication assurance
Publication elementMinimum contentMisleading practiceRequired correction
headline valuemethod, population and observation yearmethod-neutral current raterestore scope beside value
tablecounts, denominator, languages and sourcerank of unlike methodsseparate or classify comparison
uncertaintysampling and material non-sampling limitspoint estimate treated as exactadd precision and qualification
distributionlevels or score rangemean as full population accountpublish distribution
trendobserved years and breakscontinuous line through missing yearsmark observations and discontinuities
subgrouprelevance, sample and protectionunstable small-cell rankingsuppress, combine or qualify
narrativeassociation and bounded interpretationcausal claim from cross-sectionrevise to supported relationship
terminologyconstruct-respecting languageadult worth reduced to labeldescribe performance and context
correctionversion, date and effectsilent replacementissue correction notice

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

156

Reporting conclusion

Responsible reporting makes the method visible and the policy meaning bounded. It does not weaken literacy advocacy; it ensures that public action responds to evidence rather than to artefacts of measurement.

The strongest statistic is not the most definite-sounding value, but the value whose population, method and uncertainty can withstand scrutiny.

Part IX

Applied Comparative Cases

157

Status of the cases

The following cases are constructed to examine measurement decisions. They do not describe identified countries and do not establish empirical conversion factors. Each case retains the population, method and participation conditions needed to understand the calculation.

158

Case A: a census self-declaration rate

A census asks one household respondent whether each member aged 10 and over can read and write a simple message in any language. The published adult rate uses persons aged 15 and over. Of 2,480,000 adults enumerated, 2,021,200 are recorded yes, 421,600 no and 37,200 unknown or not stated.

Dividing yes responses by all enumerated adults gives 81.5 per cent. Dividing yes by known responses gives 82.8 per cent. Neither denominator is inherently correct for every use. The first implicitly treats unknown as not literate; the second assumes missing status can be excluded without bias.

Metadata show that one respondent answered for 68 per cent of other adult members. The proxy share is higher among employed men absent at interview and younger adults temporarily away. The census provides valuable small-area coverage but does not establish observed proficiency.

The publication reports 2,021,200 declared literate adults, 421,600 declared not literate and 37,200 unknown. It uses 81.5 per cent only with an explicit denominator and shows 82.8 per cent as a known-response rate. Policy maps include unknown status rather than merging it with either category.

159

Case A record

Table 9. Census self-declaration calculation
ComponentCountShare of all adultsInterpretation
enumerated adults aged 15+2,480,000100.0%population denominator
declared can read and write2,021,20081.5%self or proxy classification under question
declared cannot421,60017.0%declared category, not direct assessment
unknown or not stated37,2001.5%retained separately
known responses2,442,80098.5%alternative denominator
yes among known responses2,021,20082.8%conditional rate requiring missingness caution
records supplied by proxy for another adult1,686,40068.0%method-quality characteristic

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

160

Case A decision

The census result can support local identification of areas with high declared need, subject to proxy and missingness patterns. It cannot be compared as a proficiency rate with a direct assessment or used to infer writing and document skills beyond the question.

A validation survey should oversample areas with high unknown and proxy reporting and record language, self-response and performance separately.

161

Case B: direct sentence verification

A household survey asks adults with less than completed secondary education to read a sentence card. Adults with secondary completion are assumed able and are not tested. Among 8,000 sampled adults, 4,600 are assumed literate, 2,700 are asked to read and 700 have missing education information or cannot be routed.

Of those asked, 1,620 read the whole sentence, 540 read part, 405 cannot read it and 135 do not complete for language, vision, refusal or interruption. Reporting 4,600 assumed plus 1,620 whole-sentence readers as “literate” gives 6,220 of 8,000, or 77.8 per cent. The value combines indirect and direct methods and treats 1,780 adults outside the numerator for different reasons.

The report therefore provides a component table. It also produces a lower-bound observed whole-sentence count of 1,620 and does not present it as a population literacy rate because most adults were not assessed.

162

Case B record

Table 10. Mixed assumption and sentence-task classification
Routing resultCountMeasurement statusPublication treatment
assumed from secondary completion4,600indirect classificationpublish separately
read whole sentence1,620direct narrow taskdescribe exact performance
read part540direct partial performanceretain ordered category
unable to read sentence405direct narrow taskno claim beyond task and language
task non-completion135unobserved direct performanceanalyse reason
routing unresolved700education or eligibility missingdo not presume status
combined headline numerator6,220unlike evidence combineduse only with full method label
combined headline rate6,220 / 8,00077.8%not directly comparable with single-method rate

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

163

Case B decision

The mixed indicator may meet an immediate reporting convention but is unsuitable for evaluating proficiency or comparing with a self-declaration series. The next survey should assess a probability subsample across all education levels to test the assumption and quantify task non-completion.

Adults who read part of the sentence should not be combined arbitrarily with either extreme. Their category contains policy information about emerging or limited reading capability.

164

Case C: specialised direct assessment

A specialised survey selects 5,400 adults aged 16–65. The frame excludes collective institutions and three remote districts containing an estimated four per cent of the otherwise eligible population. Of 5,400 selected adults, 4,590 complete the background interview and 4,050 provide valid assessment evidence.

Weighted results place 18 per cent below Level 1, 31 per cent at Level 1, 34 per cent at Level 2 and 17 per cent at Level 3 or above. Standard errors range from 0.8 to 1.2 percentage points. The figures describe the covered household population aged 16–65, not all adults aged 15 and over.

Non-completion is higher among adults whose background interview reports limited use of the assessment language. Weighting adjusts for age, sex, region and education but not directly for unobserved proficiency. The report presents sampling intervals and a residual non-response limitation.

165

Case C record

Table 11. Specialised assessment participation and proficiency
MeasureResultUncertainty or coveragePermitted statement
selected adults5,400probability samplefield sample
background completers4,59085.0% of selectedcontextual evidence available
valid assessment evidence4,05075.0% of selectedachieved proficiency sample
below Level 118%SE 0.8 pointscovered population estimate
Level 131%SE 1.1 pointscovered population estimate
Level 234%SE 1.2 pointscovered population estimate
Level 3 or above17%SE 0.9 pointscovered population estimate
frame exclusionabout 4% of otherwise eligible populationinstitutions and remote districtsno whole-population claim

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

166

Case C decision

The direct assessment supports distributional analysis and relationships with background variables within the covered population. It does not provide a directly comparable replacement for the national census rate because age, coverage, method and construct differ.

The next cycle should test language-access improvements and develop evidence for excluded districts. A more precise point estimate within the current frame would not resolve those coverage limitations.

167

Case D: apparent trend after a method change

A country reports self-declared adult literacy of 76.4 per cent in an earlier census. Five years later, a household survey using a sentence-reading task reports 69.8 per cent reading the whole sentence, 8.6 per cent reading part and 21.6 per cent unable or unobserved under its published rule.

The seven-point difference is presented initially as decline. Review finds different age ceilings, proxy use in the census, testing only in two languages and exclusion of remote areas in the survey. There is no overlap sample.

The values cannot establish decline. The census describes declared reading and writing in any language; the survey describes sentence performance under a restricted language and coverage design. They remain useful as separate observations.

168

Case D comparison record

Table 12. Method-change trend assessment
DimensionEarlier censusLater household surveyComparability judgement
populationage 15+, national census scopeages 15–64, remote areas excludedmaterial difference
constructcan read and write simple messagereads whole or part of sentencerelated, not equivalent
respondent72% proxy for other adultsadult performs taskmethod break
languageany language declaredtwo assessment languagesmaterial access difference
result76.4% declared yes69.8% whole sentenceno direct subtraction
partial categoryabsent8.6%distribution not binary-equivalent
bridge studynonenonedifference cannot be allocated to method or time
trend classificationnot supportable

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

169

Case D decision

The time series should show a discontinuity and explain both methods. A future bridge study can administer the census question and the direct task to the same probability sample and examine results by language and respondent status.

Until then, programme claims should use other evidence and avoid attributing the difference to national change.

170

Case E: preferred-language assessment

A multilingual assessment offers three language versions. Of 3,200 participating adults, 1,920 select Language A, 880 Language B and 400 Language C. Translation review is complete, but item analysis finds four document tasks easier in Version B after overall proficiency is considered.

The national distribution combines versions under the original scale. The language-specific comparison is withheld pending review of those items. Removing them changes the mean for Version B by 6 scale points and the national mean by 1.2 points.

The national result may remain sufficiently stable for broad reporting, while language-group rank is not supportable. The publication documents the item decision and retains preferred-language participation as a strength of coverage.

171

Case E decision record

Table 13. Multilingual version-equivalence decision
EvidenceNational estimate effectLanguage comparison effectDecision
four flagged document tasksunder reviewVersion B unexpectedly easierinvestigate content and translation
exclusion sensitivitymean changes 1.2 pointsVersion B mean changes 6 pointsnational broad conclusion stable; subgroup rank unstable
preferred-language availabilitywider participationdifferent version exposureretain and report language choice
sample by versionA 1,920; B 880; C 400unequal precisionpublish standard errors
final releasecombined distribution with noteno ordered language comparisonissue technical qualification

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

172

Case E boundary

Different average scores by assessment language cannot be interpreted as language effects because populations choosing each language differ in education, age, location and opportunity. Version analysis addresses measurement equivalence, not the social cause of group differences.

Policy use should focus on access to adult learning and written materials in relevant languages, supported by broader evidence.

173

Case F: non-response adjustment

A direct assessment selects 10,000 adults. Valid assessment evidence is obtained from 7,200. Response is 82 per cent among adults with upper-secondary education and 61 per cent among adults with primary education or less. Base-weighted proficiency above a selected level is 58.0 per cent.

Non-response weights formed by age, sex, region and education reduce the estimate to 54.6 per cent. A follow-up of 300 initial non-respondents obtains a short task from 174 and indicates lower performance than respondents within the same education groups.

The adjusted 54.6 per cent remains potentially high because the weighting variables do not capture all response-related proficiency. The report presents the estimate, sampling error and a directional residual-bias statement rather than another precise correction from the small follow-up.

174

Case F record

Table 14. Non-response adjustment and residual bias
StageEstimate above selected levelEvidenceInterpretation
base-weighted respondents58.0%selection weights onlyrespondent distribution
adjusted main estimate54.6%age, sex, region and education adjustmentimproved population estimate
valid completions7,200 of 10,00072.0%material non-response
higher-education response82%field dispositionresponse related to education
lower-education response61%field dispositionunder-representation likely
follow-up completion174 of 300 traced casesselective small follow-updirectional evidence, not full correction
residual judgementprobable upward bias remainslower follow-up performancequalify estimate

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

175

Case F decision

The adjusted estimate is preferable to the base-weighted respondent value, but neither weighting nor follow-up proves elimination of bias. The next design should strengthen initial participation and collect auxiliary information for all selected adults.

Publication should not select 58.0 per cent because it is the least adjusted or 54.6 per cent because it appears more conservative. It should use the method judged most defensible and disclose the residual limitation.

176

Case G: programme evaluation and population indicators

An adult learning programme assesses 640 entrants and 472 completers. Mean scores rise by 18 points among completers. A national household indicator improves by two percentage points during the same period.

The programme result is affected by attrition, practice, instruction and assessment conditions. The population indicator includes non-participants and different cohorts. Neither result alone establishes the programme’s population effect.

Baseline scores are available for 143 non-completers, whose mean is lower than that of completers. Complete-case change is therefore likely to overstate the result for all entrants.

177

Case G record

Table 15. Programme and population evidence boundary
EvidencePopulation representedFinding supportedFinding not supported
640 entrant baselinesenrolled entrantsstarting distributiongeneral adult population
472 matched completersretained participantschange among completerschange among all entrants
168 without follow-upattriting entrantsattrition count and baseline differencezero change without evidence
18-point mean gainmatched completersobserved within-person differencecausal effect without comparison
national two-point changecovered household populationpopulation indicator movementprogramme contribution alone
participation recordsprogramme exposurereach and completion patternproficiency of non-participants

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

178

Case G decision

Evaluation should report entry, retention and matched change, examine attrition and use an appropriate comparison where feasible. The national indicator provides context but cannot serve as the programme counterfactual.

Programme success should include participation and equitable reach as well as assessment change.

179

Case H: local-area estimates from a national survey

A national survey is asked to publish literacy estimates for 96 districts. In 41 districts the unweighted adult sample is below 40, and several estimates have relative standard errors above 25 per cent. Direct publication would invite unstable ranking.

The survey can provide reliable estimates for six regions. Model-assisted district estimates may be developed using census covariates, but they are partly predicted and depend on model assumptions.

The report publishes regional direct estimates and classifies district results as experimental model-based planning evidence with uncertainty intervals and validation diagnostics.

180

Case H record

Table 16. Geographic-estimate release decision
Geographic productDirect sample conditionEstimation statusRelease decision
nationaladequate probability sampledirect design-basedpublish
six regionsplanned strata and acceptable precisiondirect design-basedpublish with standard errors
55 districtsat least 40 observations but variable precisiondirect estimate often unstablepublish only where quality rule passes
41 districtsfewer than 40 observationsinadequate direct precisiondo not rank or release as direct rate
model-assisted districtssurvey plus census covariatespartly predictedlabel separately with intervals
district changeone survey roundno stable time seriesdo not infer trend

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

181

Case H decision

Small-area need remains important, but unstable direct rates do not meet it. Census self-report, administrative information, qualitative evidence and model-assisted estimates can be considered together with distinct labels.

Resource allocation should not depend on an unqualified single district rank where uncertainty can reverse the order.

182

Cross-case finding

The cases show that most serious errors occur when unlike states are collapsed: unknown with not literate, assumed with assessed, non-completion with lowest proficiency, method change with trend, language version with group effect, weighted adjustment with removal of all bias, programme change with national impact, or predicted local values with direct observations.

Reliable comparison preserves these distinctions and makes the resulting policy choice explicit.

183

Table 17: cross-case error and remedy

Table 17. Cross-case measurement error and remedy
CaseCollapsed distinctionMisleading conclusionRequired remedy
census declarationunknown versus declared nomissing adults classified without evidencepublish components and denominators
sentence taskassumed versus directly observedmixed rate called assessed literacyseparate method classes
specialised assessmentcovered versus target populationpartial frame called all adultsrestrict population claim
method-change trendmeasurement versus timedecline inferred from unlike valuesmark break and bridge methods
multilingual assessmentversion effect versus population differencelanguage groups rankedtest items and withhold unsupported rank
non-responseadjusted versus unbiasedweight treated as complete correctiondisclose residual bias
programme evaluationcompleters versus entrants and populationprogramme credited with national changeanalyse attrition and comparison
district estimatedirect versus predictedunstable local ranksquality rules and separate model status

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

184

Applied-case conclusion

Adult literacy indicators gain authority when the analytical record permits a reader to reconstruct who was represented, what was observed and why a comparison was allowed. Precision without these distinctions can conceal rather than reduce uncertainty.

Part X

Policy Uses and Institutional Responsibilities

185

Measurement in service of literacy policy

The Dakar Framework identifies adult literacy as a central education commitment and calls for substantial improvement, particularly for women, together with equitable access to continuing education. Measurement should make unmet need and unequal opportunity visible and support decisions on provision.[REF-01]

The indicator is a means of public action. It should not become a substitute for investment in learning opportunities, language materials and environments in which literacy can be used.

186

National policy diagnosis

National diagnosis should combine prevalence, distribution, language, age, education, location, work and participation evidence. One total rate cannot show whether low observed proficiency arises principally among older cohorts, isolated areas, particular language communities or adults with interrupted schooling.

Policy should respond to the pattern and to direct evidence about barriers, not to a headline rank.

187

Planning adult learning provision

Direct assessment can inform the range and level of programmes required; self-reported difficulty and practice can inform access, scheduling and relevance. Small-area census data can support geographic placement where the method boundary is retained.

Provision should not restrict admission to adults who fail a particular task. Assessment for planning differs from an eligibility test.

188

Identifying language needs

Measurement should distinguish language of literacy from general ability. Data on languages used for reading and writing can guide materials, instructors and communication.

A low result in an official-language assessment may indicate a need for official-language learning, but it should not erase capability in another language or justify exclusion from services.

189

Gender equality

Sex-disaggregated results remain necessary because access to schooling, adult learning, time, mobility and literacy use may differ substantially. The 2006 global development and literacy monitoring context gives particular attention to gender disparities.[REF-02] [REF-14]

Analysis should examine survey participation and proxy response by sex as well as the substantive result.

190

Cohort analysis

Age patterns can reflect historical differences in educational access and opportunities for use. They should not be interpreted solely as individual decline. Cross-sectional differences between age groups are not the same as longitudinal skill loss.

Cohort analysis can nevertheless help plan programmes and anticipate change as younger and older cohorts enter and leave the adult population.

191

Educational attainment

Combining literacy and attainment evidence can reveal adults whose skills exceed or fall below what formal qualifications might suggest.[REF-06] [REF-07]

Policy should avoid assuming that school completion guarantees current proficiency or that adults without schooling lack all literacy. The two indicators answer different questions.

192

Employment and workplace learning

Literacy proficiency and practices are associated with employment, training and income, but relationships are shaped by opportunity and selection.[REF-06] [REF-07]

Workplace policy should expand opportunities to learn and use skills without using a population assessment to screen individual workers beyond its design.

193

Health and public services

Adults may face written demands in health, finance, transport and public administration. Literacy evidence can guide clearer communication and assisted service channels.

The appropriate response is not only to change the adult. Institutions should reduce unnecessary text complexity and make essential information accessible.

194

Civic participation

Literacy supports access to public information and participation, while civic activity also creates purposes for literacy. Measurement may include relevant practices without treating political participation as a literacy test.

Public communication of results should enable communities to deliberate about provision and resource priorities.

195

Poverty analysis

Household surveys can relate literacy indicators to consumption, income, employment and living conditions. Association should be examined with household composition, location, education and opportunity.[REF-11] [REF-13]

A poverty gradient does not establish that literacy alone caused economic status. Policy may need coordinated educational and social action.

196

Programme targeting

Geographic or group evidence can guide outreach, but targeting should avoid stigma and ecological error. A resident of a low-rate area should not be presumed to have a particular level.

Individual entry assessment should serve placement and support, with privacy and an opportunity to demonstrate learning needs through an appropriate method.

197

Resource allocation

Allocation formulas may include population size, measured need, access barriers and service cost. Indicator uncertainty and method differences should be reflected rather than hidden.

A small difference in estimated rate should not move substantial resources where sampling or non-sampling uncertainty can reverse the order.

198

Monitoring national commitments

Global and national monitoring requires stable definitions and observation dates. The international Literacy Decade and Education for All commitments emphasise strengthened literacy action and monitoring.[REF-01] [REF-15]

Monitoring should not reward countries for adopting a method that yields a higher rate. Improvements in measurement may initially lower or redistribute estimates and should be recognised as statistical progress.

199

Evaluating policy

Repeated comparable assessment can show population change but does not by itself attribute change to a policy. Evaluation should examine exposure, implementation, cohort effects, economic conditions and other plausible explanations.

Where only one pre- and post-observation exists, causal conclusions should remain limited.

200

Early-warning use

Changes in programme participation, literacy practice or a short module may provide earlier signals than a full assessment cycle. These indicators can prompt enquiry but should not be presented as substitutes for proficiency change.

An early-warning threshold should have a defined verification and response.

201

Local planning

Local authorities need evidence at a usable geographic scale. Census breadth, administrative programme data and community enquiry may support local planning even when specialised assessment is reliable only regionally.

The sources should remain separate. A locally reported declaration rate should not inherit the measurement authority of a national direct assessment.

202

Community participation

Adults, educators and language communities should contribute to construct relevance, task review, field access and interpretation. Their participation can identify practices and barriers absent from technical review.

Participation does not allow interest groups to suppress unfavourable findings. The statistical authority retains responsibility for method and fair release.

203

Statistical authority

The national statistical authority should protect professional methods, confidentiality, impartial release and public access in accordance with the Fundamental Principles of Official Statistics.[REF-09]

It should document revisions and resist political selection of definitions or dates designed to produce a preferred result.

204

Education ministry

The education ministry should define policy questions, provide programme and education-system knowledge and use findings for provision. It should not alter statistical results to align with programme claims.

Joint governance should preserve the responsibilities and independence of each body.

205

Census office

The census office can provide broad population and local-area evidence and maintain exact question and respondent metadata. It should evaluate proxy response, assumptions and non-response.

Where a specialised assessment is planned, common background variables and a validation subsample can strengthen linkage without overburdening the census.

206

Assessment body

The body responsible for direct assessment should maintain the framework, item security where required, language versions, field standards, scoring, scaling and technical documentation.

It should state whether the design supports population estimates, individual results or both. High-stakes reuse requires separate evidence.

207

Adult education providers

Providers can contribute practical knowledge about learner goals, barriers and programme assessment. Their participant data describe service users, not the unserved population.

Providers should not be asked to produce national prevalence estimates from enrolment records.

208

Research institutions

Independent research can test validity, non-response, language effects and policy relationships. Access to protected microdata should follow clear conditions and disclosure control.

Publication freedom and replication improve confidence, while confidentiality remains binding.

209

International comparison

International organisations can support common definitions, technical capacity and comparable assessments. They should also preserve national metadata and avoid implying equivalence where methods differ.

Countries should participate in methodological decisions and retain the capacity to interpret results in their own linguistic and institutional context.

210

Procurement and technical services

Where external services support sampling, printing, data collection or analysis, contracts should specify method, quality evidence, ownership, confidentiality, transfer and correction.

Technical delegation does not transfer public accountability for the indicator.

211

Governance committee

A national literacy measurement committee may coordinate policy questions, languages, population coverage, field access and release. Membership should include statistical, education, adult learning and relevant language expertise.

The committee should record decisions and conflicts without compromising the statistical authority’s professional responsibility.

212

Confidentiality governance

Literacy data can expose personal circumstances and create stigma. Access should be limited by role, identifiers separated where feasible and outputs tested for disclosure.

Individual results should not be transferred to employers, benefit authorities or enforcement bodies unless a clear lawful purpose and appropriate design exist.

213

Release calendar

The observation period, processing schedule, preliminary status and final release date should be published in advance where practicable. Equal access to principal results supports impartiality.

Delay for unresolved quality concerns should be explained; delay to secure a favourable policy narrative is inappropriate.

214

Correction authority

The statistical body should have authority to correct errors, issue revised tables and explain consequence without awaiting agreement from every policy stakeholder.

A correction affecting a public target should prompt a separate policy response rather than alteration of the statistical record.

215

Capacity development

Sustainable measurement requires sampling, field, language, data processing, psychometric, analytical and dissemination capability. Short external missions cannot replace national institutional development.

The contemporary LAMP initiative places capacity alongside improved assessment and policy relevance.[REF-03]

216

Cost and periodicity

Censuses offer breadth at long intervals; household modules can recur more often; specialised assessments provide depth at higher cost. A national architecture should combine them according to decision cycles.

Reducing sample quality or language access to obtain annual data may produce a less useful series than sound measurement at wider intervals.

217

Burden on adults

Assessment time, travel, anxiety and repeated contact are real burdens. Instruments should retain questions and tasks that serve defined analysis and avoid duplicating information available reliably elsewhere.

Burden review should include adults who did not complete, not only those who tolerated the full design.

218

Public accountability

Institutions should report method decisions, cost, coverage, response, quality findings, release and correction. Accountability concerns the integrity and usefulness of evidence, not whether the rate meets a political expectation.

An unfavourable result can be responsibly produced; an unexplained favourable result cannot.

219

Table 18: institutional responsibility map

Table 18. Institutional responsibility for adult literacy evidence
FunctionLead responsibilityRequired controlPublic output
policy questioneducation and adult-learning authoritiesdecision use and public-interest testmeasurement brief
population framestatistical authority or census officecoverage and selection recordpopulation scope
construct and tasksassessment body with literacy expertiseframework and validity evidenceassessment description
language versionsassessment and language specialistsadaptation and equivalence reviewlanguage coverage note
fieldworkstatistical or contracted field bodytraining, contact and deviationsresponse profile
weighting and estimationstatistical authorityreproducible methods and varianceestimates with uncertainty
policy interpretationresponsible ministries and communitiesclaim-source boundarypolicy response
statistical releasestatistical authorityimpartiality, metadata and correctioncontrolled publication
confidentialitydata custodianaccess and disclosure controlsprotected public tables
evaluationindependent or functionally separate reviewquestion, comparison and limitationsfindings and management response

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

220

Institutional conclusion

No single body holds all knowledge required for adult literacy measurement. Coordination is necessary, but responsibility should remain visible. Statistical integrity, educational relevance, language competence and public participation are complementary controls.

The resulting indicator should be judged by whether it supports better and more equitable learning opportunities, not by whether it produces a convenient national figure.

Part XI

A National Adult Literacy Measurement Architecture

221

Architectural principle

A national system should combine breadth, depth, context and continuity. Attempting to obtain all four from one instrument either overloads the instrument or leaves essential questions unanswered.

Each component should have a defined role and a documented relationship to the others.

222

Component 1: census literacy module

A census module can provide broad coverage, small-area evidence and a population frame. It should use tested questions, record proxy status where feasible and avoid deriving literacy solely from schooling.[REF-08]

Its output is a declared or reported classification under the census method, not a proficiency distribution.

223

Component 2: household survey module

A recurring household module can collect graded self-assessment, literacy practices, language and access to learning alongside social and economic variables.[REF-11] [REF-13]

It can monitor perceived difficulty and context more frequently than a specialised assessment, provided wording and mode remain stable.

224

Component 3: direct assessment

A periodic probability assessment can estimate proficiency distributions across defined domains and relate them to background conditions. Its sample, languages and administration require dedicated quality control.

It should be timed to policy and evaluation needs rather than forced into an annual reporting cycle.

225

Component 4: programme evidence

Adult education programmes should maintain enrolment, attendance, completion and learning evidence suited to their objectives. Common core fields can support system analysis, but local relevance remains important.

Programme data describe participants and should not fill national prevalence gaps.

226

Component 5: qualitative enquiry

Interviews, group discussions, observation and community studies can examine literacy purposes, stigma, language, service access and reasons for non-participation. A stated sampling and analysis method is required.

Qualitative evidence explains mechanisms and meanings that a scale cannot fully represent.

227

Common population concepts

Components should align core age, sex, residence, geography, language and education definitions where feasible. Differences required by purpose should be mapped.

Common labels should not conceal different eligibility or coverage rules.

228

Common identifiers and privacy

Statistical linkage may use protected identifiers under lawful authority. Public and general analytical files should remove direct identifiers and limit detail.

Linkage value should be weighed against disclosure risk and public trust.

229

Question bank

A controlled bank can retain tested self-report, practice and background questions with wording, translations, response categories and evidence. Instruments select from it under version control.

Unrecorded local alteration prevents comparison and should be avoided.

230

Assessment framework maintenance

The framework should be reviewed for social and linguistic relevance without changing the construct silently. New task types may be introduced through field trials and linking designs.

Secure items and released examples should together support integrity and public understanding.

231

Language programme

Language inclusion requires a planned process for population mapping, prioritisation, translation, adaptation, recruitment, sampling and analysis. It should not be arranged only after fieldwork begins.

Decisions to include or exclude languages carry both validity and equity consequences.

232

Master sample or coordinated frames

Coordinated area frames can reduce cost and improve consistency across household modules and direct assessment. Sample rotation should manage respondent burden and preserve independent selection.

Frame updates are necessary where migration, growth or displacement changes coverage.

233

Bridge studies

Bridge studies should be scheduled when questions, modes, languages, age limits or assessment frameworks change. Parallel administration allows method effects to be distinguished from population change.

The bridge sample must be large and representative enough for the intended reconciliation.

234

Metadata registry

The registry should contain every indicator version, exact question or framework, population, source, period, method, languages, respondent rule, weighting and limitations. Tables and public summaries refer to the registry entry.

Historical versions remain accessible so that trend breaks can be understood.

235

Observation calendar

The calendar identifies census years, household modules, assessment cycles, programme reporting and policy reviews. It permits evidence to be combined without pretending that all sources are contemporaneous.

Major policy evaluation should be scheduled around an adequate baseline and follow-up rather than an administrative anniversary alone.

236

Analysis plan

The analysis plan should be approved before principal results are known. It defines populations, indicators, weights, variance, proficiency reporting, subgroup analysis, missing data and comparison rules.

Exploratory findings may be reported as such, with protection against selective emphasis.

237

Quality assurance plan

Quality assurance covers frame, selection, instrument, translation, fieldwork, scoring, data processing, weighting, analysis, disclosure and dissemination. Responsibilities and evidence should be assigned before collection.

Quality is not a final inspection applied after methodological decisions can no longer be corrected.

238

Pilot and field trial

A pilot tests operational flow, consent, contact, burden and logistics. A field trial tests items, language versions, scoring and measurement properties. The functions overlap but are not identical.

Both should include adults likely to encounter access and participation barriers.

239

Main collection readiness

Readiness requires an approved framework, stable instruments, trained personnel, verified systems, language materials, sample release, protection arrangements and contingency. Failure of a critical condition should delay or stage collection.

Calendar pressure should not turn an unresolved language or scoring issue into accepted error.

240

Processing and reconciliation

Data processing should preserve field dispositions, original responses, edits, scoring and weights. Changes require rules and audit trails.

Counts should reconcile from selected sample through final estimates. Unexplained loss between stages prevents release.

241

Technical review

Reviewers should examine construct, sampling, field response, language, scoring, scaling, weighting, uncertainty and disclosure. Findings and management responses should be recorded.

Independence should be described accurately; internal technical challenge is not the same as external review.

242

Preliminary results

Preliminary results may support timely policy where processing is sufficiently mature. Their status, possible revision and excluded analyses should be prominent.

They should not be used for detailed ranking or target judgement before quality review is complete.

243

Final release

The final package should include the principal report, methodological documentation, tables, comparison classifications and accessible public communication. Microdata access conditions should be stated.

Release should occur only after numerical, metadata, confidentiality and interpretation checks agree.

244

Policy response

Responsible authorities should respond to the distribution, participation barriers and provision implications. The response should distinguish immediate action from matters requiring further evidence.

A measurement report should not itself claim that resources have been authorised or services delivered.

245

Evaluation of the system

After each cycle, the system should review coverage, response, language access, burden, errors, corrections, use and cost. It should identify which indicators informed decisions and which collections were unused.

Improvement should simplify as well as add. A field with no credible use should be retired.

246

Table 19: national architecture

Table 19. National adult literacy measurement architecture
ComponentPrincipal strengthPrincipal limitationRecommended periodic role
census modulepopulation breadth and local areashort self or proxy classificationlong-interval coverage benchmark
household modulecontext and recurring practicesindirect or declared capabilityinterim monitoring and diagnosis
direct assessmentproficiency distribution and domainscost, sample and participationperiodic depth and trend
programme recordsimplementation and participant progressno evidence on unserved populationcontinuing service management
qualitative enquirymeaning, barriers and mechanismsnot a prevalence estimatortargeted policy explanation
bridge studyestimates method discontinuityadditional sample and burdenwhenever major method changes
metadata registrypreserves definition and historyrequires active custodycontinuous control
public consultationrelevance and trustcannot replace statistical methoddesign and interpretation stages

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

247

Minimum viable architecture

Where resources are limited, the minimum system should maintain an improved household or census module, a probability validation subsample with a direct task, complete method metadata and a participation profile. Expansion to a multi-domain assessment should follow capacity and policy need.

The minimum should not be achieved by excluding languages or difficult-to-reach groups without disclosure.

248

Mature architecture

A mature system links periodic multi-domain assessment with census and household evidence, programme monitoring, language capacity, research access and policy evaluation. It preserves comparable trends while allowing controlled revision.

Maturity is demonstrated by evidence use, correction and public trust, not the number of instruments.

249

Architecture conclusion

A layered national architecture avoids demanding that one indicator provide small-area coverage, detailed proficiency, annual trend and programme attribution simultaneously. It assigns each method a defensible role and makes their relationships explicit.

This is the practical route from a single uncertain rate to a coherent public evidence system.

Part XII

Quality Evaluation and Decision Rules

250

Fitness for use

Quality should be judged against the decision an indicator is expected to support. A census declaration may be fit for broad local outreach and unfit for proficiency ranking; a direct assessment may be fit for national distribution and unfit for district allocation.

The quality statement should name both the supported and unsupported uses.

251

Relevance

Relevance concerns whether the construct, population, period and reporting level answer a current policy question. A technically strong measure can be irrelevant when it concerns the wrong language, age range or domain.

Policy users should participate in defining the question without controlling the result.

252

Accuracy

Accuracy concerns closeness between the reported result and the condition defined by the measurement design. It includes sampling and non-sampling error.

No single statistic establishes accuracy. Evidence comes from validation, consistency, field controls, item analysis, response patterns and comparison with appropriate external information.

253

Reliability

Reliability concerns consistency under repeated or parallel measurement conditions. It may be examined for items, scores, coders, interviewers and repeated questions.

High reliability does not establish that the correct construct was measured. A consistently applied narrow question can remain invalid for a broad claim.

254

Validity

Validity concerns whether evidence and theory support the intended interpretation and use. It encompasses content, response processes, internal structure, relationships with other variables and consequences.

Validity belongs to an interpretation, not permanently to an instrument title.

255

Timeliness

Literacy results should be released while they remain relevant, but speed should not displace necessary processing and review. Observation date and release date should both be visible.

Preliminary release is appropriate only where its scope and revision risk are clear.

256

Coherence

Coherence concerns whether related outputs use compatible definitions and tell an explainable account. Census, survey and programme results need not agree numerically, but their differences should be traceable to method and population.

Unexplained contradiction requires investigation rather than selective publication.

257

Accessibility and clarity

Users should be able to locate, understand and use the result and metadata. Technical documents, accessible summaries and machine-readable tables serve different audiences while preserving one approved meaning.

Clarity does not require omission of uncertainty.

258

Interpretability

Interpretability requires definitions, classifications, sources, methods, uncertainty and examples of task demand. Proficiency levels should be described through what the scale supports, not through moral or social labels.

Historical values require archived metadata to remain interpretable.

259

Comparability

Comparability should be rated for each principal contrast. A dataset may be internally comparable across regions but not over time after a method change.

The rating should appear in the analytical table, not only in a distant technical note.

260

Confidentiality and integrity

Quality includes protection of respondent information and resistance to unauthorised alteration. Confidentiality encourages trust and protects adults from harm; integrity protects the public record.[REF-09]

Both require assigned responsibility and documented incident response.

261

Cost-effectiveness

Cost should be considered in relation to decision value, precision, coverage and capacity developed. The least expensive indicator may become costly if it directs resources incorrectly.

Cost evaluation should include respondent burden and opportunity cost, not finance alone.

262

Error profile

An error profile records frame, sampling, non-response, measurement, processing, modelling and disclosure limitations. It identifies likely direction and magnitude only where evidence permits.

The profile should be revised as quality studies produce new information.

263

Critical errors

Some errors prevent the intended result from being released: lost sample-selection records, uncontrolled item exposure, unresolvable scoring corruption, absent population definition or a confidentiality breach affecting the publication.

Criticality should be defined before final review and linked to an authorised response.

264

Material errors

A material error could change a principal estimate, comparison, policy conclusion or public understanding. It requires correction or explicit qualification.

Materiality is substantive as well as numerical. A language omission can be material even where the national mean changes little.

265

Tolerable limitations

Some limitations can be accepted when bounded and unlikely to alter the intended decision. Acceptance should name the limitation, evidence, authority and use restriction.

Repeated acceptance should not become a substitute for improvement.

266

Verification rules

Verification should include sample selection, instrument version, response disposition, scoring, weight construction, table calculation and narrative claims. High-risk steps receive independent repetition or source checks.

Verification is complete only when a discrepancy is resolved and all affected outputs agree.

267

Replication

An authorised analyst should be able to reproduce principal estimates from protected data, documented code or calculation instructions and metadata. Replication records the software, weight, variance and exclusion rules used.

Public-use data may support wider replication where confidentiality allows.

268

Sensitivity analysis

Sensitivity analysis tests plausible alternative treatments of missing data, thresholds, flagged items, weights or population definitions. It shows whether the policy conclusion depends on one contestable choice.

Alternatives should be selected on methodological grounds, not to locate the preferred result.

269

External consistency

Literacy results may be compared with education, age, language and programme evidence to identify unexpected patterns. Consistency strengthens plausibility but does not prove correctness because related sources may share assumptions.

Unexpected differences require explanation and may reveal valuable new information.

270

Field indicators

Contact attempts, duration, refusal, break-off, assessor observations, language mismatch and deviations provide early quality evidence. They should be analysed during collection.

Field indicators should trigger support or verification without encouraging staff to pressure adults into participation.

271

Scoring indicators

Scorer agreement, missing responses, option patterns, item-total relationships and version effects should be monitored. A statistically unusual item requires content and administration review.

Automated flags do not determine deletion; an authorised methodological decision is required.

272

Weight indicators

Weight distributions, adjustment factors, effective sample sizes and contribution of extreme weights should be examined. Trimming trades variance reduction against potential bias.

The decision and effect on principal estimates should be recorded.

273

Publication indicators

Every reported percentage should reconcile to its denominator or approved estimation output. Titles, notes and narrative should use the same population and method.

Tables should be tested for suppressed cells, totals, rounding and version consistency.

274

Quality rating

A concise rating may communicate whether an estimate is fit, fit with qualification, experimental or not releasable for a stated use. The rating criteria should be public and accompanied by reasons.

One overall label should not conceal a critical failure in language, coverage or confidentiality.

275

Table 20: quality evaluation record

Table 20. Adult literacy indicator quality evaluation
DimensionEvidencePass conditionRestricted-use condition
relevancepolicy question and construct mapindicator answers stated decisionuse narrowed to actual construct
coveragetarget–frame reconciliationmaterial groups representedpopulation claim restricted
participationresponse flow and subgroup profileno unresolved severe selectivityresidual bias stated
measurementvalidity and reliability studiesinterpretation supporteddomain or language limited
estimationweights, variance and replicationprincipal estimates reproducibleexperimental model label
comparabilitymethod and population matrixintended contrast supportabledescriptive or qualified only
timelinessobservation and release datesrelevant after quality completionpreliminary status stated
claritymetadata and accessible accountmethod visible beside resultheadline withheld until corrected
confidentialitydisclosure reviewrisk controlledaggregation or access restriction
integrityapproval, version and correctionauthorised record completerelease suspended for critical failure

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

276

Release decision

Release should be authorised by the competent statistical body after technical and disclosure review. Policy disagreement with a valid result is not a quality failure.

An unreleasable estimate may still guide restricted internal investigation where confidentiality and uncertainty are controlled.

277

Correction decision

Correction is required when an error affects values, labels, comparison or interpretation materially. The notice should identify the original, corrected result and policy consequence.

New data or a later method are revisions or new observations, not corrections of the historical record.

278

Quality conclusion

Quality evaluation transforms a long list of desirable characteristics into a decision about a stated use. It protects against both uncritical publication and indefinite delay in search of impossible certainty.

The governing standard is transparent fitness for public purpose.

Part XIII

Strategic Findings and Conclusions

279

Finding 1: no single literacy rate is self-explanatory

Every rate embodies decisions about capability, population, language, respondent and threshold. A number detached from these decisions cannot support responsible comparison.

Method metadata should therefore be treated as part of the indicator itself.

280

Finding 2: methods provide different evidence

Self-report describes a response under a question; proxy report adds another person’s knowledge; schooling inference uses an associated characteristic; direct assessment observes selected performance.

Their differences are analytically useful and should not be erased by a common title.

281

Finding 3: the binary rate is an incomplete summary

A binary indicator can support communication and broad monitoring but conceals degrees, domains and uncertainty near the threshold. Contemporary literacy evidence supports a continuum and plural practices.[REF-02] [REF-04]

Policy should retain distributional evidence where available.

282

Finding 4: educational attainment is not literacy

Education contributes strongly to literacy, but attainment levels do not determine present proficiency. School quality, language, practice and time vary.[REF-06] [REF-07]

Attainment and literacy should be analysed jointly without substituting one for the other.

283

Finding 5: direct assessment has boundaries

Direct assessment improves observation and distributional analysis but remains bounded by tasks, languages, frame and participation. Technical sophistication does not remove social and operational error.

The result should identify adults whose performance was not observed.

284

Finding 6: participation can bias the estimate

Adults facing the greatest literacy barriers may be hardest to reach or least willing and able to complete an assessment. An achieved-sample rate can therefore be systematically favourable.

Response adjustment helps but does not prove removal of bias.

285

Finding 7: language defines evidence

Performance in one language is not evidence of literacy in every language. Conversely, inability to use the assessment language is not proof of absence of literacy.

Language mapping, adaptation and reporting are central validity controls.

286

Finding 8: cross-national ranking is often too strong

Accurately collected national values can remain incomparable because methods, questions and populations differ. A ranked table gives a stronger impression than the evidence warrants.

Comparability classes and separate panels provide a more responsible account.

288

Finding 10: uncertainty extends beyond standard error

Sampling uncertainty can be calculated, but coverage, non-response, language and construct limitations also affect conclusions. They require qualitative and sensitivity evidence.

Reporting only a confidence interval can create false completeness.

289

Finding 11: local need exceeds local precision

Policy often requires district evidence that a national assessment cannot estimate reliably. Census modules, model-assisted estimates and community evidence may help, with distinct status.

Unstable direct ranks should not control material allocation.

290

Finding 12: programme evidence has a narrower population

Learner assessment can show change among participants, but attrition and selection limit inference. A national indicator cannot serve as a programme comparison by coincidence of timing.

Evaluation requires a design aligned with programme exposure and intended result.

291

Finding 13: layered systems are preferable

Census breadth, household context, direct assessment depth, programme records and qualitative enquiry are complementary. A layered architecture assigns each an explicit role.

Integration should preserve method identity rather than construct one opaque composite.

292

Finding 14: statistical independence and participation coexist

Communities and policy bodies should contribute to relevance, language and interpretation. The statistical authority remains responsible for scientific method, confidentiality and impartial release.[REF-09]

Neither technical isolation nor political control produces trusted evidence.

293

Finding 15: measurement should lead to provision

The public purpose of adult literacy statistics is improved opportunity to learn, use and sustain literacy. A measurement programme that repeatedly documents disparity without informing provision is incomplete.

Policy response should be monitored separately from the indicator.

294

Immediate priority 1: classify existing indicators

Countries should inventory literacy values by self, proxy, indirect, short direct and multi-domain direct method. Each record should include population, year, language and question or framework.

This inventory can immediately prevent unsupported trend and rank claims.

295

Immediate priority 2: publish exact metadata

Current releases should place essential method information beside principal values. Historical metadata should be recovered where feasible and uncertainty acknowledged where it cannot be found.

Unknown method is itself a reason to restrict comparison.

296

Immediate priority 3: analyse participation

Existing surveys should publish contact, cooperation and assessment completion by available group. Language and disability-related non-completion should not remain inside a general missing category.

Field operations should use the findings in the next cycle.

297

Immediate priority 4: stop automatic schooling substitution

Where education is used to assign literacy, the assumed component should be separately counted and tested in a subsample. No-schooling adults should also receive an opportunity to demonstrate capability.

This protects both statistical validity and dignity.

298

Immediate priority 5: preserve trend breaks

Statistical publications should mark every material change of method and refrain from subtracting unlike values. Bridge designs should be planned before the old method is discontinued.

Improved measurement should not be discouraged by fear of an apparent setback.

299

Medium-term priority 1: direct-assessment capacity

Countries requiring proficiency evidence should develop national sampling, language, field, scoring and analysis capacity. LAMP provides a contemporary international direction for this work.[REF-03]

Capacity development should include dissemination and policy use, not collection alone.

300

Medium-term priority 2: language inclusion

Written-language evidence should guide which versions are developed and how samples support analysis. Exclusions and limitations should be public.

Language inclusion should be planned with communities and specialists from the beginning.

301

Medium-term priority 3: method-linking studies

Validation and bridge samples should examine self, proxy and direct evidence in the same population. They should test differential relationships by relevant group.

Results should explain disagreement rather than promise a universal conversion.

302

Medium-term priority 4: recurring household context

Stable household modules can monitor practices, perceived difficulty, language and access between assessment cycles. They should be concise and linked to clear decisions.

Changes in wording or mode require control.

303

Medium-term priority 5: research access

Protected access to microdata and documentation can strengthen replication and methodological research. Disclosure controls should preserve confidentiality while avoiding unnecessary barriers.

Independent analysis should be answered with evidence, not institutional defensiveness.

305

Longer-term priority 2: environments for literacy

Measurement should increasingly examine opportunities to use literacy in work, family, public services and civic life.[REF-02] [REF-04]

Policy can then address both adult learning and institutional demands that exclude people through unnecessary complexity.

306

Longer-term priority 3: evaluation of literacy policy

Population indicators should be integrated with implementation and evaluation evidence so that policy contribution can be judged without confusing association and causation.

Equity, sustainability and unintended effects should remain part of evaluation.

307

Table 21: strategic action schedule

Table 21. Strategic priorities for adult literacy measurement
HorizonPriorityResponsible institutionsEvidence of progress
immediateindicator method inventorystatistical and education authoritiescomplete metadata register
immediateparticipation profilessurvey and assessment bodiesstaged response tables
immediatetrend-break publicationstatistical publishersannotated historical series
immediateschooling-assumption disclosurecensus and survey bodiesseparate assumed counts
medium termdirect-assessment capacitynational statistical and assessment bodiesvalid national proficiency study
medium termlanguage inclusionlanguage communities and assessment bodytested versions and coverage record
medium termbridge studiesstatistical and research institutionsmethod-effect estimates
medium termhousehold context modulesurvey and policy bodiesstable recurring indicators
longer termcomparable proficiency trendnational and international partnerslinked repeated assessments
longer termpolicy evaluationresponsible ministries and independent evaluatorscontribution findings and response

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

308

Final conclusion

Adult literacy measurement is a public judgement about what capability is observed, whose experience is represented and which comparison is justified. The apparent simplicity of a national rate conceals choices about language, schooling, respondent, task and threshold. Those choices should be visible.

Self-reported data remain useful for broad coverage, perceived difficulty and literacy practices when exact wording, proxy status and missingness are known. They cannot establish the same result as a direct assessment. Direct assessment provides richer proficiency evidence but must answer for the adults, languages and settings it fails to include.

The responsible course is not to select one method as universally authoritative. It is to construct a layered system, assign each method a defined role, bridge changes where possible and classify comparisons by the strength of equivalence. Statistical uncertainty, participation and dignity belong inside the indicator, not outside it.

The resulting evidence can support the commitments of education for all and the international Literacy Decade only when it leads to learning opportunities and environments in which adults can use literacy. Measurement earns authority through clarity, restraint and public use.[REF-01] [REF-15]

Part XIV

Indicator Specifications for Practical Use

309

Purpose of the specifications

The specifications below define a limited family of indicators that may coexist without being treated as interchangeable. Each identifies the evidence, denominator and permitted decision. National authorities should adapt language and population details while preserving method identity.

310

Declared reading-and-writing status

The numerator is adults answering that they can read and write under the exact question; the denominator is all eligible adults, with unknown status separately reported. Self and proxy responses should be distinguished.

Use is limited to declared status and broad coverage analysis. It does not establish proficiency or separate reading from writing where the question combines them.

311

Known-response declaration rate

This rate divides declared yes by yes plus no, excluding unknown. It is useful for comparison with the all-eligible rate and for examining the numerical effect of missing status.

It should never replace the primary coverage account without evidence that unknown responses are ignorable.

312

Proxy-response share

The numerator is adult literacy records supplied by another household member; the denominator is all adult literacy records. Distribution by age, sex and absence status helps assess method variation.

The indicator is a quality characteristic, not a correction factor.

313

Graded self-assessed reading difficulty

Adults report whether selected reading tasks are performed easily, with difficulty or not at all. Exact materials and reference period should be stated.

This supports service-access and perceived-difficulty analysis. It is not a proficiency scale unless validated for that use.

314

Graded self-assessed writing difficulty

Writing is measured separately through defined activities such as completing a form or composing a short message. Response categories and assistance should be recorded.

The measure can reveal domain differences hidden by a combined literacy question.

315

Literacy-practice frequency

The indicator reports how often adults engage with specified texts in work, household, learning or community settings during a defined period. “Never” can reflect absence of opportunity, preference or difficulty.

Practice is relevant to maintenance and policy context but does not demonstrate task performance.

316

Whole-sentence reading performance

The numerator is adults who read the specified sentence fully under the administration and language rule; the denominator is all eligible adults offered valid access. Partial reading, non-performance and non-completion remain separate.

The claim is restricted to the sentence task and cannot include writing.

317

Partial-sentence performance

This indicator retains adults able to read part but not all of the sentence under the scoring rule. It prevents a potentially important capability range from disappearing into a binary rate.

Interpretation requires consistent probing and scorer guidance.

318

Assessment participation rate

Valid assessment completers are divided by all selected eligible adults. Background-only completion, refusal, language exclusion, accessibility exclusion and interruption are component measures.

The rate informs potential bias and should accompany every proficiency result.

319

Assessment-language coverage

The numerator is target-population adults for whom an authorised assessment language is available; the denominator is the defined target population. Estimates may use frame or survey evidence.

Availability does not prove version equivalence or actual participation.

320

Mean proficiency score

The weighted mean summarises the score distribution for a covered population under the scaling model. It should be reported with standard error and scale interpretation.

The mean should not stand alone where lower-tail need or distributional inequality drives policy.

321

Proficiency-level distribution

The indicator reports weighted shares within approved score ranges. Counts, standard errors and level descriptions accompany the distribution.

Levels are reporting categories and should not be converted into fixed identities for adults.

322

Below-selected-level share

This share may support a defined policy question, such as the population likely to face difficulty with a specified range of tasks. The selected level and rationale should be explicit.

Changing the threshold changes the population estimate; no universal deficit category should be implied.

323

Lower-tail percentile

A lower percentile score can show the position beneath which a stated share of the covered population falls. It avoids some dependence on an arbitrary level boundary.

Precision can be weak in small samples, and task interpretation remains necessary.

324

Literacy distribution by sex

The indicator reports the selected measure separately for women and men under identical method conditions, with uncertainty. Survey participation and proxy status should also be compared.

The difference describes inequality; explanation requires evidence about education, language, work and social conditions.

325

Literacy distribution by age

Age groups should be defined consistently and reported with observation date. Results describe cohort and age differences at one time unless longitudinal evidence exists.

They should not be labelled skill loss solely from cross-sectional contrast.

326

Literacy distribution by educational attainment

Assessment or declaration results are tabulated by harmonised education level. The analysis shows variation within and between attainment categories.[REF-10]

It should demonstrate why education cannot be used as an automatic literacy classification.

327

Literacy distribution by language

Results may be grouped by home language, language of literacy practice or assessment language. These variables should not be conflated.

Assessment-language groups are self-selected or assigned populations and should not be ranked causally without an adequate design.

328

Literacy distribution by location

Urban–rural or regional estimates require stable definitions, sample support and appropriate variance. Geographic categories should reflect policy and remain comparable over time.

Local publication should avoid disclosure and unstable rank.

329

Self–direct agreement rate

The indicator reports the proportion for which a defined self-report category agrees with a defined direct-assessment category in a linked sample. Both categories and the assessment threshold must be stated.

Agreement does not establish that the constructs are identical or that disagreement is dishonesty.

330

Proxy–self agreement rate

Adults’ own responses are compared with independent proxy responses under a controlled design. Results may vary by relationship and household absence.

The study informs proxy quality but should not expose disagreement within named households.

331

Method discontinuity estimate

In a bridge sample, the difference between old and new methods is estimated under aligned population and period. Subgroup variation and uncertainty should be shown.

The estimate may support historical interpretation but is not necessarily stable across future cohorts or settings.

332

Coverage ratio

The estimated frame population is divided by the target population under aligned definitions. Known omissions and duplicates are reported separately.

A high overall ratio can conceal complete omission of a small policy-relevant group.

333

Contact rate

Selected eligible units successfully contacted are divided by all units known or estimated eligible. Unknown eligibility requires a stated treatment.

The indicator separates frame access from cooperation and assessment performance.

334

Cooperation rate

Adults agreeing to the relevant interview or assessment are divided by contacted eligible adults. The outcome should distinguish informed refusal, gatekeeper refusal and other non-cooperation.

Pressure to improve the rate must not compromise voluntary participation or respect.

335

Assessment break-off rate

Adults who begin but do not provide valid completion are divided by assessment starters. Reason and point of break-off should be analysed.

The rate can identify burden, difficulty, language or administration problems but does not alone assign cause.

336

Weight-adjustment effect

The difference between base-weighted and final-weighted principal estimates shows the effect of adjustment. It should be reported with the variables and weight stages used.

A large effect prompts review; a small effect does not establish absence of non-response bias.

337

Effective sample size

The effective sample size expresses loss of precision from clustering and unequal weights relative to a simple design. The exact definition should be stated.

It supports decisions about subgroup publication and future design, not a replacement for design-based variance.

338

Relative standard error

The standard error divided by the estimate provides one indicator of relative precision. It can be unstable near zero and should be interpreted with absolute uncertainty.

Publication rules may use it alongside unweighted count and disclosure controls.

339

Programme reach

Eligible adults participating in a literacy programme are divided by the defined eligible or intended population. Estimating that denominator may require local evidence.

Reach should be disaggregated and should not be confused with completion or learning.

340

Programme retention

Participants remaining through a defined instructional point are divided by entrants. Transfers, planned exits and missing records should be classified.

Retention is a participation result and does not establish proficiency gain.

341

Matched learning change

The difference between comparable baseline and follow-up measures is reported for participants with both observations. The matched count and attrition profile accompany it.

The result applies to matched participants unless an adjustment with defensible assumptions broadens inference.

342

Public-service text access indicator

Adults may report or demonstrate access to defined essential written information, while services record availability of accessible formats and assistance. The indicator should specify the service event.

It connects literacy policy to institutional responsibility rather than locating every barrier in the adult.

343

Table 22: indicator-use register

Table 22. Practical indicator-use register
Indicator familyEvidence classPrimary useMandatory companion
declared statusself or proxy reportbroad coverage and outreachrespondent and unknown shares
graded difficultyself-assessmentservice access and perceived needexact task wording
literacy practicereported behaviourcontext and opportunityperiod and text type
sentence performanceshort direct tasknarrow verificationpartial and non-completion categories
proficiency distributionmulti-item direct assessmentpopulation skill distributionparticipation, language and uncertainty
subgroup distributionany stable measureequity analysissample, method and confidentiality
method agreementlinked measuresvalidation and transitionthreshold and overlap design
field participationoperational recordsbias analysistarget and selected populations
programme reachprogramme and population dataservice accesseligibility definition
matched changerepeated participant assessmentlearning among observed participantsattrition and comparison limits

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

344

Specification conclusion

The indicator family provides complementary evidence without a false common scale. Selection should follow the decision, and every value should retain the companion information necessary to understand its population and method.

Part XV

Decision Scenarios for Statistical Authorities

345

Scenario 1: an unexplained high rate

A newly submitted census rate is substantially higher than the previous observation. Before publication, the authority should examine question wording, schooling assumptions, proxy share, unknown treatment, age limits and geographic coverage.

The value should remain pending if the metadata cannot explain its method. Plausibility against expectation is not sufficient verification.

346

Scenario 2: a lower direct-assessment result

A direct assessment produces a lower share above a selected threshold than the existing self-declaration rate. The authority should publish the two as different measures and explain the constructs.

It should not retrospectively revise the earlier declaration rate or call the direct result a decline.

347

Scenario 3: missing language version

Field preparation reveals that an intended language version cannot meet validity and staffing requirements by the start date. The authority should consider staged collection, adjusted population scope or delay.

Using another language by convenience and coding non-performance as low skill is not acceptable.

348

Scenario 4: severe non-response in one region

National response is adequate, but one region has high contact failure and assessment break-off. Additional fieldwork should preserve selected cases, and the regional estimate should be withheld or qualified if selectivity remains.

The national estimate also requires sensitivity analysis according to the region’s population weight.

349

Scenario 5: item security breach

Evidence shows that some assessment tasks circulated before administration in selected areas. The authority should identify exposure, test performance anomalies and determine whether affected items or cases can support inference.

Replacement or exclusion requires a pre-authorised technical decision and an account of scale comparability.

350

Scenario 6: scoring disagreement

Monitoring finds low scorer agreement on constructed writing responses. Release of the affected domain should pause while rules, retraining, moderation and rescoring are completed.

Other independent domains may proceed if their validity and publication meaning are unaffected.

351

Scenario 7: political request to change the threshold

A ministry requests a lower threshold so the reported target is attained. A threshold may be reviewed for substantive policy reasons, but the original result and target remain on their approved basis.

Any new threshold produces a separate indicator and cannot rewrite historical performance.

352

Scenario 8: unstable district rankings

Local leaders request an ordered table of district rates. Where confidence intervals overlap substantially and sample counts are low, the authority should publish appropriate estimates and uncertainty without an ordinal league table.

Planning categories may be formed from multiple evidence sources under a transparent rule.

353

Scenario 9: proxy responses dominate

A census review finds that most working-age adult records were supplied by proxies. The data remain census evidence but should be labelled accordingly. A validation subsample can estimate disagreement.

The next operation should consider timing and respondent procedures that increase self-response without reducing coverage.

354

Scenario 10: programme data exceed population need

Provider records show more annual enrolments than the estimated local adult population below a literacy threshold. The figures may include repeat enrolments, commuters, unlike age ranges or an inappropriate prevalence estimate.

Reconciliation should precede any claim of universal programme reach.

355

Scenario 11: preliminary results change after weighting

Initial unweighted results differ materially from final weighted estimates. Public communication should use the final design-consistent values and explain that preliminary tabulations did not represent the population.

The difference is a reason for correct weighting, not evidence of manipulation.

356

Scenario 12: new census question

Cognitive testing supports a clearer graded question, but it differs from the previous binary item. The authority should adopt the better question where warranted and plan an overlap or bridge.

Trend continuity is valuable but should not preserve an inadequate measure indefinitely.

357

Scenario 13: adult refuses an assessment

The assessor should record informed refusal and conclude contact respectfully. No lowest score is assigned. Available background data may be retained only under the consent and confidentiality arrangements.

Weighting and bias analysis address population inference; coercion does not.

358

Scenario 14: accommodation changes mode

An adult requires a presentation or response method not covered by the standard design. The responsible specialist should determine whether the construct is preserved and whether the result belongs on the common scale.

If not, the adult’s exclusion from the standard estimate is recorded and a suitable descriptive assessment may still inform provision.

359

Scenario 15: correction after release

A coding error changes one subgroup estimate and a national total by 0.3 percentage points. Materiality depends on affected conclusions, not size alone. All tables, narrative and policy uses should be traced.

A dated correction should state which findings remain and which change.

360

Scenario 16: conflicting official values

Two public agencies release different adult literacy rates for the same nominal year. A reconciliation statement should compare observation year, population and method and assign each value an authorised use.

The goal is not necessarily to choose one number, but to prevent incompatible meanings from being presented as a factual dispute.

361

Table 23: decision authority and response

Table 23. Statistical decision scenarios and authorised responses
ScenarioImmediate authorityRequired evidenceRelease consequence
unexplained ratestatistical authoritysource and method verificationhold pending metadata
method differencestatistical and assessment bodiesconstruct and population mappublish separately
missing languagefield and technical governancevalidity and coverage optionsstage, restrict or delay
regional non-responsefield and estimation leadsresponse profile and sensitivityqualify or withhold region
item exposureassessment security and technical panelexposure and item analysisexclude, rescale or suspend
scorer disagreementscoring leadagreement and rescoring evidencepause affected domain
threshold requeststatistical authorityapproved construct and policy rationaleretain original series
unstable rankingpublication authorityuncertainty and sample rulesremove rank
coding correctioncorrection authoritytrace and recalculationissue controlled correction
conflicting valuesinter-agency statistical committeemethod reconciliationassign distinct uses

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

362

Decision-scenario conclusion

The scenarios demonstrate that quality governance depends on decisions made before and after fieldwork, not only on calculation. A predetermined authority, evidence requirement and release consequence shorten response and protect the public record.

363

Synthesis of measurement trade-offs

Adult literacy measurement presents recurring trade-offs that should be decided openly. Breadth can be increased through short census questions, but a shorter instrument offers less evidence about domains and proficiency. Depth can be increased through direct assessment, but time, language development, sampling and participation become more demanding. Frequency can be increased through recurring modules, but repeated design change can weaken the series more than an additional observation strengthens it.

There is also a trade-off between standardisation and relevant access. A common instrument supports comparison, yet literal uniformity can introduce irrelevant difficulty where language, script or document conventions differ. Adaptation improves access only if the intended construct and task demand remain sufficiently equivalent. The decision requires linguistic, substantive and empirical evidence; neither universal sameness nor unrestricted local variation is adequate.

Small-area demand creates another tension. National assessment samples can provide reliable distributions but insufficient local precision. Censuses can provide local declaration data without direct proficiency. Model-assisted estimates can fill part of the geographic gap but depend on assumptions. A responsible system presents these sources as complementary layers and does not grant the strongest methodological label to the estimate with the smallest geographic unit.

Continuity and improvement also compete. Stable wording protects trend, while an inadequate question can perpetuate weak evidence. Method renewal should proceed through testing, overlap and a visible break. A country should not be penalised analytically for improving measurement, but it should not convert the new result into an artificial historical trend.

Finally, measurement ambition must be balanced with respondent dignity and public trust. Longer assessment and repeated contact can add information, but burden and anxiety may selectively reduce participation. Confidentiality encourages cooperation but does not justify secrecy about methods. Accessible communication expands public use but must retain qualifications.

These trade-offs do not imply that every design is equally defensible. They require a reasoned choice linked to purpose, resources and consequence. The acceptable design is the one that preserves the essential population and construct, measures uncertainty, protects participants and states where interpretation must stop.

364

The standard for a defensible national estimate

A national estimate is defensible when its population can be reconstructed, its evidence method is named, its language conditions are visible, its participation losses are examined and its numerical uncertainty is reported. These conditions are cumulative. A precise calculation from a partial frame does not become national through weighting alone; an inclusive frame does not establish proficiency through an undefined question.

The standard is proportionate to use. Broad programme outreach may proceed from a stable census declaration with known limitations. A claim about the distribution of adult proficiency requires direct evidence and a probability design. A judgement about programme effect requires evidence of exposure and a comparison suited to causation. The same literacy label does not equalise these thresholds.

Defensibility also requires institutional independence and public explanation. Technical decisions should be made by the competent statistical authority, recorded before results are known where feasible, and open to methodological scrutiny. Policy bodies may set priorities and respond to findings, but they should not select the value or threshold that best satisfies an existing commitment.

Where the evidence falls short, the correct response is a restricted claim, a qualified estimate or further measurement. Withholding a precise but unsupported rank is not a failure of statistics. It is evidence that the system distinguishes what is known from what remains uncertain.

365

Closing public-interest test

Before an adult literacy indicator is adopted, the responsible authority should ask whether the measure will improve understanding of educational need, whether the population most affected can participate, whether language and disability are treated as measurement conditions, and whether the published result can be used without unjustified stigma. A method that is efficient but systematically excludes the intended beneficiaries fails this test.

The authority should also identify the decision that follows. If no institution is responsible for provision, communication or further investigation, repeated measurement may create visibility without remedy. Conversely, an urgent local response may be justified by credible evidence even when a nationally comparable proficiency estimate is not yet available.

The public-interest test therefore joins statistical fitness with institutional use. It does not allow policy urgency to weaken evidence, nor does it allow technical uncertainty to become a reason for inaction. It requires an honest account of what is known, a proportionate decision and a plan to close material evidence gaps.

366

Final evidentiary boundary

This report supports decisions about the design, interpretation and institutional use of adult literacy indicators. It does not supply a universal conversion between declaration and assessment, a single threshold valid for every policy, or a substitute for national linguistic and population evidence. Application should preserve these boundaries and record any departure required by local conditions.

The central requirement remains stable: a public result should identify whose literacy was considered, how evidence was obtained, which language and task conditions applied, what uncertainty remains and which decision the evidence can reasonably support.

Part XVI

Measurement Conditions across Diverse Settings

367

Purpose of contextual differentiation

Global comparison should recognise that national statistical capacity, linguistic plurality, settlement patterns and adult-learning systems differ. Common principles can govern definitions and disclosure while operational designs respond to context.

Context should explain design choices; it should not excuse an unsupported population claim.

368

High census coverage with limited assessment capacity

Where a census is the principal source, priority should be given to exact wording, proxy identification, language rules, unknown status and a small validation study. This can improve interpretation without immediately requiring a large specialised assessment.

The census result remains declared or reported literacy.

369

Established household-survey systems

Countries with recurring household surveys can add a concise literacy module, rotate practice questions and periodically administer direct tasks to a probability subsample.[REF-11] [REF-13]

Stable core questions protect trend, while controlled modules can address new policy needs.

370

Strong assessment capacity

Where technical and financial capacity support a multi-domain assessment, the design should still examine frame coverage, language access and selective completion. High psychometric quality within the achieved sample does not settle population representation.

Investment should include analysis and policy use, not fieldwork alone.

371

Multilingual national settings

Language selection should follow written-language evidence, population size, policy purpose and consequences of exclusion. Several versions may require larger samples and longer development.

A national result can combine valid versions where the scale supports it, while language-specific interpretations remain conditional on population differences.

372

Low-density rural populations

Sampling remote communities may be expensive but substantively necessary. Oversampling, extended field periods and local-language assessors may be required.

Excluding remote areas and labelling the result national can misdirect precisely the policies concerned with unequal access.

373

Highly mobile populations

Seasonal work, migration and unstable residence weaken household contact and usual-residence classification. Field timing and repeated contact should reflect known mobility.

The report should distinguish frame omission, temporary absence and movement outside the target territory.

374

Post-conflict and disrupted settings

Population frames, infrastructure, language relations and trust may be weak. A staged design may begin with area mapping, focused surveys and qualitative evidence before a national estimate is credible.

Measurement should not delay urgent adult-learning provision supported by existing evidence.

375

Small island and small-population settings

A census or near-census assessment may be feasible, but confidentiality and respondent burden become acute. Sampling error may be small while non-response and disclosure risk remain material.

Regional pooling requires attention to language, population and method equivalence.

376

Urban informal settlements

Rapid growth and irregular addresses can cause frame undercoverage. Updated area listing, local mapping and varied contact hours may improve inclusion.

Service-based samples can inform barriers but do not replace probability population estimates without a defensible frame.

377

Populations with limited schooling

Routing adults away from assessment because they lack schooling reproduces an untested assumption. Instruments should include accessible entry tasks that permit capability to be demonstrated.

Task difficulty should extend low enough to describe emerging proficiency without humiliation.

378

Populations with high formal attainment

High school completion does not remove the need for assessment where policy concerns adult proficiency and practice.[REF-06] [REF-07]

Ceiling coverage should be adequate to distinguish advanced performance, and assumed literacy should not replace observation.

379

Ageing populations

Upper-age exclusions may remove a growing share of adults and weaken service planning. Assessment burden, sensory access and cohort education require careful treatment.

Age differences should be separated from longitudinal change and reported with the actual age range.

380

Youth transition to adulthood

Results for ages 15–24 connect school quality, early work and continuing learning. School attendance and assessment language may affect interpretation.

Youth literacy should not be used as a direct substitute for adult rates or school-learning measures.

381

Gender-restricted participation environments

Interview timing, interviewer sex, privacy and household permission can affect women’s or men’s participation. Field protocols should protect direct informed response and record gatekeeper refusal.

An apparently small sex gap may reflect differential proxy reporting or exclusion.

382

Disability-inclusive population estimates

Frame and instrument design should include adults with disabilities through accessible contact and administration. Exclusions should be counted by reason and should qualify the population estimate.

Alternative modes belong on the common scale only where construct and measurement evidence support that use.

383

Limited statistical resources

A smaller probability sample with sound selection, tested questions and complete metadata is preferable to a larger uncontrolled collection. Existing survey infrastructure can reduce cost if its population and field methods are fit.

Technical assistance should transfer reproducible capability and documentation.

384

Decentralised statistical systems

Regional bodies may collect literacy data under different languages and procedures. A national framework should establish common core definitions, version control and quality evidence while allowing justified local modules.

Central aggregation should occur only after method reconciliation.

385

Administrative dependence on programme data

Where programme records dominate the evidence base, authorities should distinguish participation, completion and assessed learning from population prevalence. Unserved adults are absent by definition.

A household or community sampling component is required to examine unmet need.

386

Transition from binary to continuum measures

Countries moving toward direct proficiency assessment should preserve the binary series as a historical declared indicator and conduct an overlap study. The new distribution can answer stronger questions without being forced into the old categories.

Public targets may require formal rebasing.

387

Table 37: contextual design choices

Table 37. Measurement conditions and proportionate design choices
Setting conditionPriority design responseEvidence retainedClaim requiring restraint
census is principal sourceimprove wording, proxy record and validationlocal declared statusdirect proficiency
recurring household surveysstable module and direct subsamplecontext and linked methoduniversal conversion
multilingual populationplanned versions and language samplinglanguage-specific accesscausal language ranking
remote populationoversample and extend fieldworkterritorial coveragenational claim after exclusion
mobile populationupdated frame and varied contactabsence and movement statusstable-resident assumption
small populationintensive coverage and disclosure controldetailed national evidenceunsafe local cells
limited schoolingaccessible low-demand tasksemerging proficiencyschooling-based assignment
ageing populationinclusive age and accommodationolder-adult distributionage difference as skill loss
limited resourcesfocused probability designdefensible bounded estimatelarge uncontrolled rate
decentralised collectioncommon core and reconciliationregional method recordaggregation before equivalence

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

388

Contextual decision rule

Contextual adaptation is justified when it improves representation or construct relevance without obscuring method. The authority should record the condition, chosen response, evidence of adequacy and effect on comparison.

An adaptation that changes the observed capability may still be useful for local policy, but it should receive a separate indicator identity.

389

Contextual conclusion

Internationally credible measurement does not require every country to use one instrument. It requires common discipline in population definition, method identity, language access, uncertainty and public reporting.

These controls allow diversity of setting to enter the evidence without making every national value incomparable by default.

Part XVII

Interpretations That the Evidence Does Not Support

390

Literacy as a fixed personal identity

No indicator in this report supports treating literacy as an immutable identity. Capabilities can develop, be sustained or weaken according to learning and opportunities for use.[REF-02] [REF-06]

Categories describe evidence under defined conditions and should not determine an adult’s social worth or capacity to learn.

391

One task as complete capability

Successful reading of a sentence does not establish writing, document use, numeracy or engagement with extended text. Failure does not establish absence of every literacy capability.

The task statement should remain as narrow as the evidence.

392

Self-report as deliberate misstatement

Disagreement with direct assessment can arise from different constructs, thresholds, language, context and reference standards. It should not be interpreted automatically as deception.

Validation should explain patterns and improve questions rather than assign blame.

393

School completion as demonstrated proficiency

Formal attainment documents educational participation and completion under a system. It does not directly observe current adult literacy.[REF-06] [REF-10]

Using it as a convenient assumption changes the method and should remain visible.

394

Group difference as inherent deficiency

Differences by language, sex, location, age or poverty do not establish an inherent characteristic of the group. Historical educational access, opportunity, institutional demand, migration and measurement conditions may contribute.

Interpretation should direct attention to remediable barriers and provision.

395

Assessment-language score as language quality

Average performance among adults assessed in different languages does not establish that one language is more capable of supporting literacy. Populations, versions and opportunities differ.

Version validity and social explanation are separate analytical questions.

396

Association as individual destiny

Population associations between literacy and employment, income or health do not determine the outcome of a particular adult. They also do not show that raising one score alone will produce the associated outcome.

Public communication should avoid deterministic language.

397

Statistical significance as policy importance

A precise small difference may be statistically detectable and educationally minor. A substantively important difference may be estimated imprecisely in a small group.

Policy judgement should consider magnitude, uncertainty, affected population, equity and consequence together.

398

Absence of significance as equivalence

Failure to reject a statistical null does not demonstrate that two groups or periods are equivalent. The sample may lack power or the interval may include important differences.

Equivalence requires a defined margin and suitable design where that claim is needed.

399

National average as universal experience

A national mean or rate can conceal a lower tail, local exclusion and language disparity. It cannot describe every adult or community.

Distributional and participation evidence should accompany aggregate policy use.

400

Better measurement as declining literacy

Introduction of direct assessment, improved coverage or removal of schooling assumptions may lower a reported value. The difference may reflect correction of the evidence base rather than deterioration.

The series should mark the method change and avoid assigning it to population time.

401

Higher response as absence of bias

A high response rate reduces but does not eliminate non-response bias. A small non-responding group can be highly distinctive, and frame undercoverage exists before response is measured.

Response level, pattern and relationship to literacy all matter.

402

Weighting as full representation

Weights align respondents with known population variables under assumptions. They cannot create observations for an omitted language, inaccessible mode or absent frame group.

Residual bias should remain part of the conclusion.

403

Proficiency level as individual diagnosis

Population assessments may provide limited precision for one adult, especially under item-sampling designs. A level estimate should not be reused for individual placement, employment or benefit decisions without separate validity and due process.

Population monitoring and individual assessment are different purposes.

404

Programme completion as learning

Completion establishes participation through a defined point. It does not establish that intended capability changed or was sustained.

Learning evidence, assessment coverage and attrition remain necessary.

405

Population change as programme impact

A national indicator can change through cohorts, migration, schooling, social conditions and multiple programmes. Coincidence with one intervention does not establish attribution.

Programme evaluation requires evidence closer to exposure and a credible alternative account.

406

Comparability label as permission for every use

A pair of values may be comparable for broad distribution and not for precise ranking, subgroup analysis or causal inference. The authorised use must accompany the classification.

Comparability is always tied to the question.

407

International standardisation as uniform language

Common concepts and methods do not require the same language or culturally irrelevant documents. Adaptation can be necessary for equivalent access and task demand.

The evidence standard is disciplined equivalence, not identical appearance.

408

Technical detail as public inaccessibility

Complex survey and assessment methods require full technical records, but principal public meaning can still be communicated clearly. Method, population and uncertainty should not disappear from accessible accounts.

Clarity and rigour are compatible responsibilities.

409

Uncertainty as a reason for no action

Policy often proceeds under bounded uncertainty. Credible evidence of severe exclusion may justify immediate provision while stronger measurement is developed.

The action should be proportionate, its evidence basis explicit and its result monitored.

410

Final interpretive safeguard

The recurring safeguard is to keep the public claim at the level of the evidence. When a result concerns declared status, it should say so; when it concerns performance on selected tasks, it should name them; when the population is restricted, it should identify the restriction.

This precision is the foundation of internationally credible adult literacy reporting.

411

Implications for the next measurement cycle

The next cycle should begin from the weaknesses identified in the current evidence rather than from a general ambition to collect more data. Where proxy response is extensive, field timing and respondent selection should be improved. Where language coverage is incomplete, version development and sampling should begin early. Where direct assessment completion is selective, burden, access and trust require operational attention before another estimate is commissioned.

Continuity should be protected through stable core definitions and planned bridge evidence. Renewal remains necessary where a binary question, schooling assumption or narrow task no longer answers the policy question. The old and new measures can coexist during transition, with each assigned a clear status and use.

Institutional arrangements also require continuity. Sampling, language, scoring, analysis and public communication should not depend on knowledge held by temporary personnel. Controlled documentation, national technical teams and transparent review preserve capacity between cycles.

The next measurement cycle will be successful only if its evidence is used. Authorities should state which findings led to changes in adult-learning provision, accessible public communication, language materials or outreach, and which questions remain unresolved. This closes the relationship between statistical observation and the public commitment to literacy.

Results from the next cycle should be compared with this evidence only after the population, construct, language and method have been reconciled. A more recent value is not necessarily a more comparable value. The integrity of the series depends on preserving the observation that actually occurred and explaining the relationship between successive designs.

This discipline permits measurement to improve without losing institutional memory.

It also ensures that policy progress is judged against evidence of the same meaning, not against a numerical resemblance created by incompatible methods. Where equivalence is absent, the responsible conclusion is a documented break and a new baseline. That conclusion should remain visible in every subsequent public comparison using the series.

A

Literacy indicator metadata dictionary

Each indicator receives a unique identifier, title, version, responsible body, approval date, observation period and release status. Superseded versions remain available. The identifier follows the value into every table and analytical extract.

The record states reading, writing, numeracy or other domain; declared, practised or demonstrated capability; text and context; response or task demand; classification; and permitted interpretation. An undefined entry is not replaced by the general word literacy.

Age, upper-age limit, usual-residence rule, territory, household or institutional scope and material exclusions are specified. The target, frame, selected, responding and represented populations receive separate counts or estimates.

Method is classified as self-declaration, proxy declaration, indirect assignment, short direct verification or multi-item assessment. Mixed indicators identify each component and routing rule. Question text, item framework and administration mode are linked.

The dictionary records languages permitted by the concept, languages offered, language selected, translation version, script and interpreter or accommodation rules. The population unable to access an authorised version remains visible.

Sampling design, base weight, non-response adjustment, calibration, imputation, variance method, rounding and suppression rules are stated. Modelled estimates identify covariates, model version and validation.

Coverage, contact, cooperation, completion, reliability, validity, standard error, non-sampling limitations and comparability class are recorded. The quality statement includes the authorised use and any restricted use.

Table 24. Adult literacy indicator metadata dictionary
Metadata groupRequired fieldsRelease testFailure response
identityidentifier, title, version and ownerone controlled definitionresolve duplicate or conflicting record
constructdomain, evidence and interpretationclaim matches observed conditionnarrow title and use
populationage, residence, coverage and exclusionsinference population reproduciblerestrict population statement
methodquestion, respondent, task and routingmethod class explicitwithhold method-neutral rate
languageconcept, offered versions and accesslanguage condition visiblequalify or redesign
estimationweights, variance and missing-data rulesestimate reproduciblecorrect calculation
qualityresponse, uncertainty and limitationsfitness for use statedassign restricted status
historyobservation, release and change datestrend breaks traceablerestore version record

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

Every change states old and new field, reason, evidence, effective date and effect on comparison. Editorial correction that does not alter meaning is distinguished from methodological change. Historical outputs retain the version under which they were released.

Essential population, method, language, period and uncertainty fields should accompany the public value. More detailed records may be provided separately, but complexity cannot excuse omission of the indicator’s basic meaning.

B

Sampling, response and weighting protocol

The sampling plan states the target population and quantifies frame coverage from the most reliable demographic evidence. Areas, institutions or persons omitted by design are listed with estimated size where possible. Duplicates and out-of-date units receive correction procedures.

Strata, stages, primary units, households and adult selection rules are documented with probabilities. Interviewer discretion in selection is prohibited. Reserve or replacement units are used only under an approved probability design.

Every selected unit receives a final disposition: ineligible, unknown eligibility, not located, no contact, temporary absence, refusal, language barrier, accessibility barrier, partial, complete or other defined status. Attempts, dates and modes are retained.

The background interview and direct assessment have separate outcomes. Assessment disposition identifies not offered, refused, unable to access language, accommodation unavailable, started, interrupted, partially scoreable and valid completion.

Contact, cooperation, background completion and valid-assessment rates are calculated from documented denominators. Unknown eligibility receives a stated estimation rule. Rates are presented overall and for planned strata and relevant groups.

The base weight is the inverse final probability of selecting the adult. Non-response adjustment uses variables observed for both respondents and non-respondents. Calibration aligns with reliable controls under consistent population definitions.

The distribution, minimum, maximum, percentiles, design effect and effective sample size are examined. Extreme weights prompt source review before trimming. The effect of every adjustment stage on principal estimates is recorded.

Table 25. Sampling, response and weighting control
Control stageRequired evidenceDiagnosticDecision
framecoverage and exclusion maptarget–frame ratio by groupsupplement or restrict scope
selectionprobabilities at every stageweights reproduce selected countscorrect selection record
contactattempt and disposition historypatterned non-contactextend or vary fieldwork
cooperationinformed response outcomerefusal by group or interviewerimprove information and supervision
assessmentstart, break-off and valid completionliteracy-related lossanalyse bias and redesign
base weightinverse probabilityextreme or missing valuesreconcile sample selection
adjustmentclass or model variableslarge estimate movementexamine assumptions
calibrationpopulation controlsinconsistent definitionsreplace control or restrict use
variancestrata, cluster and weight methodimplausible precisioncorrect design estimation

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

Where material bias is plausible, the authority should use field observations, frame variables, shortened follow-up or linked administrative evidence under proper authority. Follow-up participation is itself selective and should not be treated as a complete correction.

The final record reconciles selected adults through valid completion and weighted population. Unexplained count loss, unidentified selection probabilities or a critical unrepresented group prevents the intended whole-population release.

C

Language adaptation and accessibility protocol

The authority defines whether the assessment concerns any written language, selected national languages or a specified language. Population evidence, policy purpose, resources and consequences guide version selection. Exclusion is documented as a limitation, not treated as absence of literacy.

Each version uses translators with command of the source and target language, literacy and measurement specialists, and reviewers familiar with written use in the target population. Conflicts and decisions are recorded.

Before translation, the team identifies the capability, textual feature, cognitive operation and difficulty driver in each task. Elements essential to comparability are separated from surface features that may be adapted.

Independent drafts are reconciled through evidence and recorded decisions. The review examines vocabulary, syntax, script, layout, document convention, number format, names and contextual familiarity. Back translation is supporting evidence only.

Cognitive interviews and field trials include adults across relevant proficiency, age and language-practice groups. Review examines interpretation of instructions, engagement with text, response process, time and distress.

Item difficulty, discrimination, missingness and differential functioning are examined across versions with adequate samples. A flagged item receives linguistic and substantive review before retention, modification or exclusion.

The protocol defines accessible consent, communication, presentation and response methods. Each adjustment is assessed against the construct. Standard-scale inclusion and descriptive alternative assessment are distinguished.

Table 26. Language adaptation and accessibility decision record
StageEvidenceAcceptance conditionUnresolved consequence
language selectionpopulation and policy mapscope serves intended populationrestrict claim or add version
construct analysistask-demand specificationessential demand identifiedtask not translated
translationindependent drafts and reconciliationmeaning and demand preservedrevise wording
contextual adaptationdocumented substitutionno irrelevant advantagetest or exclude item
cognitive testingparticipant response evidenceintended process observedredesign task
field trialtiming, missingness and performanceoperational and measurement adequacydelay main fieldwork
statistical analysiscross-version item evidenceno material unexplained differencequalify, rescale or remove
accessibilityconstruct-preservation reviewbarrier removed without changed demandseparate result or exclusion record

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

Reports state offered languages, choice or assignment rule, version sample sizes, exclusions and known equivalence limits. Language-group results are not interpreted as effects of language without an appropriate causal design.

Later orthographic, terminology or layout changes receive controlled version identifiers and testing proportionate to their likely effect. A change that alters task demand creates a comparison break unless linking evidence supports continuity.

D

Self-report and proxy module protocol

The module should identify whether it supports broad declared status, perceived difficulty, literacy practices, service access or validation. Questions not connected to an analysis or policy decision should be removed.

Self-response is preferred for personal capability and practice. Proxy response may be accepted for coverage under a defined rule and is flagged at item level. The proxy’s relationship and reason for substitution are recorded.

Questions specify activity, material, language where relevant, difficulty and reference period. Reading and writing are separated where domain-specific evidence is needed. Response categories include unknown or unable to judge for proxies.

Testing examines how adults understand literacy, difficulty categories, “simple” material, language and social consequence. It includes adults with limited schooling and varied language backgrounds.

Interviewers use exact wording and neutral clarification, protect privacy and do not infer answers from education or occupation. Assistance provided to understand the interview is recorded where it affects the question.

Attendance, highest level and completion are collected under standard education classifications. They remain explanatory variables unless a published indicator explicitly and transparently uses an assumption rule.[REF-10]

Reading and writing frequency is collected across work, household, learning and community contexts with locally relevant examples. Non-use is followed by a reason where burden permits.

Table 27. Self-report and proxy literacy module
FieldResponse designQuality flagPermitted use
respondent statusself or identified proxyproxy knowledge unknownmethod profile
reading declarationyes, no, unknown under exact wordingcombined with writingbroad declared status
writing declarationseparate statusproxy uncertaintydomain description
reading difficultygraded defined tasksunstated reference standardservice need
writing difficultygraded defined tasksassistance not recordedservice need
languagelanguages read and writtenspoken language substitutedprovision planning
practicesfrequency by text and settingopportunity confused with skillcontext and maintenance
educationstandard level and completionused as automatic literacyassociation only
access to learningavailability, participation and barrierenrolment treated as proficiencyprogramme planning

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

A probability subsample can complete direct tasks. Analysis compares self and proxy responses with assessment under the same population and period, reporting disagreement and uncertainty by relevant group.

The publication reproduces exact core wording and respondent shares. Declared status, perceived difficulty and reported practice are labelled separately. No result is described as demonstrated proficiency unless a direct task supports it.

E

Direct-assessment administration and scoring protocol

The assessor confirms selected adult, private setting, authorised language, accessible information, materials and version. Another household member does not replace the selected adult or answer assessment tasks.

The adult receives purpose, expected duration, confidentiality, voluntary status where applicable and contact information in an understandable form. The assessor explains that tasks are not a judgement of personal worth and that results are reported statistically.

Practice items establish understanding of response procedures without teaching assessed content. Difficulty with instructions prompts authorised clarification; inability to access the mode triggers the accommodation protocol.

Instructions, order, timing where relevant and permissible prompts are defined. Assessors do not translate spontaneously, explain vocabulary in assessed text or signal correctness. Interruptions and deviations are recorded.

The adult may pause or discontinue. The assessor records the point and reason without assigning unattempted items as incorrect unless the scoring framework expressly and validly requires it.

Objective responses follow controlled keys. Constructed responses use criteria, examples, double scoring and adjudication at a risk-based rate. Scorers are monitored for drift and unusual severity.

Optical or manual capture includes version and disposition controls. Critical fields receive verification. Edits retain original value, rule, authorisation and date.

Table 28. Direct-assessment administration and scoring control
StageRequired controlDeviation recordDecision consequence
identity and selectionselected eligible adult confirmedwrong or substituted adultinvalidate case
languageauthorised version and ruleunauthorised translationreview affected tasks
settingprivacy and adequate conditionsinterruption or third-party helpflag administration
instructionsstandard deliveryadditional substantive explanationvalidity review
accessapproved accommodationunavailable or altered constructseparate or exclude with count
completionstart, break and final statusreason for missing tasksapply scoring rule and bias analysis
scoringcontrolled key and agreementscorer discrepancymoderate or rescore
captureversion and response verificationmissing or impossible codecorrect from source

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

Supervisors observe sessions, recontact a protected sample for procedural verification and examine duration, completion and result patterns by assessor. Investigation separates assignment differences from performance concerns.

The adult receives appropriate closing information without an improvised individual proficiency judgement. Materials and identifiers are secured, and any welfare or access concern is referred only under the disclosed procedure.

F

Estimation, variance and sensitivity protocol

Every calculation begins with the population quantity sought: proportion, mean, percentile, distribution, difference or association. The eligible population, domain and observation period are fixed before the estimator is selected.

For a binary or category indicator, the weighted proportion is the sum of final weights for adults in the category divided by the sum of final weights for all adults in the defined denominator. Unknown and inapplicable records follow the approved rule.

The weighted mean is the sum of each final weight multiplied by the relevant score, divided by the sum of weights for adults with valid score evidence. The scored population and assessment non-completion are reported.

Variance estimation reflects stratification, clustering and weighting through an approved linearisation or replication method. The technical record identifies strata, primary units, replicate weights and any certainty selections.

Differences between groups or periods require covariance where samples or linking are related. A significance test should not replace consideration of educational magnitude and method comparability.

The interval communicates sampling uncertainty under the design and estimator. It does not include every coverage or measurement error. This boundary should accompany principal intervals.

Predefined analyses may vary missing-data treatment, weight trimming, flagged items, proficiency threshold and uncertain population eligibility. Results are compared for magnitude and decision stability.

Table 29. Estimation and sensitivity record
EstimateRequired inputsIndependent checkLimitation retained
declared ratecategory, denominator and final weightcomponent totalsproxy and missing status
proficiency meanscores or population estimates and weightsreplicate calculationassessment participation
level sharecut points, values and weightsshares total within roundingboundary uncertainty
subgroup differencecomparable estimates and covariancereversed subtraction and intervalnon-causal relationship
trend differencelinked construct and populationsbridge and method-break checkobservation interval
percentileweighted score distributionalternative computationtail precision
model-assisted local valuedirect data, covariates and modelvalidation and residualspredicted status
sensitivity rangedefensible alternative assumptionsdocumented rerunsnot a probability interval

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

Unrounded calculations are retained; public rounding is consistent across narrative and tables. Percentages reconcile within stated rounding. Counts derived from weighted estimates are labelled estimates rather than achieved sample counts.

The statistician confirms arithmetic and design, while the responsible analyst confirms that narrative claims match population, method and uncertainty. Both are required because a correct calculation can support an incorrect interpretation.

G

Comparability and trend protocol

The analyst records the exact proposed comparison and decision: magnitude, rank, distribution, trend or association. This prevents a general declaration of comparability from being extended to a stronger use.

For each value, the record contains construct, population, age, coverage, period, method, question or framework, respondent, language, assumption, sampling, response and uncertainty.

Each dimension is classified equivalent, bounded difference, material difference or unknown. Unknown is not treated as equivalent. The analyst explains whether bounded differences could reverse the proposed conclusion.

Common items, parallel methods, overlap samples or statistical linking may strengthen comparison. The design, population and uncertainty of the link are recorded. Correlation alone does not establish interchangeable scale units.

The final comparison is strong, qualified, descriptive only or not supportable. A strong classification requires no material unresolved difference for the intended use. Qualified comparison states the exact limitation.

The series lists observation date, indicator version and breaks. Changes in question, response, schooling assumption, language, mode, frame, assessment or scaling receive markers and notes.

Strong and qualified results may appear in analytical comparisons with appropriate notes. Descriptive-only values are juxtaposed without subtraction or rank. Unsupported comparisons should not be displayed in a way that invites the prohibited inference.

Table 30. Comparability and trend authorisation
Review outcomeEvidence conditionPermitted presentationProhibited presentation
strongconstruct, population and method alignedmagnitude, distribution and trend with uncertaintycausal claim without design
qualifiedbounded known differencesbroad contrast with adjacent qualificationprecise rank or small change claim
descriptive onlyrelated but materially different measuresseparate panels and contextual accountsubtraction or common scale
not supportablefundamental or unknown differencemethod description onlycomparative conclusion
linked breakoverlap evidence estimates discontinuitymarked series and bridge resultsilent splicing
unlinked breakmethod changed without bridgeseparate seriesprogress or decline across break

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

The technical comparison decision should precede policy interpretation. Disagreement is recorded and resolved by the competent statistical authority under published principles, not by selecting a convenient source.

New methods may clarify earlier differences but do not alter the contemporaneous status of historical observations. Reanalysis should be dated and should preserve the original publication and method.

H

Publication, confidentiality and correction protocol

The release inventory lists report, tables, metadata, methodological documentation, accessible summary and authorised data products. Each carries document number, version, date and observation cutoff.

Every principal claim is traced to a table or cited source. The reviewer tests population, method, time, modality and causal language. Qualifications appear beside the proposition they limit.

Headers state measure and population; notes state method, language, observation year and uncertainty. Rows and columns reconcile, suppressed cells remain protected and totals are not misleading after suppression.

Direct identifiers are removed from public outputs. Small cells, rare languages, geography and combined characteristics are assessed for re-identification and community harm. Restricted analytical detail remains available only to authorised users.

Public summaries use clear language, defined terms, readable tables and suitable oral or translated forms. They preserve denominators, uncertainty and method distinctions. Adults represented by the findings should be able to understand the public meaning.

Principal results are released under an announced procedure providing equal public access. Pre-release access, where authorised for operational preparation, should be limited, recorded and unable to alter results.

A reported error is logged, assessed for materiality, corrected across every affected format and announced with date and consequence. The correction authority acts independently of whether the revised result is favourable.

Table 31. Publication, confidentiality and correction assurance
AssuranceEvidenceRelease conditionCorrection if failed
document identityversion, date and observation periodone approved releasewithdraw conflicting edition
claim alignmentclaim-to-table and source tracewording within evidencerevise claim
numerical parityindependent calculations and format comparisonall values agreecorrect all representations
method visibilitypopulation, method and language notesessential meaning beside valuerestore metadata
uncertaintydesign and non-sampling accountprecision not overstatedadd interval and limitation
confidentialitydisclosure-risk reviewrisk controlledsuppress, combine or restrict
accessibilityclear and appropriate public formsqualifications preservedrevise communication
equal accessrelease logauthorised schedule followeddisclose and correct process
correctionissue, decision and change noticematerial error traceableissue revised controlled version

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

The archive retains source definitions, instruments, translations, field manuals, scoring, weights, calculations, review and released versions under protection. Preservation enables later interpretation without disclosing respondent information.

The competent authority signs the release only when statistical, interpretive and confidentiality reviews pass. Policy bodies may respond to the findings after release, but their response should remain distinct from the statistical account.

I

Fieldwork monitoring and incident protocol

Field supervisors receive selected, contacted, completed and outstanding counts by stratum, interviewer and authorised language. Reports distinguish background completion from valid assessment. Rates are used to identify operational problems, not to impose coercive quotas.

Duration, refusal, proxy use, missing items, break-off and score distributions are examined against assignment. Outliers trigger review of workload, geography, language and practice before conduct is inferred.

Incidents include wrong respondent, unauthorised substitution, language mismatch, loss of materials, confidentiality exposure, item exposure, unsafe setting, substantive third-party assistance, system failure and participant distress. Severity and immediate protection are recorded.

The supervisor protects the adult and information, stops affected work where necessary, secures materials and reports through the defined route. A serious confidentiality or safety incident is not deferred to routine quality review.

The technical authority determines whether an incident affects the whole case, selected tasks or only operational metadata. The original responses and decision record are retained. Field personnel do not delete a case to improve completion figures.

Action may include retraining, observation, reassignment, repeat contact with consent, replacement materials, added language capacity or suspension. Re-interview is authorised only where it does not create undue pressure or compromised item exposure.

Table 32. Fieldwork incident and response record
IncidentImmediate controlTechnical decisionPopulation implication
wrong adultstop session and secure recordinvalidate or restart authorised selectionselection integrity
language mismatchdo not score as low performancereschedule or record inaccessiblelanguage undercoverage
third-party helprecord tasks and circumstancesflag or invalidate affected evidencepossible score bias
item exposuresecure material and identify extentitem or area analysiscomparability and security
confidentiality lossprotect and notify authorityincident response and access reviewtrust and participation risk
break-off clusterexamine burden and administrationrevise field supportselective completion
system failurepreserve source and versionrecover under controlled rulemissingness and delay

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

At fieldwork close, every selected unit has a disposition and every incident has a decision or continuing owner. The incident profile is included in the quality report and informs the next design.

J

Proficiency-level description protocol

Level descriptions translate score ranges into the kinds of tasks that adults at different positions on the scale are likely to perform. They support interpretation and should not become labels of intelligence, worth or fixed capacity.

Descriptions use tasks located within the scale, their content and the probability criterion adopted by the assessment. The procedure and criterion are documented. A few released examples should illustrate rather than define the entire level.

Specialists examine text form, operation, inference, competing information, length and context across tasks. The description identifies increasing demand without claiming that every task within a range shares one feature.

Adults near a cut point have similar estimated performance despite different level labels. Reports should state this and provide standard errors for level shares. “At Level 1” means classified within the approved range under the model.

The lowest category should describe the limited evidence available rather than state that adults have no literacy. Assessment may contain too few very easy items to characterise performance precisely below the scale.

The top category is also limited by task coverage. It should not be described as mastery of all literacy demands. Sparse samples at the upper tail require suitable precision notes.

Table 33. Proficiency-level description assurance
ElementRequired basisWording controlMisuse prevented
score rangeapproved scale cut pointsnumerical range statedhidden threshold
task demandmultiple located itemslikely performance under criterionsingle item defines level
boundarymeasurement uncertaintyadjacent adults may be similarcategorical discontinuity
lowest categoryavailable easy-task evidencebelow assessed thresholdno literacy identity
highest categoryupper task coverageperformance on demanding assessed tasksuniversal mastery
population shareweights and varianceestimate with standard errorexact population count
policy usedefined task relevanceselected-level rationaleuniversal deficit threshold

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

Level descriptions change only through controlled review of scale and task evidence. Editorial simplification should not broaden the claim. A changed cut point creates a new indicator version and comparison record.

K

Programme-learning measurement protocol

Programme assessment begins with the learning objective and intended participants. It should identify reading, writing, numeracy, language or applied practice expected to change and the instructional period.

All entrants are counted with eligibility, start date and baseline status. Adults unable to complete the standard baseline remain in the participation record and receive an accessible route where valid.

Tasks should reflect intended learning without reproducing only practised examples. A national population assessment may provide a reference framework but is not automatically sensitive to programme progress.

Completion, withdrawal, transfer, absence and assessment non-response are separated. Matched change is reported with the number and characteristics lost to follow-up.

Where policy requires an effect estimate, the comparison should represent what would plausibly have occurred without the programme. Assignment, matching or phased implementation may assist, subject to ethics and design.

Observed change may reflect instruction, practice, maturation, selection and measurement. The conclusion should distinguish change among completers, change among entrants after adjustment and causal contribution.

Table 34. Programme literacy measurement record
Evidence stageRequired count or measureInterpretationLimitation
eligible populationdefined intended adultspotential reachdenominator may be estimated
entrantsenrolled and startedactual reachself-selection
valid baselinecomparable entry evidencestarting distributionbaseline non-completion
exposureattendance, duration and instructional contenttreatment receivedquantity not quality alone
valid follow-upcomparable later evidenceobserved follow-up distributionselective retention
matched changesame adults and scalechange among matched adultsattrition
comparison resultaligned non-participant or phased groupevidence on contributionresidual confounding
sustained resultlater comparable evidencemaintenancelater loss to follow-up

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

Providers should publish reach, retention, assessment coverage and learning findings without presenting participant results as community prevalence. Negative and mixed findings remain part of programme accountability.

L

National measurement cycle schedule

The responsible authorities approve policy questions, governance, target population, component roles, languages, budget and release intention. The observation date and principal comparisons are fixed.

The statistical and assessment bodies prepare frames, instruments, translations, sampling, protection and analysis. Cognitive testing, pilot and field trial have distinct decision gates.

Readiness confirms valid instruments, language versions, trained field staff, verified selection, functioning capture, incident response, scoring and secure logistics. Critical failure delays or stages fieldwork.

Fieldwork follows varied contact schedules and continuous quality monitoring. Design changes during collection require central authority and a record of affected cases and comparability.

The sequence includes disposition reconciliation, capture verification, scoring, edit, weighting, variance, item and language review, disclosure and replication. Provisional results remain restricted until the relevant gate.

Technical review, public tables, accessible communication and policy briefing are prepared from the same approved estimates. Observation and release dates are distinguished.

Education and other responsible bodies issue action responses. Programme design, language provision and resource decisions refer to the relevant indicator and population. Statistical findings are not rewritten as commitments.

Table 35. National adult literacy measurement cycle
StageControlled productDecision gatePrincipal risk
initiationpolicy and measurement briefpurpose and governance approvedinstrument without decision use
constructframework and indicator mapinterpretations approvedbroad claim from narrow evidence
languagetested versions and coverageaccess adequate for scopeexcluded communities
sampleframe and selection planpopulation inference supportedunrepresented target group
field trialoperational and measurement evidencemain collection readyunresolved task or burden
collectiondispositions, responses and incidentscoverage and participation acceptableselective completion
processingscores, weights and replicated estimatesnumerical quality passedhidden processing error
interpretationcomparability and limitation recordclaims supportablefalse rank or trend
releasecontrolled statistical publicationconfidentiality and parity passedinconsistent public values
policy responseactions, resources and evaluation planresponsible authority actsmeasurement without provision

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

Stable household questions and programme indicators may continue between direct assessments. They should not be used to estimate unobserved proficiency annually. Their role is context, access and early warning.

The next cycle begins with an evaluation of participation, language, burden, method use and policy consequence. Change is introduced through testing and linking so that improvement in relevance does not silently destroy comparability.

M

Indicator reconciliation and source-status protocol

Where several official literacy values exist, reconciliation establishes why they differ and which use each can support. It does not force unlike methods into one preferred number. The process should begin before values enter a common table or public target.

For every value, the record identifies producing body, publication, observation date, release date, population, question or assessment, respondent, language, assumption, weighting, uncertainty and revision status. Secondary reproductions are traced to the primary source.

A value is classified as observed declaration, observed direct task, indirect classification, mixed-method estimate, modelled estimate or projection. Preliminary, final, corrected and superseded status are also recorded.

Age ranges are recalculated to a common range where source data permit. Territory, household status and exclusions are compared. If harmonisation requires removing a material population, both original and harmonised scopes remain visible.

Exact questions, routing and tasks are compared. Shared words such as read, write, simple and understand are not treated as equivalent without operational evidence. Proxy and schooling assumptions are quantified.

Published numerators, denominators, weights and rounding are reproduced. Where only a rounded rate exists, an exact population count should not be inferred. Revisions are tied to their source notice.

Table 36. Indicator reconciliation and source-status decision
Reconciliation fieldEvidenceCompatible outcomeIncompatible outcome
source statusprimary release and revision historysame controlled versionpreliminary mixed with final
observation datefieldwork periodaligned or policy-relevant intervalunknown or distant year labelled current
populationage, territory and residencecommon recalculated scopeunresolved material difference
constructdomain and interpretationsame defined capabilitydeclaration equated with proficiency
respondentself, proxy or performed taskcomparable response processproxy share unknown and material
languagepermitted and administered formsequivalent accessone-language result treated as any-language
estimationweight, model and uncertaintyreproducible common statisticmodelled value called observation
trendstable version or bridgelinked seriessilent method break

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

If two values remain incompatible, the public reconciliation note states each source and use. One may support local declared need and another national proficiency distribution. Both can be valid within their boundaries.

An error correction changes the affected historical record through a dated notice. A new method does not correct the earlier value; it creates a new series or a linked series where bridge evidence exists.

Values submitted for international monitoring should include the full method class and observation year. If the requested indicator does not match the national source, the difference should be declared rather than concealed through relabelling.

A national target identifies the indicator version against which progress will be judged. If measurement improves, the target may require formal rebasing, with old and new evidence shown. Performance is not recalculated by simply applying the old numerical threshold to a different method.

The statistical authority approves the source-status map and permitted comparisons. Policy bodies acknowledge the evidence basis of targets and actions. The completed record allows later users to understand why more than one official literacy value may properly exist.

An aggregate should use observed or approved estimated values, a population weight consistent with the indicator and a rule for missing countries. The covered population share and observation-year range should accompany the result. A regional value should not be produced when missing populations are large enough to make the result materially unrepresentative under the approved rule.

Mixed-method aggregation requires particular restraint. Weighting self-declaration and direct-assessment rates into one regional number does not make their constructs equivalent. Where the monitoring purpose requires both, results should be grouped by method or presented as a qualified series with the limitation prominent.

Female, male and total values should reconcile under the population weights and categories used. A total derived from one source should not be combined with sex-specific values from another observation without disclosure. Unknown sex records and non-binary national categories, where collected, require a stated treatment consistent with the source and protection requirements.

Gender disparity measures should preserve direction, denominator and method. A difference in percentage points is not the same as a ratio, and neither alone explains the institutional barriers producing inequality.

The reconciliation is closed only when each value has an authorised status, comparisons have classifications and unresolved conflicts are disclosed. Closure does not require numerical agreement. It requires that no user be invited to treat unlike evidence as one observation.

The final map should be reviewed whenever a correction, new survey or method change enters the series. This continuing custody is essential to the integrity of long-term adult literacy monitoring.

Where publication space is limited, the minimum statement should still identify the observation year, covered population, evidence method, respondent rule, language scope and whether the value is observed, assumed or modelled. It should state any material break from the preceding value and provide the location of the full source record.

A short statement must not compress incompatible evidence into an apparently continuous rate. If comparison is descriptive only, that classification should appear with the figures. If comparison is not supportable, the values should be separated and no difference calculated.

The responsible statistical officer should confirm that the abbreviated statement preserves the conclusion of the full reconciliation. Brevity is acceptable only when it does not change the public meaning.

The statement should also retain the principal uncertainty measure and the share of the intended population not represented. Neither element may be omitted merely because the result is presented in a summary table. A comparison without these controls should be returned for qualification before publication or policy use.

N

Public table specification and interpretive notes

Every table has a number, substantive title, indicator version, observation period and covered population. The title states whether results are declared, assessed, modelled or programme-based.

Column headers identify count, weighted estimate, percentage, score, standard error or interval. Units are not left to surrounding prose. Age, sex, language and geography categories retain their source definitions.

The first note states evidence method, respondent rule, language scope and material assumption. A mixed-method value identifies its components. Direct-assessment tables state the domains and target population.

Sample estimates carry standard errors or confidence intervals and a marker for estimates failing the approved precision rule. Unweighted sample count is provided where it assists assessment of stability and does not create disclosure risk.

Excluded territories, institutions, ages and language groups are stated. Assessment completion and any residual non-response limitation appear with proficiency tables.

Trend and cross-national tables state comparability class. Method breaks are marked at the relevant value. Descriptive-only values are not joined by lines, differences or ranking.

Table 38. Public literacy table specification
Table elementMandatory contentInterpretation protected
number and titlesubject, method and populationtable cannot be detached from meaning
observation fieldreference date or fieldwork periodrelease year not mistaken for data year
value headerunit and indicator versionrate, score and count not conflated
population noteage, residence and exclusionsscope not overstated
respondent noteself, proxy or assessed adultdeclaration not called direct evidence
language notepermitted and administered languageslanguage-limited result not universalised
uncertaintystandard error and non-sampling limitationpoint estimate not treated as exact
comparabilitystrong, qualified, descriptive or unsupportedrank and trend constrained
sourceprimary producing body and publicationsecondary value traceable

Source and methodological notes are stated immediately below the table in the authoritative Markdown text.

The text identifies the principal distribution, material disparity and uncertainty without repeating every cell. It distinguishes finding from explanation. Any causal interpretation cites evidence beyond the table.

Layout, font, contrast and reading order should support access. A concise explanation may use ordinary language and examples while preserving denominator, method and uncertainty. Symbols and abbreviations are defined.

A correction updates the table, structured data and every narrative claim using it. The notice identifies the cells changed and whether the substantive conclusion is altered. The superseded table remains archived under control.

O

Summary release statement

The summary release should state in one place the observation period, population, method, languages, achieved participation, principal estimate, sampling uncertainty and material non-sampling limitation. It should identify whether comparison with the preceding value is strong, qualified, descriptive only or unsupported.

The statement should describe the educational meaning of the result without calling adults below a threshold incapable or treating a group difference as inherent. It should identify the policy question and the responsible public response.

The primary statistical source, methodological documentation and accessible public explanation should be identified. A corrected release should point to the correction notice and retain the original observation date.

The statistical authority approves the numerical and methodological account. The responsible education authority may provide a separate policy response. Keeping the two functions distinct protects both impartial evidence and accountable action.

The summary statement should retain any limitation capable of changing the principal interpretation, even where the full technical account is available elsewhere. Coverage, language and method breaks are substantive conditions and should not be reduced to an unreferenced footnote.

Approval confirms that the shortened account remains faithful to the complete evidence record.

It also confirms that every cited value can be traced to an authorised source and that no unsupported comparative or causal conclusion has been introduced through abbreviation.

The approved statement remains part of the permanent public statistical record.

References

  1. REF-01

    World Education Forum. The Dakar Framework for Action: Education for All — Meeting Our Collective Commitments. 2000.

    The global commitment to halve adult illiteracy, improve equitable access to continuing education and monitor measurable progress.

    https://unesdoc.unesco.org/ark:/48223/pf0000121147
  2. REF-02

    UNESCO. Education for All Global Monitoring Report 2006: Literacy for Life. 2005.

    The principal contemporary synthesis of literacy concepts, global estimates, measurement limitations, policy conditions and adult literacy priorities.

    https://unesdoc.unesco.org/ark:/48223/pf0000141639
  3. REF-03

    UNESCO Institute for Statistics. Literacy Assessment and Monitoring Programme (LAMP): Information Brochure. 2006.

    The contemporary development of direct literacy assessment, background information, proficiency distributions and improved national measurement capacity.

    https://uis.unesco.org/sites/default/files/documents/literacy-assessment-and-monitoring-programme-lamp-information-brochure-en.pdf
  4. REF-04

    UNESCO. The Plurality of Literacy and Its Implications for Policies and Programmes. 2004.

    The treatment of literacy as plural, socially situated and related to language, purpose, culture and practice.

    https://unesdoc.unesco.org/ark:/48223/pf0000136246
  5. REF-05

    UNESCO Institute for Statistics. Aspects of Literacy Assessment: Topics and Issues from the UNESCO Expert Meeting, 10–12 June 2003. 2005.

    Technical and conceptual issues in literacy assessment, including construct, language, context, administration and interpretation.

    https://uis.unesco.org/sites/default/files/documents/aspects-of-literacy-assessment-topics-and-issues-from-the-unesco-expert-meeting-2005-en_0.pdf
  6. REF-06

    OECD and Statistics Canada. Literacy in the Information Age: Final Report of the International Adult Literacy Survey. 2000.

    Comparative direct assessment of prose, document and quantitative literacy and the distribution of adult proficiency.

    https://www.oecd.org/en/publications/literacy-in-the-information-age_9789264181762-en.html
  7. REF-07

    OECD and Statistics Canada. Learning a Living: First Results of the Adult Literacy and Life Skills Survey. 2005.

    The contemporary extension of adult literacy assessment to prose, document, numeracy and problem-solving domains, with background and participation analysis.

    https://www.oecd.org/en/publications/learning-a-living_9789264010390-en.html
  8. REF-08

    United Nations Statistics Division. Principles and Recommendations for Population and Housing Censuses, Revision 1. 1998.

    The contemporaneous census framework for population characteristics, literacy enquiry, concepts, classifications and census operations.

    https://unstats.un.org/unsd/demographic-social/standards-and-methods/
  9. REF-09

    United Nations Statistical Commission. Fundamental Principles of Official Statistics. 1994.

    Professional independence, scientific methods, transparency, confidentiality and public access in official statistics.

    https://unstats.un.org/unsd/dnss/gp/fundprinciples.aspx
  10. REF-10

    UNESCO. International Standard Classification of Education: ISCED 1997. 1997.

    Comparability controls for educational attainment and the separation of educational level from directly observed literacy proficiency.

    https://uis.unesco.org/sites/default/files/documents/international-standard-classification-of-education-1997-en_0.pdf
  11. REF-11

    World Bank. Designing Household Survey Questionnaires for Developing Countries: Lessons from 15 Years of the Living Standards Measurement Study. 2000.

    Questionnaire design, respondent selection, field operations, non-response, household reporting and data-quality controls.

    https://documents.worldbank.org/en/publication/documents-reports/documentdetail/452741468778781879/
  12. REF-12

    United Nations Department of Economic and Social Affairs, Statistics Division. Household Sample Surveys in Developing and Transition Countries. 2005.

    Sampling, coverage, field implementation, non-response, weighting, error and dissemination in household surveys.

    https://unstats.un.org/unsd/hhsurveys/pdf/Household_surveys.pdf
  13. REF-13

    UNESCO Institute for Statistics. Guide to the Analysis and Use of Household Survey and Census Education Data. 2004.

    Analysis of education variables from household and census sources, including definitions, denominators, disaggregation and comparability.

    https://uis.unesco.org/sites/default/files/documents/guide-to-the-analysis-and-use-of-household-survey-and-census-education-data-en_0.pdf
  14. REF-14

    United Nations. The Millennium Development Goals Report 2006. 2006.

    The contemporary global development-monitoring context, including education and gender disparities relevant to literacy participation and reporting.

    https://unstats.un.org/unsd/mdg/resources/static/products/progress2006/mdgreport2006.pdf
  15. REF-15

    United Nations General Assembly. International Plan of Action for the United Nations Literacy Decade. 2002.

    The international policy context for literacy as a continuum, strengthened monitoring, partnerships and attention to diverse learning needs.

    https://undocs.org/A/57/218