
Publication record
This is the controlled English edition. Evidence and institutional status are stated as at the evidence cut-off date.
Executive summary
Adult literacy statistics are indispensable to education policy and unusually difficult to interpret. A national rate may be drawn from a census question answered by the individual, a statement made by another household member, an assumption based on schooling, a short reading task or a broader assessment of performance across several kinds of text. These methods do not necessarily measure the same construct, cover the same population or produce comparable values.
The distinction is material. A self-reported response can describe how a person identifies or functions in a familiar setting, but it is affected by question wording, social expectations, language, privacy, proxy response and the threshold understood by interviewer and respondent. A direct assessment observes performance under specified conditions, but it is affected by task selection, translation, administration, sampling, non-response and the relationship between the assessment setting and adults’ everyday literacy practices.
Literacy is not adequately represented as a permanent binary possession. Contemporary international analysis describes a continuum of capabilities used for different purposes and in different social, linguistic and economic settings. Direct international assessments have demonstrated substantial distributions of prose, document, quantitative and numeracy proficiency within populations and within levels of educational attainment.
This does not render census and household-survey indicators without value. Censuses can provide broad population coverage and small-area information that a specialised assessment cannot ordinarily match. Household surveys can connect literacy responses to poverty, work, health, language, migration and participation. Their value depends on a clear account of what was asked, who answered, which languages were available, how schooling was treated and who was omitted.
The report establishes four classes of literacy evidence: self-declaration by the person concerned; proxy declaration by another household member; indirect classification through educational participation or attainment; and direct performance assessment. It also distinguishes a short verification task from a multi-domain assessment. No class is declared universally superior. Each answers a narrower set of questions and carries characteristic error.
Comparability requires more than applying the same label. Concept, population, age range, language, reference period, question, response categories, mode, respondent, task demands, scoring, sampling, weighting and non-response must be examined together. A time series can break when one of these changes. A cross-national table can contain accurately reported national values that remain unsuitable for direct ranking.
Participation is both an object of measurement and a condition of measurement quality. Adults with weaker literacy, unfamiliarity with the survey language, disability, insecure residence, long working hours, institutional residence or fear of official contact may be less likely to be reached or to complete an assessment. If their exclusion is not measured, the published distribution may overstate population proficiency.
The report proposes an indicator-use protocol rather than one universal literacy rate. Every published value should carry a method class, population, age range, languages, respondent rule, reference date, coverage statement and explicit interpretation boundary. Comparisons should be classified as strong, qualified, descriptive only or not supportable.
Where countries require both extensive coverage and stronger evidence of skill, a layered design is appropriate: a concise census or household module for breadth; direct assessment in a probability sample for proficiency; and a background questionnaire for language, education, literacy practices and participation. The contemporary development of the Literacy Assessment and Monitoring Programme reflects the need for more relevant and reliable evidence while building national statistical capacity.
The public interest is not served by replacing one uncertain estimate with an opaque assessment score. Responsible measurement must preserve the definition, sampling, uncertainty, accessibility, confidentiality and policy meaning of the result. Literacy statistics should identify barriers and inform provision; they should not stigmatise adults, languages or communities.
Key findings
- “Literacy rate” is not a method-neutral term. Its meaning depends on the concept, question or task, population, respondent, language and threshold used.
- Self-declaration, proxy declaration, schooling-based inference, short verification and multi-domain direct assessment measure overlapping but non-identical conditions.
- A binary response may support broad monitoring but cannot describe the distribution of literacy practices and proficiency within the population.
- Educational attainment is related to literacy but is not an adequate substitute for observed proficiency. Adults with the same formal level may have substantially different skills.
- Direct assessment reduces reliance on perception and proxy reporting but introduces its own construct, task, translation, sampling, administration and participation requirements.
- Proxy reporting should be separately identified. A household member may not know another adult’s reading and writing practices, particularly where ability is concealed or used outside the home.
- Question wording can change the threshold. “Can read and write,” “can read a simple statement,” and “reads without difficulty” do not create one interchangeable indicator.
- Language is part of measurement validity. Failure in an unavailable or unfamiliar language cannot be interpreted automatically as absence of literacy in all languages.
- Survey participation is socially distributed. Coverage and non-response analysis should examine adults least likely to be reached or assessed, not only the achieved sample total.
- Cross-national comparison requires method metadata beside the value. A common title or denominator is insufficient.
- Trend analysis should identify breaks caused by changes in question, respondent rule, language, frame, age range, assumption based on schooling or direct-assessment method.
- National aggregates should be accompanied by sex, age, location and other policy-relevant distributions where sample and confidentiality permit. The choice should respond to participation and provision questions.
- Thresholds and proficiency levels are reporting devices. They should not conceal score distributions, uncertainty or variation near a cut point.
- Census breadth, household-survey context and specialised assessment depth are complementary when linked through transparent definitions and appropriate samples.
- Published literacy evidence should support educational provision and equal participation, not rank communities through measures that are not comparable.
Scope and method
The report examines adult literacy measurement as at 7 October 2006. Its principal population is persons aged 15 years and over, consistent with common international adult-literacy reporting, while recognising that national systems may use other age ranges. It addresses censuses, general household surveys, specialised literacy surveys and international comparative assessments.
The evidence base comprises the Dakar Framework and international Literacy Decade action; the 2006 global monitoring report on literacy; UNESCO analysis of literacy plurality and assessment; the UIS Literacy Assessment and Monitoring Programme; the International Adult Literacy Survey and Adult Literacy and Life Skills Survey; United Nations census and household-survey guidance; principles of official statistics; international education classification; and contemporary global monitoring.
The report compares methods, not countries. Numerical cases are constructed to demonstrate denominator, non-response, standard error, classification and trend decisions. They do not estimate any identified national population.
“Self-report” means an adult’s answer about that adult’s own literacy. “Proxy report” means an answer supplied by another person. “Indirect classification” means assignment from another characteristic, such as completed schooling. “Direct assessment” means performance on one or more specified tasks. “Comparability” means sufficient equivalence of concept, population and measurement to support the stated comparison; it does not require complete identity in every operational detail.
Part I
What an Adult Literacy Indicator Claims
Indicator purpose
feasibility and require, different, measures in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 1. Indicator purpose to connect evidence with delivery, unequal effect and correction. Broad population monitoring, local programme planning, identification of service barriers and evaluation of a learning programme may require different measures.
The method should be selected after the question. A readily available rate should not determine the purpose retrospectively.
Literacy as capability and practice
capacity and language, setting, opportunity in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 2. Literacy as capability and practice to connect evidence with delivery, unequal effect and correction. It is shaped by purpose, language, text, setting and opportunity.[REF-02] [REF-04]
A measure that observes one task should state the domain it represents. It should not claim every social and communicative dimension of literacy.
The binary convention
equity and levels, domains, categories in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 3. The binary convention to connect evidence with delivery, unequal effect and correction. The convention can provide a concise count but compresses different levels, domains and uses into two categories.
Movement near the operational threshold can alter the rate without representing a sharp division in adults’ capabilities.
A continuum of proficiency
timing and analysis, distribution, demand in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 4. A continuum of proficiency to connect evidence with delivery, unequal effect and correction. This allows analysis of distribution and task demand.[REF-06] [REF-07]
coverage and treated, substantively, discontinuous in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 4. A continuum of proficiency to connect evidence with delivery, unequal effect and correction. Adults within a level are not identical, and performance near adjacent cut points should not be treated as substantively discontinuous.
Reading and writing
uncertainty and observed, inference, justified in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 5. Reading and writing to connect evidence with delivery, unequal effect and correction. A reading task cannot establish writing performance unless writing is separately observed or the inference is justified.
Published metadata should identify whether reading, writing or both entered the classification.
Numeracy
comparability and distinct, assessed, domain in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 6. Numeracy to connect evidence with delivery, unequal effect and correction. The Adult Literacy and Life Skills Survey treats numeracy as a distinct assessed domain.[REF-07]
A combined label should not conceal different constructs or thresholds. Policy responses may differ where the principal difficulty concerns text, quantity or their interaction.
Documents and prose
authority and locating, integrating, information in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 7. Documents and prose to connect evidence with delivery, unequal effect and correction. Continuous text, forms, schedules, maps and tables may place different demands on locating, integrating and using information.[REF-06]
A short sentence-reading item provides little evidence about the use of complex documents, even where it is useful for a minimum verification purpose.
Functional interpretation
coverage and educational, cultural, settings in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 8. Functional interpretation to connect evidence with delivery, unequal effect and correction. The term requires specification because functions and demands differ across work, family, civic, educational and cultural settings.
It should not be used as an undefined higher threshold added to a basic rate.
Skills and opportunities for use
Observed proficiency reflects learning and opportunities to use skills. Adults may gain, maintain or lose fluency as demands and practices change.[REF-06] [REF-07]
Policy should therefore examine the environments that support literacy, not treat the measured result solely as an individual attribute.
Population specification
distribution and population, insufficient, metadata in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 11. Population specification to connect evidence with delivery, unequal effect and correction. “Adult population” is insufficient metadata.
timing and educational, opportunity, cohort in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 11. Population specification to connect evidence with delivery, unequal effect and correction. Values from different upper-age rules can diverge where proficiency and educational opportunity vary by cohort.
Reference period
feasibility and migration, household, change in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 12. Reference period to connect evidence with delivery, unequal effect and correction. Fieldwork extending across months may require a defined reference date and treatment of birthdays, migration and household change.
Modelled estimates and projections should be identified separately from observed results.
Unit of observation
capacity and access, replace, record in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 13. Unit of observation to connect evidence with delivery, unequal effect and correction. Household-level conditions may explain practice and access, but they do not replace the adult record.
Weights and denominators should correspond to persons, not interviewed households, when the reported result is an adult literacy rate.
Table 1: claim specification
| Claim field | Required statement | Question protected | Misinterpretation prevented |
|---|---|---|---|
| purpose | monitoring, planning, diagnosis or evaluation | why the value is produced | available statistic determines policy question |
| construct | reading, writing, numeracy, domain and use | what capability is represented | broad literacy inferred from one task |
| classification | binary, ordered category or scale | how performance becomes a result | threshold treated as natural divide |
| population | age, residence and coverage | to whom the result applies | unlike populations compared |
| language and script | permitted and administered forms | in what language performance was observed | language mismatch treated as no literacy |
| respondent | self, proxy or assessed person | who supplied the evidence | proxy answer treated as self-report |
| method | question, indirect rule or direct task | how evidence was obtained | method-neutral “rate” assumed |
| period | fieldwork and reference date | when the condition was observed | publication year treated as measurement year |
| uncertainty | sampling and material non-sampling limits | how precisely result is known | point estimate treated as exact |
| use boundary | claims not supported | where interpretation must stop | indicator used for unsupported ranking |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Public meaning
An indicator becomes public evidence when users can understand its scope and limitations. Technical complexity does not justify publication of an unexplained headline rate.
The public meaning should be accurate enough to protect policy and respectful enough to avoid presenting adults below a threshold as incapable of learning or participation.
Part II
Self-Reported and Proxy-Reported Literacy
Nature of self-report
remedy and questions, designed, purposes in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 16. Nature of self-report to connect evidence with delivery, unequal effect and correction. It can be rapid, inexpensive and suitable for large population instruments. It can also capture perceived difficulty or actual practice when questions are designed for those purposes.
It does not directly observe performance. Its validity depends on the relationship between the reported judgement and the construct used in policy.
Question wording
continuity and difficulty, provides, differentiation in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 17. Question wording to connect evidence with delivery, unequal effect and correction. ” leaves the language, material, understanding, ease and threshold to interpretation. A question about reading a simple message is narrower; a graded question about difficulty provides more differentiation.
Wording changes can alter responses even when underlying ability is unchanged. Exact question text should accompany comparisons.
Combined questions
distribution and understanding, expected, category in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 18. Combined questions to connect evidence with delivery, unequal effect and correction. Respondents may answer according to the stronger skill, the weaker skill or their understanding of the expected category.
Separate questions improve diagnostic value, though they do not remove self-evaluation error.
Response categories
feasibility and anchors, consistent, administration in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 19. Response categories to connect evidence with delivery, unequal effect and correction. Ordered categories such as easily, with difficulty and not at all can retain useful variation but require clear anchors and consistent administration.
An additional “unknown” or “not stated” category should not be combined with non-literacy. Its distribution may reveal respondent or fieldwork problems.
Reference standard
equity and standard, formal, writing in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 21. Reference standard to connect evidence with delivery, unequal effect and correction. A person who manages familiar work documents may answer yes; another with similar performance may answer no because the understood standard is formal writing.
Comparability is weakened when respondents supply both the evidence and an unstated threshold.
Literacy practices as self-report
timing and practice, opportunity, maintenance in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 22. Literacy practices as self-report to connect evidence with delivery, unequal effect and correction.[REF-06] [REF-07]
Practice questions should state the period and purpose and allow for locally relevant materials. Frequency is not a substitute for quality or understanding.
Proxy reporting
uncertainty and uncertainty, knowledge, judgement in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 23. Proxy reporting to connect evidence with delivery, unequal effect and correction. Proxy reporting reduces field cost and permits completion when members are absent. It also introduces uncertainty about knowledge and judgement.
The dataset should identify whether the answer was self or proxy wherever operationally possible.
Knowledge within the household
comparability and stigma, household, relations in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 24. Knowledge within the household to connect evidence with delivery, unequal effect and correction. Ability may also be concealed because of stigma or household relations.
Proxy accuracy should not be presumed from relationship alone. Survey evaluation can compare self and proxy answers in a subsample.
Household hierarchy
authority and systematically, favour, particular in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 25. Household hierarchy to connect evidence with delivery, unequal effect and correction. Interview scheduling and respondent rules can systematically favour particular ages or sexes.
Field protocols should identify the most knowledgeable eligible respondent and record who answered each literacy item.
Inference from schooling
coverage and demonstrate, current, performance in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 26. Inference from schooling to connect evidence with delivery, unequal effect and correction. Educational attainment is associated with literacy but does not demonstrate current performance.[REF-06]
feasibility and convert, literacy, observations in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 26. Inference from schooling to connect evidence with delivery, unequal effect and correction. International classification helps describe education levels but does not convert them into literacy observations.[REF-10]
Assumed literacy and assumed illiteracy
remedy and illiteracy, misclassify, adults in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 27. Assumed literacy and assumed illiteracy to connect evidence with delivery, unequal effect and correction. Both can misclassify adults.
Every assumed classification should be separately identified in metadata and sensitivity analysis.
Non-response
continuity and estimates, biased, upward in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 28. Non-response to connect evidence with delivery, unequal effect and correction. If adults with weaker literacy are less likely to respond personally, complete-case estimates may be biased upward.
Substitution by proxy may improve coverage while changing the measurement method. The trade-off should be documented rather than hidden in an overall response rate.
Interviewer effects
distribution and procedures, treatment, uncertainty in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 29. Interviewer effects to connect evidence with delivery, unequal effect and correction. Training should address neutrality, exact wording, language procedures and treatment of uncertainty.[REF-11] [REF-12]
Unusually high or low literacy rates by interviewer may indicate assignment differences or field error and require controlled review.
Mode effects
feasibility and without, changing, meaning in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 30. Mode effects to connect evidence with delivery, unequal effect and correction. A self-completed literacy question can exclude the very respondent whose evidence is sought unless accessible assistance is provided without changing meaning.
Mode changes in a time series require evaluation and, where material, a break in comparability.
Language of interview
capacity and rather, literal, wording in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 31. Language of interview to connect evidence with delivery, unequal effect and correction. Translation should preserve the construct and response thresholds rather than literal wording alone.
The language actually used should be recorded, together with interpreter use and any cases in which an eligible adult could not be interviewed in an adequate language.
Table 2: self- and proxy-report risk register
| Risk | Mechanism | Evidence needed | Reporting response |
|---|---|---|---|
| unstated threshold | respondent defines “literate” | cognitive testing and question text | restrict interpretation to declared status |
| social desirability | stigma or perceived benefit affects answer | privacy review and validation subsample | disclose probable direction and uncertainty |
| combined skills | reading and writing collapsed | separate-item comparison | do not infer domain-specific ability |
| proxy knowledge | respondent does not observe another adult’s practice | self–proxy re-interview | identify proxy share and disagreement |
| schooling assumption | attainment substitutes for present skill | direct or self-report comparison | publish assumed component separately |
| missing language | eligible adult cannot use interview language | language non-interview count | qualify coverage and improve provision |
| mode change | response process differs over time | bridge study or parallel administration | classify trend as qualified or broken |
| interviewer variation | probing or coding differs | interviewer-level quality review | retrain, verify and correct where supported |
| non-response | participation related to literacy | response by group and follow-up evidence | weight cautiously and retain residual bias |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Appropriate uses
timing and relevant, service, access in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 33. Appropriate uses to connect evidence with delivery, unequal effect and correction. Graded questions can identify perceived difficulty relevant to service access.
It is weaker for estimating proficiency distributions, evaluating instruction or ranking populations with different languages and response conventions.
Interpretation boundary
A self-reported literacy rate is evidence of classification under the survey question and field conditions. It should not be described as an observed demonstration of literacy.
This boundary is not a dismissal. It is the condition for using the measure honestly and for deciding when complementary assessment is necessary.
Part III
Direct Assessment of Adult Literacy
What direct assessment adds
continuity and education, literacy, practices in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 35. What direct assessment adds to connect evidence with delivery, unequal effect and correction. It can distinguish levels and domains, examine distributions and relate performance to education, work and literacy practices.[REF-06] [REF-07]
The result remains conditional on the construct, tasks, language, administration and population reached. “Direct” does not mean complete or error-free.
Construct definition
distribution and rather, available, collection in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 36. Construct definition to connect evidence with delivery, unequal effect and correction. Content should follow the construct rather than an available collection of items.
A broad policy term such as functional literacy requires a narrower assessment statement before results can be interpreted.
Domain coverage
Prose, document, writing, numeracy and problem solving are related but distinct. An assessment may cover one or several.[REF-06] [REF-07]
The report title should not imply domains omitted by design. Composite reporting should retain domain-level evidence where policy conclusions differ.
Task authenticity
capacity and assessment, intends, measure in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 38. Task authenticity to connect evidence with delivery, unequal effect and correction. Familiar appearance does not alone establish validity; a task must require the capability the assessment intends to measure.
Everyday context can improve relevance while creating cultural or occupational differences that require review.
Task difficulty
equity and defensible, increases, demand in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 39. Task difficulty to connect evidence with delivery, unequal effect and correction. A progression should reflect defensible increases in demand.
Empirical performance assists calibration, but statistical difficulty does not by itself explain what capability a task requires.
Short verification tasks
timing and adults, simple, sentence in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 40. Short verification tasks to connect evidence with delivery, unequal effect and correction. It is feasible within a broader household survey and may identify adults able to read all, part or none of a simple sentence.
It does not measure writing, extended comprehension, documents or the full continuum of proficiency. Its economy depends on a strict interpretation boundary.
Multi-item assessment
uncertainty and capacity, respondent, cooperation in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 41. Multi-item assessment to connect evidence with delivery, unequal effect and correction. They can support scaled scores and proficiency distributions but require more time, technical capacity and respondent cooperation.
The burden should be considered in sample design and field protocol. Longer assessment is not automatically more valid if non-completion becomes selective.
Screening and routing
A screening stage may route adults to tasks suited to their demonstrated entry performance. Routing can reduce frustration and improve information at lower proficiency.
The rule, measurement consequence and treatment of non-attempts must be explicit. A screening failure should not create an unobserved group outside the reported distribution.
Item sampling
authority and scores, become, precise in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 43. Item sampling to connect evidence with delivery, unequal effect and correction. This expands content coverage while each person completes only part of the pool. Population proficiency can be estimated, but individual scores become less precise.
Published results should distinguish population inference from diagnostic claims about a named adult.
Translation and adaptation
coverage and convention, changes, difficulty in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 44. Translation and adaptation to connect evidence with delivery, unequal effect and correction. Literal equivalence may fail where grammar, script, word length or document convention changes difficulty.
Adaptation decisions require bilingual, subject and measurement review, field testing and documentation.
Multiple languages
remedy and comparisons, versions, differ in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 45. Multiple languages to connect evidence with delivery, unequal effect and correction. Each approach affects the population meaning. Choice can show capability in a preferred language but complicate comparisons where versions differ.
Results should record assessment language and avoid classifying performance in one language as literacy in every language.
Script and orthography
continuity and equivalence, psychometric, equivalence in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 46. Script and orthography to connect evidence with delivery, unequal effect and correction. Item design should not treat surface equivalence as psychometric equivalence.
Where oral and written language relationships differ, instructions and practice items require particular attention.
Cultural and contextual review
distribution and neither, possible, desirable in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 47. Cultural and contextual review to connect evidence with delivery, unequal effect and correction. Removal of every contextual feature is neither possible nor desirable.
The objective is to minimise irrelevant difficulty and document the context retained.
Administration conditions
feasibility and preserve, intended, comparison in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 48. Administration conditions to connect evidence with delivery, unequal effect and correction. Standardisation should be practical and sufficient to preserve the intended comparison.
Deviations and interrupted sessions should be recorded rather than coded automatically as low proficiency.
Interviewer and assessor competence
capacity and language, without, authorisation in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 49. Interviewer and assessor competence to connect evidence with delivery, unequal effect and correction. They should not teach the task, signal correctness or change language without authorisation.
Certification of field competence should be based on observed administration and correction, not attendance alone.
Accessibility
equity and changing, capability, measured in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 50. Accessibility to connect evidence with delivery, unequal effect and correction. An accommodation should remove irrelevant barriers without changing the capability being measured.
Exclusion from assessment should be reported by reason. An assessment that omits a population cannot support a whole-population claim without qualification.
Anxiety and stigma
timing and absence, individual, sanction in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 51. Anxiety and stigma to connect evidence with delivery, unequal effect and correction. Information should explain purpose, confidentiality, voluntary participation where applicable and absence of individual sanction.
Respectful administration is an ethical requirement and a data-quality condition.
Scoring
Scoring rules should define full, partial and incorrect responses, omissions and invalid administrations. Constructed responses require scorer training and reliability checks.
Changes made after fieldwork should be documented and applied consistently. Ambiguous items should be reviewed before population estimates are finalised.
Scaling
Scaled scores place performance from different item sets on a common continuum under a statistical model. Model fit, item behaviour and uncertainty require technical examination.
A scale is not a physical quantity with an absolute zero. Differences should be interpreted through task demand and sampling error.
Proficiency levels
Levels translate a score continuum into descriptions of tasks likely to be completed. They aid communication but create boundaries not present in the underlying scale.
Reports should provide distributions and standard errors and avoid describing all adults below a chosen level as having no literacy.
Plausible values and population use
coverage and method, estimates, variance in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 55. Plausible values and population use to connect evidence with delivery, unequal effect and correction. Users should follow the technical method for group estimates and variance.
Simple averaging of an individual-score field may understate uncertainty or produce biased relationships.
Quality control
Field monitoring, response checks, item analysis, scoring reliability, data cleaning and weighting should form one documented quality system.[REF-05] [REF-12]
Quality flags should lead to verification and defined treatment. Removal of inconvenient records without a rule can bias the distribution.
Table 3: direct-assessment validity record
| Domain | Required evidence | Principal risk | Release limitation |
|---|---|---|---|
| construct | framework and intended interpretation | tasks narrower than claim | name assessed domain only |
| content | item specification and review | irrelevant knowledge determines response | qualify or remove affected items |
| language | translation, adaptation and administration record | version difficulty differs | test equivalence and report language |
| sampling | frame, selection and weights | population not represented | restrict population claim |
| participation | contact, response and completion by group | weakest adults under-represented | non-response analysis required |
| accessibility | accommodation and exclusion record | disability becomes test barrier | qualify coverage and redesign access |
| administration | training, observation and deviations | conditions alter performance | flag or exclude under fixed rule |
| scoring | rules and agreement evidence | scorer variation | moderate and rescore |
| scaling | model, fit and uncertainty | false precision | report standard error and diagnostics |
| interpretation | level descriptions and boundaries | below-level equated with no literacy | publish continuum and task meaning |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Appropriate uses
distribution and change, designs, comparable in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 58. Appropriate uses to connect evidence with delivery, unequal effect and correction.
It is not automatically suited to individual certification, diagnosis or high-stakes allocation. The sampling and measurement design governs permissible use.
Direct-assessment boundary
feasibility and qualitative, knowledge, barriers in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 59. Direct-assessment boundary to connect evidence with delivery, unequal effect and correction. It does not observe every literacy practice, language or context and should not displace qualitative knowledge about barriers and use.
Its public authority depends on transparent design, participation and uncertainty, not on technical complexity alone.
Part IV
Comparability Across Methods, Populations and Time
Comparability as a claim
comparability and inadequate, ranking, estimates in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 60. Comparability as a claim to connect evidence with delivery, unequal effect and correction. It is a judgement about whether they can support a particular contrast. Two values may be adequate for broad description but inadequate for ranking or small trend estimates.
The intended use should therefore accompany every comparability decision.
Concept equivalence
authority and identical, comprehension, sentence in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 61. Concept equivalence to connect evidence with delivery, unequal effect and correction. A binary declaration of ability to read and write is not conceptually identical to a scale of prose comprehension or a short sentence task.
They may be examined together as different evidence about literacy, but their numerical results should not be treated as interchangeable.
Population equivalence
coverage and cohorts, differ, materially in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 62. Population equivalence to connect evidence with delivery, unequal effect and correction. A survey of ages 16–65 should not be compared directly with a census rate for all persons aged 15 and over where older cohorts differ materially.
Recalculation to a common age range may improve comparability if microdata and weights permit.
Time equivalence
Literacy data are collected irregularly. A table labelled with one year may combine observations from different years or modelled values. The observation year should be prominent.
Economic, educational and cohort change can make distant observations unsuitable for a current comparison.
Method equivalence
Self, proxy, indirect and direct methods should be coded explicitly. A method change can alter a time series independently of real population change.
Bridge studies using parallel methods on the same sample can estimate the direction and magnitude of the discontinuity.
Question equivalence
Small wording differences can change difficulty and threshold. Comparability review should preserve exact wording, response categories, routing and interviewer instructions.
A translated label in an international table cannot demonstrate equivalent national questions.
Language equivalence
feasibility and permitting, written, language in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 66. Language equivalence to connect evidence with delivery, unequal effect and correction. A national rate based on one official language differs in meaning from a rate permitting any written language.
The comparison should state whether it concerns literacy in a specified language or literacy demonstrated in any assessed language.
Respondent equivalence
A self-response series cannot be assumed comparable with a later household-proxy series. The proxy share may also change with interview timing, migration or employment patterns.
Respondent status should enter the method metadata and, where possible, published quality tables.
Educational assumption equivalence
equity and different, learning, opportunities in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 68. Educational assumption equivalence to connect evidence with delivery, unequal effect and correction. Even a common grade label may represent different learning opportunities.
Attainment-based estimates should be separated from reported or assessed literacy and should not be used to fill missing observations without disclosure.
Threshold equivalence
timing and thresholds, thresholds, equivalent in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 69. Threshold equivalence to connect evidence with delivery, unequal effect and correction. Similar percentages below two thresholds do not show that the thresholds are equivalent.
Linking requires common items, common persons or another defensible design with uncertainty.
Sampling equivalence
Different frames may omit remote areas, collective households, migrants or persons without stable addresses. Sampling stages and clustering affect precision.
Comparable point estimates remain misleading when one design excludes populations central to the policy question.
Weighting equivalence
Weights reflect selection probability, non-response adjustment and calibration. Differences can improve representativeness but also signal different assumptions.
Analysts should apply design weights and appropriate variance methods. Unweighted comparisons of complex samples are generally insufficient.
Coverage error
authority and examined, demographic, controls in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 72. Coverage error to connect evidence with delivery, unequal effect and correction. Coverage ratios should be examined with demographic controls.
Adjustment cannot fully correct a group that has no frame representation and no reliable external total.
Non-response equivalence
Overall response rates can be similar while group patterns differ. Contact failure, refusal and assessment non-completion have different implications and should be separated.
Adjustment models reduce known imbalance but cannot guarantee removal of bias related to unobserved literacy.
Standard errors
remedy and presented, definitive, ranking in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 74. Standard errors to connect evidence with delivery, unequal effect and correction. A difference smaller than its uncertainty should not be presented as a definitive ranking.
Non-sampling error remains outside the interval and requires separate discussion.
Rounding and rank
continuity and substantively, meaningful, difference in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 75. Rounding and rank to connect evidence with delivery, unequal effect and correction. Published tables should discourage ordinal claims unsupported by statistically and substantively meaningful difference.
Country order is a display choice, not a finding.
Trend breaks
distribution and during, bridge, period in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 76. Trend breaks to connect evidence with delivery, unequal effect and correction. The old and new series may overlap during a bridge period.
Silently joining the values can produce a false improvement or decline.
Comparability classes
feasibility and unlikely, reverse, conclusion in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 77. Comparability classes to connect evidence with delivery, unequal effect and correction. Qualified comparability permits bounded differences that are documented and unlikely to reverse the broad conclusion.
Descriptive-only comparison shows context without supporting magnitude or rank. Not supportable applies where differences are fundamental or unknown.
Table 4: comparability decision matrix
| Dimension | Strong | Qualified | Descriptive only | Not supportable |
|---|---|---|---|---|
| construct | same domain and interpretation | bounded domain difference | related literacy concept | unrelated or unspecified |
| population | harmonised age and coverage | small documented difference | material difference retained | population unknown |
| method | same or linked design | evaluated method variation | different known method | method unknown |
| question or task | equivalent and stable | minor tested adaptation | different known demand | wording or task unavailable |
| language | equivalent access and adaptation | bounded language difference | different language rule disclosed | language condition unknown |
| period | same or policy-relevant interval | modest documented interval | distant observation used as context | observation date unknown |
| participation | adequate and similarly distributed | residual bounded difference | material difference disclosed | selective participation unexamined |
| uncertainty | design-based and sufficient | approximate but decision-stable | point estimates only | precision cannot be judged |
| permitted use | magnitude, distribution and trend | broad magnitude with qualification | contextual juxtaposition | no comparative conclusion |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Comparison record
equity and should, policy, relevant in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 79. Comparison record to connect evidence with delivery, unequal effect and correction. The reason for inclusion should be policy-relevant.
Users should be able to identify why two values appear together and which inference is authorised.
Comparability conclusion
timing and assembled, unlike, measures in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 80. Comparability conclusion to connect evidence with delivery, unequal effect and correction. A transparent descriptive comparison can be more authoritative than an exact rank assembled from unlike measures.
The absence of full comparability should direct investment in better evidence, not justify treating current values as equivalent.
Part V
Coverage, Sampling and Participation
Participation as a statistical condition
remedy and population, represented, result in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 81. Participation as a statistical condition to connect evidence with delivery, unequal effect and correction. Failure at any stage can alter the population represented by the result.
Participation should therefore be analysed as part of measurement, not reported only as an operational total.
Target population
continuity and another, defined, population in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 82. Target population to connect evidence with delivery, unequal effect and correction. It should state age, residence, territory and any exclusions. A study may target usual residents, de facto residents or another defined population.
The choice affects migrants, temporary workers, displaced persons and adults who divide time between households.
Frame population
distribution and institutional, geographic, omissions in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 83. Frame population to connect evidence with delivery, unequal effect and correction. It may differ from the target population because of age, address, institutional or geographic omissions.
A coverage assessment should quantify known differences and identify groups for which no reliable estimate is possible.
Household-based coverage
feasibility and histories, literacy, access in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 84. Household-based coverage to connect evidence with delivery, unequal effect and correction. These groups may have distinctive educational histories and literacy access.
The published population should name exclusions. An estimate should not be called national whole-population evidence where important residents are outside the design.
Remote and sparsely populated areas
capacity and exclusion, substantively, important in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 85. Remote and sparsely populated areas to connect evidence with delivery, unequal effect and correction. Language diversity and educational access can make this exclusion substantively important.
Cost decisions should be explicit, and national summaries should not imply equal territorial coverage.
Adults in institutions
equity and affect, inclusion, privacy in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 86. Adults in institutions to connect evidence with delivery, unequal effect and correction. Institutional gatekeeping can affect both inclusion and privacy.
Where these adults are excluded, the result applies to the household population. Separate studies may be needed for policy concerning institutional education.
Homeless and mobile populations
timing and solely, because, absent in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 87. Homeless and mobile populations to connect evidence with delivery, unequal effect and correction. Mobility also affects repeated contact and weighting. Their omission should not be assumed negligible solely because they are absent from the frame.
Supplementary location-based or service-based approaches may inform planning, but estimates derived from them require their own population definition.
Sample size and precision
uncertainty and insufficient, linguistic, regional in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 88. Sample size and precision to connect evidence with delivery, unequal effect and correction. A large national sample may remain insufficient for a small linguistic or regional group.
Precision requirements should follow policy use. Publication of numerous unstable subgroup rates is not improved transparency.
Stratification
comparability and weights, reflect, design in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 89. Stratification to connect evidence with delivery, unequal effect and correction. Selection probabilities and analysis weights must reflect the design.
Post hoc division of a sample into many groups does not provide the same assurance as planned representation.
Clustering
Area-based surveys often select clusters of households. Adults within clusters may be more similar than adults selected independently, increasing variance for a given sample size.
Variance estimation should use the actual design. Treating clustered observations as a simple random sample produces unjustified precision.
Selection within households
coverage and available, confident, person in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 91. Selection within households to connect evidence with delivery, unequal effect and correction. The rule should prevent interviewers or households from choosing the most available or confident person.
Substitution of another adult changes selection probability and can bias the result. It should not be permitted without a controlled design.
Contact procedures
remedy and refusal, language, barrier in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 92. Contact procedures to connect evidence with delivery, unequal effect and correction. Repeated visits and varied timing improve inclusion. The contact record should distinguish unavailable address, no eligible adult, temporary absence, refusal and language barrier.
These categories support targeted field correction and later bias analysis.
Informed participation
continuity and require, literacy, assessed in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 93. Informed participation to connect evidence with delivery, unequal effect and correction. Information should not require the literacy level being assessed.
Oral explanation and accessible formats may be necessary. Consent obtained through unreadable material is not an adequate participation safeguard.
Refusal
Refusal may arise from time, distrust, fear, stigma or assessment burden. Conversion efforts should remain respectful and should not become coercive.
Refusal rates should be examined by area and observable characteristics. A low overall rate can conceal concentration in a group central to literacy policy.
Break-off and partial completion
feasibility and language, concern, performance in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 95. Break-off and partial completion to connect evidence with delivery, unequal effect and correction. Reasons include fatigue, difficulty, interruption, disability, language and concern about performance.
Partial completion should be categorised under predetermined rules. Treating every non-attempt as the lowest score may confound proficiency with access and participation.
Weighting for unequal selection
Base weights reflect inverse selection probabilities. Further adjustments may address non-response and align the sample with reliable population totals.[REF-12]
Weights can correct known imbalances under assumptions; they cannot recreate information for a wholly absent population or guarantee removal of literacy-related bias.
Non-response adjustment classes
comparability and create, unstable, weights in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 100. Non-response adjustment classes to connect evidence with delivery, unequal effect and correction. Excessively broad classes leave bias; excessively narrow classes can create unstable weights.
The procedure and weight distribution should be documented, including trimming and its effect.
Calibration
authority and consistently, survey, population in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 101. Calibration to connect evidence with delivery, unequal effect and correction. Controls should be reliable, temporally appropriate and defined consistently with the survey population.
Agreement on calibration variables does not ensure agreement on unobserved literacy. Residual bias remains an interpretation issue.
Imputation
coverage and because, principal, outcome in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 102. Imputation to connect evidence with delivery, unequal effect and correction. Imputing literacy proficiency or declared status requires stronger justification because the value is the principal outcome.
Observed and imputed shares should be reported. Imputation should not convert non-participation into apparently direct evidence.
Response-rate components
remedy and stages, conceals, mechanism in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 103. Response-rate components to connect evidence with delivery, unequal effect and correction. One percentage combining these stages conceals the mechanism of loss.
Comparable surveys should use consistent disposition rules and denominators.
Table 5: participation flow
| Stage | Required count | Quality question | Interpretation if lost |
|---|---|---|---|
| target population | estimated eligible adults | who is intended to be represented | defines inference |
| frame coverage | adults represented by frame | which groups are absent or duplicated | coverage error |
| selected sample | selected eligible units | was probability selection preserved | selection integrity |
| contacted sample | households or adults reached | are contact failures patterned | potential availability bias |
| cooperating adults | eligible adults agreeing | are refusals related to trust or burden | cooperation bias |
| background completers | usable contextual interview | can non-assessment be analysed | auxiliary evidence available |
| assessment starters | adults beginning tasks | did access or anxiety prevent start | participation barrier |
| valid completers | scoreable assessment evidence | are break-offs selective | achieved proficiency sample |
| weighted population | represented adults after adjustment | do weights rely on defensible controls | bounded population inference |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Participation profile
distribution and assessment, completion, separately in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 105. Participation profile to connect evidence with delivery, unequal effect and correction. It should also examine contact, refusal and assessment completion separately.
Differences guide field improvement and qualify substantive estimates. They should not be used to blame groups for non-participation.
Fieldwork correction
feasibility and authorised, language, support in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 106. Fieldwork correction to connect evidence with delivery, unequal effect and correction. Corrective action may include additional visits, reassignment, extended hours or authorised language support.
Correction should preserve probability selection and respondent rights. Replacing difficult cases with convenient ones is not correction.
Participation conclusion
capacity and reasoned, account, material in Adult Literacy Measurement: Comparability, Participation and the Limits of Self-Reported Data require ### 107. Participation conclusion to connect evidence with delivery, unequal effect and correction. Reliable adult literacy measurement requires an explicit chain from target population to valid completion and a reasoned account of every material loss.
Where participation is plausibly related to literacy, uncertainty extends beyond the sampling error and must constrain the public claim.
Part VI
Language, Culture and Accessible Measurement
Language as part of the construct
Literacy is exercised through particular languages and scripts. Assessment language is therefore not merely a delivery choice; it helps define what capability is observed.
Reports should specify whether the policy question concerns literacy in any language, an official language, a language of schooling or another defined set.
Language mapping
Design should begin with evidence on languages read, written and used by the target population. Spoken-language prevalence alone is insufficient because written use and script may differ.
The map should inform instrument languages, sample needs, recruitment and public interpretation.
Language choice
Respondent choice can respect capability and improve participation. It may also select among versions with different task characteristics. A controlled choice procedure should record preferred language, administered language and reason where they differ.
Forced use of one language answers a narrower question and should be labelled accordingly.
Translation process
Translation requires construct analysis, forward drafting, independent review, reconciliation and field testing. Back translation can identify some differences but cannot alone establish functional equivalence.
Reviewers should examine vocabulary, syntax, text genre, numerical convention and expected familiarity.
Adaptation register
Every departure from the source task should be recorded with reason, affected demand, reviewers and empirical evidence. Adaptation may be necessary when an institution, object or layout has no equivalent.
The objective is comparable demand, not identical surface form.
Item functioning across languages
Items that show unexpected differences between language groups after relevant proficiency is considered require investigation. Statistical evidence can signal a problem but cannot identify whether the cause is translation, culture, curriculum or genuine variation.
Content review and field evidence should accompany the analysis.
Multilingual adults
Adults may distribute reading and writing practices across languages. One language may be used at home, another at work and another for official documents.
A single-language assessment captures part of this repertoire. Background questions should record relevant practices without converting every language into a full assessment requirement.
Minority and indigenous languages
Excluding a written minority language can understate capability and obscure demand for literacy materials and education. Inclusion may require script, terminology and sampling expertise not available in a central instrument.
The decision and its limitations should be public, and development should involve competent language communities.
Oral languages and emerging written conventions
Some languages may have limited standardised written use or several orthographies. A conventional reading assessment may then measure exposure to one standard as much as general literacy.
Policy should distinguish oral capability, written-language development and literacy in other languages rather than assign a global deficit.
Sign languages and communication access
Instructions and consent may require sign-language access, while the assessed construct concerns engagement with written text. Interpretation should not supply answers or alter text demand.
The protocol should separate communication accommodation from substantive task modification.
Visual accessibility
Print size, contrast, lighting and layout can create irrelevant difficulty. Large print or other presentation may be provided where it preserves text and response demand.
Braille assessment, oral presentation or assistive technology may represent a related but different mode requiring separate validity analysis.
Physical access
Writing or page-handling tasks may disadvantage adults with motor impairments. A response accommodation can preserve the intended literacy construct if motor production is not part of that construct.
The use and effect of accommodation should be recorded without publicly identifying the adult.
Cognitive and learning differences
Instructions, pace and task load may affect adults with learning or cognitive impairments. Simplifying assessed text would change difficulty; providing accessible instructions and appropriate time may remove a procedural barrier.
No one adjustment is universally valid. The decision follows the construct and intended population claim.
Cultural familiarity
Documents, transactions and topics differ across settings. An unfamiliar form can create difficulty unrelated to the intended information-processing demand, while over-familiar content can advantage a subgroup.
Balanced task pools and expert review should reduce systematic irrelevant differences.
Gendered practices
Access to schooling, paid work, official documents and leisure reading may differ by sex and social setting. Tasks drawn mainly from one domain can reflect opportunity as well as capability.
Assessment should not erase these conditions. Background evidence helps interpret them and supports policy on access and use.
Rural and urban contexts
Text exposure and common documents may vary between rural and urban settings. Adaptation should preserve comparable cognitive demand without assuming that urban administrative materials define universal functionality.
Sampling and reporting should allow location differences to be examined where precision permits.
Respect and non-stigmatisation
Field language should not describe adults as deficient persons. It should describe observed performance under defined conditions and recognise the possibility of learning and varied practice.
Public categories should be tested for unintended stigma, particularly where results are published for small linguistic communities.
Table 6: language and accessibility assurance
| Assurance area | Evidence before fieldwork | Field record | Reporting boundary |
|---|---|---|---|
| language population | spoken and written language map | preferred and administered language | scope limited to offered languages |
| translation | construct-based review and testing | version identifier | no presumed equivalence without evidence |
| adaptation | reason and demand analysis | item version used | material change disclosed |
| multilingual practice | background questions | languages of use | one-language score not whole repertoire |
| cultural context | expert and participant review | difficulty observations | irrelevant familiarity considered |
| communication access | accessible information and instructions | support provided | support not treated as item performance |
| visual access | presentation and accommodation rules | format used and exclusions | whole-population claim qualified |
| physical access | response-method protocol | authorised adjustment | motor barrier not confused with literacy |
| non-stigmatisation | terminology review | complaint or distress record | categories describe performance, not worth |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Accessible reporting
Literacy findings should themselves be communicated in forms accessible to adults with varied literacy. Oral briefings, clear graphics and community-language explanations can accompany the technical report.
Simplification should preserve uncertainty and method. Accessibility is not permission to turn a qualified result into an absolute statement.
Language and accessibility conclusion
A literacy measure cannot be separated from the language, script and mode through which performance is elicited. Comparability requires equivalent construct access, not uniform administration that excludes part of the population.
The report should state who could demonstrate capability, under which conditions, and whose capability remains unobserved.
Part VII
Linking Self-Report and Direct Assessment
Purpose of linking
Collecting self-report and direct assessment from the same adults can show disagreement, examine reporting thresholds and support adjustment of broad survey indicators. It can also connect perceived difficulty and practice to observed performance.
The exercise should not assume that one measure contains every truth and the other only error.
Validation subsample
A probability subsample of respondents to a census-linked or household survey module may complete direct assessment. Selection probability, non-response and language access require separate weights.
Volunteers are unlikely to represent adults who decline or fear assessment and should not be used to calibrate a national rate without qualification.
Cross-classification
Self-reported categories can be cross-tabulated against proficiency levels or task outcomes. The table should show counts, weighted proportions and uncertainty.
Disagreement is expected because constructs and thresholds differ. It should be analysed, not labelled automatically as false reporting.
Sensitivity and specificity
If a direct threshold is adopted for a defined purpose, self-report sensitivity describes the share above that threshold reporting literacy; specificity describes the share below reporting non-literacy. Both depend on the chosen threshold and population.
They should not turn a contested literacy boundary into a natural fact.
Predictive models
Background variables and self-reported practices may predict assessed proficiency. Models should be developed and validated on separate data where feasible and should report error across groups.
Predicted proficiency is not directly assessed proficiency. Publication should not obscure that distinction.
Differential reporting
The relationship between self-report and assessment may differ by age, sex, language, education, location and social expectations. A single national correction factor can therefore misclassify subgroups.
Analysis should test interaction and retain adequate sample size and protection.
Proxy validation
Where both adult self-response and household proxy response are available, agreement can be analysed before direct assessment is considered. Three-way comparison can distinguish proxy disagreement from self-assessment disagreement.
Re-interview conditions should minimise learning or disclosure effects.
Method transition
A country moving from self-report to direct assessment should conduct an overlap or bridge study. Both methods are applied under controlled conditions during at least one period.
The published series should mark the transition. A difference between methods should not be presented as population change.
Composite systems
A layered system can use a short household module for broad coverage, a direct-assessment subsample for proficiency and qualitative enquiry for literacy practices and barriers. Each component retains its own inference.
Integration occurs through common identifiers, definitions and analysis, not by collapsing all evidence into one rate.
Table 7: method-linking record
| Linking element | Required design | Valid inference | Prohibited inference |
|---|---|---|---|
| overlap sample | probability selection and response weights | population relationship between methods | volunteer agreement as national validity |
| common population | aligned age, residence and language rules | method difference within defined group | difference attributed to time trend |
| cross-classification | weighted cells and uncertainty | pattern of agreement and disagreement | self-report labelled dishonest |
| threshold analysis | explicit policy threshold | conditional sensitivity and specificity | threshold treated as universal literacy divide |
| subgroup analysis | adequate protected samples | differential reporting patterns | unstable cells ranked |
| prediction | model validation and error | bounded estimate for stated use | predicted value called observed skill |
| bridge period | parallel administration | estimated series discontinuity | spliced series without marker |
| publication | separate results and interpretation | complementary evidence | one method silently replaces another |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Ethical linking
Linking requires clear authority, confidentiality and purpose. Adults should not experience a declaration as a commitment that triggers an undisclosed test or personal consequence.
Identifiers should be protected and removed from analytical files where no longer necessary.
Linking conclusion
Linked evidence can improve understanding of method effects and support a responsible transition to stronger measurement. It cannot make unlike constructs identical.
The value lies in explaining difference and uncertainty, not manufacturing a single definitive number.
Part VIII
Reporting Literacy Indicators Responsibly
Publication purpose
Literacy statistics should inform education, language, labour, social and community policy while respecting the people represented. Publication should permit independent understanding of the measure and its limits.[REF-09]
A headline rate without method and observation year is incomplete public evidence.
Core metadata beside the value
Every principal table should state source, observation period, population age, geographic coverage, method class, respondent rule, language condition and whether schooling assumptions were used.
Detailed technical notes may follow, but essential meaning should not be separated from the number.
Counts and denominators
Rates should be accompanied by weighted population counts where appropriate and by unweighted sample counts for survey transparency. These serve different purposes and should be labelled.
Small unweighted cells require suppression or caution even where weighted populations appear large.
Precision
Sample estimates should show standard errors, confidence intervals or another accepted measure of sampling uncertainty. Tables should indicate where design or sample size makes an estimate unstable.
Decimal places should reflect precision. Additional digits do not create information.
Non-sampling limitations
Coverage, non-response, language, question, mode, scoring and model limitations should be stated near the affected result. Sampling intervals do not capture these errors.
A concise direction-of-bias statement is useful where evidence supports it; otherwise uncertainty should not be assigned a convenient sign.
Distribution before averages
Direct assessments should report proficiency distributions and relevant percentiles or level shares, not only a mean. A common average can conceal different lower and upper tails.
Policy concerning minimum access requires attention to adults facing the greatest task difficulty without reducing them to a permanent label.
Disaggregation
Sex, age, location, language, education, work and socio-economic measures may reveal unequal opportunity and participation. Selection should follow policy relevance and sample capacity.
Disaggregation is not a licence for an unlimited table of unstable comparisons.
Time series
Charts should mark observation years and method breaks. Lines should not imply annual measurement where only intermittent observations exist.
Modelled interpolation and projection should be visually and textually distinct from observed values.
Cross-national tables
Country values should be grouped or annotated by method comparability. Alphabetical presentation is often preferable to an unsupported rank.
Where method differences are fundamental, separate panels are more honest than a single ordered column.
Threshold language
“Below the selected proficiency level” states a measurement result. “Illiterate population” may imply an absolute condition not established by the assessment, especially where one language or domain was tested.
Labels should follow construct and avoid stigma.
Associations and causes
Relationships between literacy and income, employment, health or civic participation are policy-relevant. Cross-sectional association does not establish that measured literacy alone caused the difference.[REF-06] [REF-07]
Education, background, opportunity and selection may contribute. Claims should match design.
Change and programme effect
Population change between surveys can reflect cohorts, migration, education, practice, economic conditions and measurement. It should not be attributed to one literacy programme without an appropriate evaluation.
Programme participants are also not a random population sample; their change answers a narrower question.
Public communication
Briefings should explain what adults were asked or required to do, who was included and what a level means in practical task terms. Visual material should preserve denominators and uncertainty.
Communication should identify actionable barriers and provision needs rather than sensationalise a population count.
Confidentiality
Small cells, rare languages and local areas can expose identities or stigmatise communities. Disclosure control should consider direct and inferential risk.
Suppression and aggregation should be applied consistently and should not conceal material inequality from authorised policy review.
Corrections
Errors in weights, coding, labels or tables should be corrected publicly with date, affected results and interpretive consequence. The original release remains documented.
Quiet replacement weakens official-statistics accountability.
Table 8: publication assurance
| Publication element | Minimum content | Misleading practice | Required correction |
|---|---|---|---|
| headline value | method, population and observation year | method-neutral current rate | restore scope beside value |
| table | counts, denominator, languages and source | rank of unlike methods | separate or classify comparison |
| uncertainty | sampling and material non-sampling limits | point estimate treated as exact | add precision and qualification |
| distribution | levels or score range | mean as full population account | publish distribution |
| trend | observed years and breaks | continuous line through missing years | mark observations and discontinuities |
| subgroup | relevance, sample and protection | unstable small-cell ranking | suppress, combine or qualify |
| narrative | association and bounded interpretation | causal claim from cross-section | revise to supported relationship |
| terminology | construct-respecting language | adult worth reduced to label | describe performance and context |
| correction | version, date and effect | silent replacement | issue correction notice |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Reporting conclusion
Responsible reporting makes the method visible and the policy meaning bounded. It does not weaken literacy advocacy; it ensures that public action responds to evidence rather than to artefacts of measurement.
The strongest statistic is not the most definite-sounding value, but the value whose population, method and uncertainty can withstand scrutiny.
Part IX
Applied Comparative Cases
Status of the cases
The following cases are constructed to examine measurement decisions. They do not describe identified countries and do not establish empirical conversion factors. Each case retains the population, method and participation conditions needed to understand the calculation.
Case A: a census self-declaration rate
A census asks one household respondent whether each member aged 10 and over can read and write a simple message in any language. The published adult rate uses persons aged 15 and over. Of 2,480,000 adults enumerated, 2,021,200 are recorded yes, 421,600 no and 37,200 unknown or not stated.
Dividing yes responses by all enumerated adults gives 81.5 per cent. Dividing yes by known responses gives 82.8 per cent. Neither denominator is inherently correct for every use. The first implicitly treats unknown as not literate; the second assumes missing status can be excluded without bias.
Metadata show that one respondent answered for 68 per cent of other adult members. The proxy share is higher among employed men absent at interview and younger adults temporarily away. The census provides valuable small-area coverage but does not establish observed proficiency.
The publication reports 2,021,200 declared literate adults, 421,600 declared not literate and 37,200 unknown. It uses 81.5 per cent only with an explicit denominator and shows 82.8 per cent as a known-response rate. Policy maps include unknown status rather than merging it with either category.
Case A record
| Component | Count | Share of all adults | Interpretation |
|---|---|---|---|
| enumerated adults aged 15+ | 2,480,000 | 100.0% | population denominator |
| declared can read and write | 2,021,200 | 81.5% | self or proxy classification under question |
| declared cannot | 421,600 | 17.0% | declared category, not direct assessment |
| unknown or not stated | 37,200 | 1.5% | retained separately |
| known responses | 2,442,800 | 98.5% | alternative denominator |
| yes among known responses | 2,021,200 | 82.8% | conditional rate requiring missingness caution |
| records supplied by proxy for another adult | 1,686,400 | 68.0% | method-quality characteristic |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Case A decision
The census result can support local identification of areas with high declared need, subject to proxy and missingness patterns. It cannot be compared as a proficiency rate with a direct assessment or used to infer writing and document skills beyond the question.
A validation survey should oversample areas with high unknown and proxy reporting and record language, self-response and performance separately.
Case B: direct sentence verification
A household survey asks adults with less than completed secondary education to read a sentence card. Adults with secondary completion are assumed able and are not tested. Among 8,000 sampled adults, 4,600 are assumed literate, 2,700 are asked to read and 700 have missing education information or cannot be routed.
Of those asked, 1,620 read the whole sentence, 540 read part, 405 cannot read it and 135 do not complete for language, vision, refusal or interruption. Reporting 4,600 assumed plus 1,620 whole-sentence readers as “literate” gives 6,220 of 8,000, or 77.8 per cent. The value combines indirect and direct methods and treats 1,780 adults outside the numerator for different reasons.
The report therefore provides a component table. It also produces a lower-bound observed whole-sentence count of 1,620 and does not present it as a population literacy rate because most adults were not assessed.
Case B record
| Routing result | Count | Measurement status | Publication treatment |
|---|---|---|---|
| assumed from secondary completion | 4,600 | indirect classification | publish separately |
| read whole sentence | 1,620 | direct narrow task | describe exact performance |
| read part | 540 | direct partial performance | retain ordered category |
| unable to read sentence | 405 | direct narrow task | no claim beyond task and language |
| task non-completion | 135 | unobserved direct performance | analyse reason |
| routing unresolved | 700 | education or eligibility missing | do not presume status |
| combined headline numerator | 6,220 | unlike evidence combined | use only with full method label |
| combined headline rate | 6,220 / 8,000 | 77.8% | not directly comparable with single-method rate |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Case B decision
The mixed indicator may meet an immediate reporting convention but is unsuitable for evaluating proficiency or comparing with a self-declaration series. The next survey should assess a probability subsample across all education levels to test the assumption and quantify task non-completion.
Adults who read part of the sentence should not be combined arbitrarily with either extreme. Their category contains policy information about emerging or limited reading capability.
Case C: specialised direct assessment
A specialised survey selects 5,400 adults aged 16–65. The frame excludes collective institutions and three remote districts containing an estimated four per cent of the otherwise eligible population. Of 5,400 selected adults, 4,590 complete the background interview and 4,050 provide valid assessment evidence.
Weighted results place 18 per cent below Level 1, 31 per cent at Level 1, 34 per cent at Level 2 and 17 per cent at Level 3 or above. Standard errors range from 0.8 to 1.2 percentage points. The figures describe the covered household population aged 16–65, not all adults aged 15 and over.
Non-completion is higher among adults whose background interview reports limited use of the assessment language. Weighting adjusts for age, sex, region and education but not directly for unobserved proficiency. The report presents sampling intervals and a residual non-response limitation.
Case C record
| Measure | Result | Uncertainty or coverage | Permitted statement |
|---|---|---|---|
| selected adults | 5,400 | probability sample | field sample |
| background completers | 4,590 | 85.0% of selected | contextual evidence available |
| valid assessment evidence | 4,050 | 75.0% of selected | achieved proficiency sample |
| below Level 1 | 18% | SE 0.8 points | covered population estimate |
| Level 1 | 31% | SE 1.1 points | covered population estimate |
| Level 2 | 34% | SE 1.2 points | covered population estimate |
| Level 3 or above | 17% | SE 0.9 points | covered population estimate |
| frame exclusion | about 4% of otherwise eligible population | institutions and remote districts | no whole-population claim |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Case C decision
The direct assessment supports distributional analysis and relationships with background variables within the covered population. It does not provide a directly comparable replacement for the national census rate because age, coverage, method and construct differ.
The next cycle should test language-access improvements and develop evidence for excluded districts. A more precise point estimate within the current frame would not resolve those coverage limitations.
Case D: apparent trend after a method change
A country reports self-declared adult literacy of 76.4 per cent in an earlier census. Five years later, a household survey using a sentence-reading task reports 69.8 per cent reading the whole sentence, 8.6 per cent reading part and 21.6 per cent unable or unobserved under its published rule.
The seven-point difference is presented initially as decline. Review finds different age ceilings, proxy use in the census, testing only in two languages and exclusion of remote areas in the survey. There is no overlap sample.
The values cannot establish decline. The census describes declared reading and writing in any language; the survey describes sentence performance under a restricted language and coverage design. They remain useful as separate observations.
Case D comparison record
| Dimension | Earlier census | Later household survey | Comparability judgement |
|---|---|---|---|
| population | age 15+, national census scope | ages 15–64, remote areas excluded | material difference |
| construct | can read and write simple message | reads whole or part of sentence | related, not equivalent |
| respondent | 72% proxy for other adults | adult performs task | method break |
| language | any language declared | two assessment languages | material access difference |
| result | 76.4% declared yes | 69.8% whole sentence | no direct subtraction |
| partial category | absent | 8.6% | distribution not binary-equivalent |
| bridge study | none | none | difference cannot be allocated to method or time |
| trend classification | — | — | not supportable |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Case D decision
The time series should show a discontinuity and explain both methods. A future bridge study can administer the census question and the direct task to the same probability sample and examine results by language and respondent status.
Until then, programme claims should use other evidence and avoid attributing the difference to national change.
Case E: preferred-language assessment
A multilingual assessment offers three language versions. Of 3,200 participating adults, 1,920 select Language A, 880 Language B and 400 Language C. Translation review is complete, but item analysis finds four document tasks easier in Version B after overall proficiency is considered.
The national distribution combines versions under the original scale. The language-specific comparison is withheld pending review of those items. Removing them changes the mean for Version B by 6 scale points and the national mean by 1.2 points.
The national result may remain sufficiently stable for broad reporting, while language-group rank is not supportable. The publication documents the item decision and retains preferred-language participation as a strength of coverage.
Case E decision record
| Evidence | National estimate effect | Language comparison effect | Decision |
|---|---|---|---|
| four flagged document tasks | under review | Version B unexpectedly easier | investigate content and translation |
| exclusion sensitivity | mean changes 1.2 points | Version B mean changes 6 points | national broad conclusion stable; subgroup rank unstable |
| preferred-language availability | wider participation | different version exposure | retain and report language choice |
| sample by version | A 1,920; B 880; C 400 | unequal precision | publish standard errors |
| final release | combined distribution with note | no ordered language comparison | issue technical qualification |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Case E boundary
Different average scores by assessment language cannot be interpreted as language effects because populations choosing each language differ in education, age, location and opportunity. Version analysis addresses measurement equivalence, not the social cause of group differences.
Policy use should focus on access to adult learning and written materials in relevant languages, supported by broader evidence.
Case F: non-response adjustment
A direct assessment selects 10,000 adults. Valid assessment evidence is obtained from 7,200. Response is 82 per cent among adults with upper-secondary education and 61 per cent among adults with primary education or less. Base-weighted proficiency above a selected level is 58.0 per cent.
Non-response weights formed by age, sex, region and education reduce the estimate to 54.6 per cent. A follow-up of 300 initial non-respondents obtains a short task from 174 and indicates lower performance than respondents within the same education groups.
The adjusted 54.6 per cent remains potentially high because the weighting variables do not capture all response-related proficiency. The report presents the estimate, sampling error and a directional residual-bias statement rather than another precise correction from the small follow-up.
Case F record
| Stage | Estimate above selected level | Evidence | Interpretation |
|---|---|---|---|
| base-weighted respondents | 58.0% | selection weights only | respondent distribution |
| adjusted main estimate | 54.6% | age, sex, region and education adjustment | improved population estimate |
| valid completions | 7,200 of 10,000 | 72.0% | material non-response |
| higher-education response | 82% | field disposition | response related to education |
| lower-education response | 61% | field disposition | under-representation likely |
| follow-up completion | 174 of 300 traced cases | selective small follow-up | directional evidence, not full correction |
| residual judgement | probable upward bias remains | lower follow-up performance | qualify estimate |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Case F decision
The adjusted estimate is preferable to the base-weighted respondent value, but neither weighting nor follow-up proves elimination of bias. The next design should strengthen initial participation and collect auxiliary information for all selected adults.
Publication should not select 58.0 per cent because it is the least adjusted or 54.6 per cent because it appears more conservative. It should use the method judged most defensible and disclose the residual limitation.
Case G: programme evaluation and population indicators
An adult learning programme assesses 640 entrants and 472 completers. Mean scores rise by 18 points among completers. A national household indicator improves by two percentage points during the same period.
The programme result is affected by attrition, practice, instruction and assessment conditions. The population indicator includes non-participants and different cohorts. Neither result alone establishes the programme’s population effect.
Baseline scores are available for 143 non-completers, whose mean is lower than that of completers. Complete-case change is therefore likely to overstate the result for all entrants.
Case G record
| Evidence | Population represented | Finding supported | Finding not supported |
|---|---|---|---|
| 640 entrant baselines | enrolled entrants | starting distribution | general adult population |
| 472 matched completers | retained participants | change among completers | change among all entrants |
| 168 without follow-up | attriting entrants | attrition count and baseline difference | zero change without evidence |
| 18-point mean gain | matched completers | observed within-person difference | causal effect without comparison |
| national two-point change | covered household population | population indicator movement | programme contribution alone |
| participation records | programme exposure | reach and completion pattern | proficiency of non-participants |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Case G decision
Evaluation should report entry, retention and matched change, examine attrition and use an appropriate comparison where feasible. The national indicator provides context but cannot serve as the programme counterfactual.
Programme success should include participation and equitable reach as well as assessment change.
Case H: local-area estimates from a national survey
A national survey is asked to publish literacy estimates for 96 districts. In 41 districts the unweighted adult sample is below 40, and several estimates have relative standard errors above 25 per cent. Direct publication would invite unstable ranking.
The survey can provide reliable estimates for six regions. Model-assisted district estimates may be developed using census covariates, but they are partly predicted and depend on model assumptions.
The report publishes regional direct estimates and classifies district results as experimental model-based planning evidence with uncertainty intervals and validation diagnostics.
Case H record
| Geographic product | Direct sample condition | Estimation status | Release decision |
|---|---|---|---|
| national | adequate probability sample | direct design-based | publish |
| six regions | planned strata and acceptable precision | direct design-based | publish with standard errors |
| 55 districts | at least 40 observations but variable precision | direct estimate often unstable | publish only where quality rule passes |
| 41 districts | fewer than 40 observations | inadequate direct precision | do not rank or release as direct rate |
| model-assisted districts | survey plus census covariates | partly predicted | label separately with intervals |
| district change | one survey round | no stable time series | do not infer trend |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Case H decision
Small-area need remains important, but unstable direct rates do not meet it. Census self-report, administrative information, qualitative evidence and model-assisted estimates can be considered together with distinct labels.
Resource allocation should not depend on an unqualified single district rank where uncertainty can reverse the order.
Cross-case finding
The cases show that most serious errors occur when unlike states are collapsed: unknown with not literate, assumed with assessed, non-completion with lowest proficiency, method change with trend, language version with group effect, weighted adjustment with removal of all bias, programme change with national impact, or predicted local values with direct observations.
Reliable comparison preserves these distinctions and makes the resulting policy choice explicit.
Table 17: cross-case error and remedy
| Case | Collapsed distinction | Misleading conclusion | Required remedy |
|---|---|---|---|
| census declaration | unknown versus declared no | missing adults classified without evidence | publish components and denominators |
| sentence task | assumed versus directly observed | mixed rate called assessed literacy | separate method classes |
| specialised assessment | covered versus target population | partial frame called all adults | restrict population claim |
| method-change trend | measurement versus time | decline inferred from unlike values | mark break and bridge methods |
| multilingual assessment | version effect versus population difference | language groups ranked | test items and withhold unsupported rank |
| non-response | adjusted versus unbiased | weight treated as complete correction | disclose residual bias |
| programme evaluation | completers versus entrants and population | programme credited with national change | analyse attrition and comparison |
| district estimate | direct versus predicted | unstable local ranks | quality rules and separate model status |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Applied-case conclusion
Adult literacy indicators gain authority when the analytical record permits a reader to reconstruct who was represented, what was observed and why a comparison was allowed. Precision without these distinctions can conceal rather than reduce uncertainty.
Part X
Policy Uses and Institutional Responsibilities
Measurement in service of literacy policy
The Dakar Framework identifies adult literacy as a central education commitment and calls for substantial improvement, particularly for women, together with equitable access to continuing education. Measurement should make unmet need and unequal opportunity visible and support decisions on provision.[REF-01]
The indicator is a means of public action. It should not become a substitute for investment in learning opportunities, language materials and environments in which literacy can be used.
National policy diagnosis
National diagnosis should combine prevalence, distribution, language, age, education, location, work and participation evidence. One total rate cannot show whether low observed proficiency arises principally among older cohorts, isolated areas, particular language communities or adults with interrupted schooling.
Policy should respond to the pattern and to direct evidence about barriers, not to a headline rank.
Planning adult learning provision
Direct assessment can inform the range and level of programmes required; self-reported difficulty and practice can inform access, scheduling and relevance. Small-area census data can support geographic placement where the method boundary is retained.
Provision should not restrict admission to adults who fail a particular task. Assessment for planning differs from an eligibility test.
Identifying language needs
Measurement should distinguish language of literacy from general ability. Data on languages used for reading and writing can guide materials, instructors and communication.
A low result in an official-language assessment may indicate a need for official-language learning, but it should not erase capability in another language or justify exclusion from services.
Gender equality
Sex-disaggregated results remain necessary because access to schooling, adult learning, time, mobility and literacy use may differ substantially. The 2006 global development and literacy monitoring context gives particular attention to gender disparities.[REF-02] [REF-14]
Analysis should examine survey participation and proxy response by sex as well as the substantive result.
Cohort analysis
Age patterns can reflect historical differences in educational access and opportunities for use. They should not be interpreted solely as individual decline. Cross-sectional differences between age groups are not the same as longitudinal skill loss.
Cohort analysis can nevertheless help plan programmes and anticipate change as younger and older cohorts enter and leave the adult population.
Educational attainment
Combining literacy and attainment evidence can reveal adults whose skills exceed or fall below what formal qualifications might suggest.[REF-06] [REF-07]
Policy should avoid assuming that school completion guarantees current proficiency or that adults without schooling lack all literacy. The two indicators answer different questions.
Employment and workplace learning
Literacy proficiency and practices are associated with employment, training and income, but relationships are shaped by opportunity and selection.[REF-06] [REF-07]
Workplace policy should expand opportunities to learn and use skills without using a population assessment to screen individual workers beyond its design.
Health and public services
Adults may face written demands in health, finance, transport and public administration. Literacy evidence can guide clearer communication and assisted service channels.
The appropriate response is not only to change the adult. Institutions should reduce unnecessary text complexity and make essential information accessible.
Civic participation
Literacy supports access to public information and participation, while civic activity also creates purposes for literacy. Measurement may include relevant practices without treating political participation as a literacy test.
Public communication of results should enable communities to deliberate about provision and resource priorities.
Poverty analysis
Household surveys can relate literacy indicators to consumption, income, employment and living conditions. Association should be examined with household composition, location, education and opportunity.[REF-11] [REF-13]
A poverty gradient does not establish that literacy alone caused economic status. Policy may need coordinated educational and social action.
Programme targeting
Geographic or group evidence can guide outreach, but targeting should avoid stigma and ecological error. A resident of a low-rate area should not be presumed to have a particular level.
Individual entry assessment should serve placement and support, with privacy and an opportunity to demonstrate learning needs through an appropriate method.
Resource allocation
Allocation formulas may include population size, measured need, access barriers and service cost. Indicator uncertainty and method differences should be reflected rather than hidden.
A small difference in estimated rate should not move substantial resources where sampling or non-sampling uncertainty can reverse the order.
Monitoring national commitments
Global and national monitoring requires stable definitions and observation dates. The international Literacy Decade and Education for All commitments emphasise strengthened literacy action and monitoring.[REF-01] [REF-15]
Monitoring should not reward countries for adopting a method that yields a higher rate. Improvements in measurement may initially lower or redistribute estimates and should be recognised as statistical progress.
Evaluating policy
Repeated comparable assessment can show population change but does not by itself attribute change to a policy. Evaluation should examine exposure, implementation, cohort effects, economic conditions and other plausible explanations.
Where only one pre- and post-observation exists, causal conclusions should remain limited.
Early-warning use
Changes in programme participation, literacy practice or a short module may provide earlier signals than a full assessment cycle. These indicators can prompt enquiry but should not be presented as substitutes for proficiency change.
An early-warning threshold should have a defined verification and response.
Local planning
Local authorities need evidence at a usable geographic scale. Census breadth, administrative programme data and community enquiry may support local planning even when specialised assessment is reliable only regionally.
The sources should remain separate. A locally reported declaration rate should not inherit the measurement authority of a national direct assessment.
Community participation
Adults, educators and language communities should contribute to construct relevance, task review, field access and interpretation. Their participation can identify practices and barriers absent from technical review.
Participation does not allow interest groups to suppress unfavourable findings. The statistical authority retains responsibility for method and fair release.
Education ministry
The education ministry should define policy questions, provide programme and education-system knowledge and use findings for provision. It should not alter statistical results to align with programme claims.
Joint governance should preserve the responsibilities and independence of each body.
Census office
The census office can provide broad population and local-area evidence and maintain exact question and respondent metadata. It should evaluate proxy response, assumptions and non-response.
Where a specialised assessment is planned, common background variables and a validation subsample can strengthen linkage without overburdening the census.
Assessment body
The body responsible for direct assessment should maintain the framework, item security where required, language versions, field standards, scoring, scaling and technical documentation.
It should state whether the design supports population estimates, individual results or both. High-stakes reuse requires separate evidence.
Adult education providers
Providers can contribute practical knowledge about learner goals, barriers and programme assessment. Their participant data describe service users, not the unserved population.
Providers should not be asked to produce national prevalence estimates from enrolment records.
Research institutions
Independent research can test validity, non-response, language effects and policy relationships. Access to protected microdata should follow clear conditions and disclosure control.
Publication freedom and replication improve confidence, while confidentiality remains binding.
International comparison
International organisations can support common definitions, technical capacity and comparable assessments. They should also preserve national metadata and avoid implying equivalence where methods differ.
Countries should participate in methodological decisions and retain the capacity to interpret results in their own linguistic and institutional context.
Procurement and technical services
Where external services support sampling, printing, data collection or analysis, contracts should specify method, quality evidence, ownership, confidentiality, transfer and correction.
Technical delegation does not transfer public accountability for the indicator.
Governance committee
A national literacy measurement committee may coordinate policy questions, languages, population coverage, field access and release. Membership should include statistical, education, adult learning and relevant language expertise.
The committee should record decisions and conflicts without compromising the statistical authority’s professional responsibility.
Confidentiality governance
Literacy data can expose personal circumstances and create stigma. Access should be limited by role, identifiers separated where feasible and outputs tested for disclosure.
Individual results should not be transferred to employers, benefit authorities or enforcement bodies unless a clear lawful purpose and appropriate design exist.
Release calendar
The observation period, processing schedule, preliminary status and final release date should be published in advance where practicable. Equal access to principal results supports impartiality.
Delay for unresolved quality concerns should be explained; delay to secure a favourable policy narrative is inappropriate.
Capacity development
Sustainable measurement requires sampling, field, language, data processing, psychometric, analytical and dissemination capability. Short external missions cannot replace national institutional development.
The contemporary LAMP initiative places capacity alongside improved assessment and policy relevance.[REF-03]
Cost and periodicity
Censuses offer breadth at long intervals; household modules can recur more often; specialised assessments provide depth at higher cost. A national architecture should combine them according to decision cycles.
Reducing sample quality or language access to obtain annual data may produce a less useful series than sound measurement at wider intervals.
Burden on adults
Assessment time, travel, anxiety and repeated contact are real burdens. Instruments should retain questions and tasks that serve defined analysis and avoid duplicating information available reliably elsewhere.
Burden review should include adults who did not complete, not only those who tolerated the full design.
Public accountability
Institutions should report method decisions, cost, coverage, response, quality findings, release and correction. Accountability concerns the integrity and usefulness of evidence, not whether the rate meets a political expectation.
An unfavourable result can be responsibly produced; an unexplained favourable result cannot.
Table 18: institutional responsibility map
| Function | Lead responsibility | Required control | Public output |
|---|---|---|---|
| policy question | education and adult-learning authorities | decision use and public-interest test | measurement brief |
| population frame | statistical authority or census office | coverage and selection record | population scope |
| construct and tasks | assessment body with literacy expertise | framework and validity evidence | assessment description |
| language versions | assessment and language specialists | adaptation and equivalence review | language coverage note |
| fieldwork | statistical or contracted field body | training, contact and deviations | response profile |
| weighting and estimation | statistical authority | reproducible methods and variance | estimates with uncertainty |
| policy interpretation | responsible ministries and communities | claim-source boundary | policy response |
| statistical release | statistical authority | impartiality, metadata and correction | controlled publication |
| confidentiality | data custodian | access and disclosure controls | protected public tables |
| evaluation | independent or functionally separate review | question, comparison and limitations | findings and management response |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Institutional conclusion
No single body holds all knowledge required for adult literacy measurement. Coordination is necessary, but responsibility should remain visible. Statistical integrity, educational relevance, language competence and public participation are complementary controls.
The resulting indicator should be judged by whether it supports better and more equitable learning opportunities, not by whether it produces a convenient national figure.
Part XI
A National Adult Literacy Measurement Architecture
Architectural principle
A national system should combine breadth, depth, context and continuity. Attempting to obtain all four from one instrument either overloads the instrument or leaves essential questions unanswered.
Each component should have a defined role and a documented relationship to the others.
Component 1: census literacy module
A census module can provide broad coverage, small-area evidence and a population frame. It should use tested questions, record proxy status where feasible and avoid deriving literacy solely from schooling.[REF-08]
Its output is a declared or reported classification under the census method, not a proficiency distribution.
Component 2: household survey module
A recurring household module can collect graded self-assessment, literacy practices, language and access to learning alongside social and economic variables.[REF-11] [REF-13]
It can monitor perceived difficulty and context more frequently than a specialised assessment, provided wording and mode remain stable.
Component 3: direct assessment
A periodic probability assessment can estimate proficiency distributions across defined domains and relate them to background conditions. Its sample, languages and administration require dedicated quality control.
It should be timed to policy and evaluation needs rather than forced into an annual reporting cycle.
Component 4: programme evidence
Adult education programmes should maintain enrolment, attendance, completion and learning evidence suited to their objectives. Common core fields can support system analysis, but local relevance remains important.
Programme data describe participants and should not fill national prevalence gaps.
Component 5: qualitative enquiry
Interviews, group discussions, observation and community studies can examine literacy purposes, stigma, language, service access and reasons for non-participation. A stated sampling and analysis method is required.
Qualitative evidence explains mechanisms and meanings that a scale cannot fully represent.
Common population concepts
Components should align core age, sex, residence, geography, language and education definitions where feasible. Differences required by purpose should be mapped.
Common labels should not conceal different eligibility or coverage rules.
Common identifiers and privacy
Statistical linkage may use protected identifiers under lawful authority. Public and general analytical files should remove direct identifiers and limit detail.
Linkage value should be weighed against disclosure risk and public trust.
Question bank
A controlled bank can retain tested self-report, practice and background questions with wording, translations, response categories and evidence. Instruments select from it under version control.
Unrecorded local alteration prevents comparison and should be avoided.
Assessment framework maintenance
The framework should be reviewed for social and linguistic relevance without changing the construct silently. New task types may be introduced through field trials and linking designs.
Secure items and released examples should together support integrity and public understanding.
Language programme
Language inclusion requires a planned process for population mapping, prioritisation, translation, adaptation, recruitment, sampling and analysis. It should not be arranged only after fieldwork begins.
Decisions to include or exclude languages carry both validity and equity consequences.
Master sample or coordinated frames
Coordinated area frames can reduce cost and improve consistency across household modules and direct assessment. Sample rotation should manage respondent burden and preserve independent selection.
Frame updates are necessary where migration, growth or displacement changes coverage.
Bridge studies
Bridge studies should be scheduled when questions, modes, languages, age limits or assessment frameworks change. Parallel administration allows method effects to be distinguished from population change.
The bridge sample must be large and representative enough for the intended reconciliation.
Metadata registry
The registry should contain every indicator version, exact question or framework, population, source, period, method, languages, respondent rule, weighting and limitations. Tables and public summaries refer to the registry entry.
Historical versions remain accessible so that trend breaks can be understood.
Observation calendar
The calendar identifies census years, household modules, assessment cycles, programme reporting and policy reviews. It permits evidence to be combined without pretending that all sources are contemporaneous.
Major policy evaluation should be scheduled around an adequate baseline and follow-up rather than an administrative anniversary alone.
Analysis plan
The analysis plan should be approved before principal results are known. It defines populations, indicators, weights, variance, proficiency reporting, subgroup analysis, missing data and comparison rules.
Exploratory findings may be reported as such, with protection against selective emphasis.
Quality assurance plan
Quality assurance covers frame, selection, instrument, translation, fieldwork, scoring, data processing, weighting, analysis, disclosure and dissemination. Responsibilities and evidence should be assigned before collection.
Quality is not a final inspection applied after methodological decisions can no longer be corrected.
Pilot and field trial
A pilot tests operational flow, consent, contact, burden and logistics. A field trial tests items, language versions, scoring and measurement properties. The functions overlap but are not identical.
Both should include adults likely to encounter access and participation barriers.
Main collection readiness
Readiness requires an approved framework, stable instruments, trained personnel, verified systems, language materials, sample release, protection arrangements and contingency. Failure of a critical condition should delay or stage collection.
Calendar pressure should not turn an unresolved language or scoring issue into accepted error.
Processing and reconciliation
Data processing should preserve field dispositions, original responses, edits, scoring and weights. Changes require rules and audit trails.
Counts should reconcile from selected sample through final estimates. Unexplained loss between stages prevents release.
Technical review
Reviewers should examine construct, sampling, field response, language, scoring, scaling, weighting, uncertainty and disclosure. Findings and management responses should be recorded.
Independence should be described accurately; internal technical challenge is not the same as external review.
Preliminary results
Preliminary results may support timely policy where processing is sufficiently mature. Their status, possible revision and excluded analyses should be prominent.
They should not be used for detailed ranking or target judgement before quality review is complete.
Final release
The final package should include the principal report, methodological documentation, tables, comparison classifications and accessible public communication. Microdata access conditions should be stated.
Release should occur only after numerical, metadata, confidentiality and interpretation checks agree.
Policy response
Responsible authorities should respond to the distribution, participation barriers and provision implications. The response should distinguish immediate action from matters requiring further evidence.
A measurement report should not itself claim that resources have been authorised or services delivered.
Evaluation of the system
After each cycle, the system should review coverage, response, language access, burden, errors, corrections, use and cost. It should identify which indicators informed decisions and which collections were unused.
Improvement should simplify as well as add. A field with no credible use should be retired.
Table 19: national architecture
| Component | Principal strength | Principal limitation | Recommended periodic role |
|---|---|---|---|
| census module | population breadth and local area | short self or proxy classification | long-interval coverage benchmark |
| household module | context and recurring practices | indirect or declared capability | interim monitoring and diagnosis |
| direct assessment | proficiency distribution and domains | cost, sample and participation | periodic depth and trend |
| programme records | implementation and participant progress | no evidence on unserved population | continuing service management |
| qualitative enquiry | meaning, barriers and mechanisms | not a prevalence estimator | targeted policy explanation |
| bridge study | estimates method discontinuity | additional sample and burden | whenever major method changes |
| metadata registry | preserves definition and history | requires active custody | continuous control |
| public consultation | relevance and trust | cannot replace statistical method | design and interpretation stages |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Minimum viable architecture
Where resources are limited, the minimum system should maintain an improved household or census module, a probability validation subsample with a direct task, complete method metadata and a participation profile. Expansion to a multi-domain assessment should follow capacity and policy need.
The minimum should not be achieved by excluding languages or difficult-to-reach groups without disclosure.
Mature architecture
A mature system links periodic multi-domain assessment with census and household evidence, programme monitoring, language capacity, research access and policy evaluation. It preserves comparable trends while allowing controlled revision.
Maturity is demonstrated by evidence use, correction and public trust, not the number of instruments.
Architecture conclusion
A layered national architecture avoids demanding that one indicator provide small-area coverage, detailed proficiency, annual trend and programme attribution simultaneously. It assigns each method a defensible role and makes their relationships explicit.
This is the practical route from a single uncertain rate to a coherent public evidence system.
Part XII
Quality Evaluation and Decision Rules
Fitness for use
Quality should be judged against the decision an indicator is expected to support. A census declaration may be fit for broad local outreach and unfit for proficiency ranking; a direct assessment may be fit for national distribution and unfit for district allocation.
The quality statement should name both the supported and unsupported uses.
Relevance
Relevance concerns whether the construct, population, period and reporting level answer a current policy question. A technically strong measure can be irrelevant when it concerns the wrong language, age range or domain.
Policy users should participate in defining the question without controlling the result.
Accuracy
Accuracy concerns closeness between the reported result and the condition defined by the measurement design. It includes sampling and non-sampling error.
No single statistic establishes accuracy. Evidence comes from validation, consistency, field controls, item analysis, response patterns and comparison with appropriate external information.
Reliability
Reliability concerns consistency under repeated or parallel measurement conditions. It may be examined for items, scores, coders, interviewers and repeated questions.
High reliability does not establish that the correct construct was measured. A consistently applied narrow question can remain invalid for a broad claim.
Validity
Validity concerns whether evidence and theory support the intended interpretation and use. It encompasses content, response processes, internal structure, relationships with other variables and consequences.
Validity belongs to an interpretation, not permanently to an instrument title.
Timeliness
Literacy results should be released while they remain relevant, but speed should not displace necessary processing and review. Observation date and release date should both be visible.
Preliminary release is appropriate only where its scope and revision risk are clear.
Coherence
Coherence concerns whether related outputs use compatible definitions and tell an explainable account. Census, survey and programme results need not agree numerically, but their differences should be traceable to method and population.
Unexplained contradiction requires investigation rather than selective publication.
Accessibility and clarity
Users should be able to locate, understand and use the result and metadata. Technical documents, accessible summaries and machine-readable tables serve different audiences while preserving one approved meaning.
Clarity does not require omission of uncertainty.
Interpretability
Interpretability requires definitions, classifications, sources, methods, uncertainty and examples of task demand. Proficiency levels should be described through what the scale supports, not through moral or social labels.
Historical values require archived metadata to remain interpretable.
Comparability
Comparability should be rated for each principal contrast. A dataset may be internally comparable across regions but not over time after a method change.
The rating should appear in the analytical table, not only in a distant technical note.
Confidentiality and integrity
Quality includes protection of respondent information and resistance to unauthorised alteration. Confidentiality encourages trust and protects adults from harm; integrity protects the public record.[REF-09]
Both require assigned responsibility and documented incident response.
Cost-effectiveness
Cost should be considered in relation to decision value, precision, coverage and capacity developed. The least expensive indicator may become costly if it directs resources incorrectly.
Cost evaluation should include respondent burden and opportunity cost, not finance alone.
Error profile
An error profile records frame, sampling, non-response, measurement, processing, modelling and disclosure limitations. It identifies likely direction and magnitude only where evidence permits.
The profile should be revised as quality studies produce new information.
Critical errors
Some errors prevent the intended result from being released: lost sample-selection records, uncontrolled item exposure, unresolvable scoring corruption, absent population definition or a confidentiality breach affecting the publication.
Criticality should be defined before final review and linked to an authorised response.
Material errors
A material error could change a principal estimate, comparison, policy conclusion or public understanding. It requires correction or explicit qualification.
Materiality is substantive as well as numerical. A language omission can be material even where the national mean changes little.
Tolerable limitations
Some limitations can be accepted when bounded and unlikely to alter the intended decision. Acceptance should name the limitation, evidence, authority and use restriction.
Repeated acceptance should not become a substitute for improvement.
Verification rules
Verification should include sample selection, instrument version, response disposition, scoring, weight construction, table calculation and narrative claims. High-risk steps receive independent repetition or source checks.
Verification is complete only when a discrepancy is resolved and all affected outputs agree.
Replication
An authorised analyst should be able to reproduce principal estimates from protected data, documented code or calculation instructions and metadata. Replication records the software, weight, variance and exclusion rules used.
Public-use data may support wider replication where confidentiality allows.
Sensitivity analysis
Sensitivity analysis tests plausible alternative treatments of missing data, thresholds, flagged items, weights or population definitions. It shows whether the policy conclusion depends on one contestable choice.
Alternatives should be selected on methodological grounds, not to locate the preferred result.
External consistency
Literacy results may be compared with education, age, language and programme evidence to identify unexpected patterns. Consistency strengthens plausibility but does not prove correctness because related sources may share assumptions.
Unexpected differences require explanation and may reveal valuable new information.
Field indicators
Contact attempts, duration, refusal, break-off, assessor observations, language mismatch and deviations provide early quality evidence. They should be analysed during collection.
Field indicators should trigger support or verification without encouraging staff to pressure adults into participation.
Scoring indicators
Scorer agreement, missing responses, option patterns, item-total relationships and version effects should be monitored. A statistically unusual item requires content and administration review.
Automated flags do not determine deletion; an authorised methodological decision is required.
Weight indicators
Weight distributions, adjustment factors, effective sample sizes and contribution of extreme weights should be examined. Trimming trades variance reduction against potential bias.
The decision and effect on principal estimates should be recorded.
Publication indicators
Every reported percentage should reconcile to its denominator or approved estimation output. Titles, notes and narrative should use the same population and method.
Tables should be tested for suppressed cells, totals, rounding and version consistency.
Quality rating
A concise rating may communicate whether an estimate is fit, fit with qualification, experimental or not releasable for a stated use. The rating criteria should be public and accompanied by reasons.
One overall label should not conceal a critical failure in language, coverage or confidentiality.
Table 20: quality evaluation record
| Dimension | Evidence | Pass condition | Restricted-use condition |
|---|---|---|---|
| relevance | policy question and construct map | indicator answers stated decision | use narrowed to actual construct |
| coverage | target–frame reconciliation | material groups represented | population claim restricted |
| participation | response flow and subgroup profile | no unresolved severe selectivity | residual bias stated |
| measurement | validity and reliability studies | interpretation supported | domain or language limited |
| estimation | weights, variance and replication | principal estimates reproducible | experimental model label |
| comparability | method and population matrix | intended contrast supportable | descriptive or qualified only |
| timeliness | observation and release dates | relevant after quality completion | preliminary status stated |
| clarity | metadata and accessible account | method visible beside result | headline withheld until corrected |
| confidentiality | disclosure review | risk controlled | aggregation or access restriction |
| integrity | approval, version and correction | authorised record complete | release suspended for critical failure |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Release decision
Release should be authorised by the competent statistical body after technical and disclosure review. Policy disagreement with a valid result is not a quality failure.
An unreleasable estimate may still guide restricted internal investigation where confidentiality and uncertainty are controlled.
Correction decision
Correction is required when an error affects values, labels, comparison or interpretation materially. The notice should identify the original, corrected result and policy consequence.
New data or a later method are revisions or new observations, not corrections of the historical record.
Quality conclusion
Quality evaluation transforms a long list of desirable characteristics into a decision about a stated use. It protects against both uncritical publication and indefinite delay in search of impossible certainty.
The governing standard is transparent fitness for public purpose.
Part XIII
Strategic Findings and Conclusions
Finding 1: no single literacy rate is self-explanatory
Every rate embodies decisions about capability, population, language, respondent and threshold. A number detached from these decisions cannot support responsible comparison.
Method metadata should therefore be treated as part of the indicator itself.
Finding 2: methods provide different evidence
Self-report describes a response under a question; proxy report adds another person’s knowledge; schooling inference uses an associated characteristic; direct assessment observes selected performance.
Their differences are analytically useful and should not be erased by a common title.
Finding 3: the binary rate is an incomplete summary
A binary indicator can support communication and broad monitoring but conceals degrees, domains and uncertainty near the threshold. Contemporary literacy evidence supports a continuum and plural practices.[REF-02] [REF-04]
Policy should retain distributional evidence where available.
Finding 4: educational attainment is not literacy
Education contributes strongly to literacy, but attainment levels do not determine present proficiency. School quality, language, practice and time vary.[REF-06] [REF-07]
Attainment and literacy should be analysed jointly without substituting one for the other.
Finding 5: direct assessment has boundaries
Direct assessment improves observation and distributional analysis but remains bounded by tasks, languages, frame and participation. Technical sophistication does not remove social and operational error.
The result should identify adults whose performance was not observed.
Finding 6: participation can bias the estimate
Adults facing the greatest literacy barriers may be hardest to reach or least willing and able to complete an assessment. An achieved-sample rate can therefore be systematically favourable.
Response adjustment helps but does not prove removal of bias.
Finding 7: language defines evidence
Performance in one language is not evidence of literacy in every language. Conversely, inability to use the assessment language is not proof of absence of literacy.
Language mapping, adaptation and reporting are central validity controls.
Finding 8: cross-national ranking is often too strong
Accurately collected national values can remain incomparable because methods, questions and populations differ. A ranked table gives a stronger impression than the evidence warrants.
Comparability classes and separate panels provide a more responsible account.
Finding 9: trends break when methods change
A new question or direct assessment may improve measurement while creating a discontinuity. The resulting numerical difference should not be labelled progress or decline without a bridge.
Historical series gain credibility when breaks are preserved.
Finding 10: uncertainty extends beyond standard error
Sampling uncertainty can be calculated, but coverage, non-response, language and construct limitations also affect conclusions. They require qualitative and sensitivity evidence.
Reporting only a confidence interval can create false completeness.
Finding 11: local need exceeds local precision
Policy often requires district evidence that a national assessment cannot estimate reliably. Census modules, model-assisted estimates and community evidence may help, with distinct status.
Unstable direct ranks should not control material allocation.
Finding 12: programme evidence has a narrower population
Learner assessment can show change among participants, but attrition and selection limit inference. A national indicator cannot serve as a programme comparison by coincidence of timing.
Evaluation requires a design aligned with programme exposure and intended result.
Finding 13: layered systems are preferable
Census breadth, household context, direct assessment depth, programme records and qualitative enquiry are complementary. A layered architecture assigns each an explicit role.
Integration should preserve method identity rather than construct one opaque composite.
Finding 14: statistical independence and participation coexist
Communities and policy bodies should contribute to relevance, language and interpretation. The statistical authority remains responsible for scientific method, confidentiality and impartial release.[REF-09]
Neither technical isolation nor political control produces trusted evidence.
Finding 15: measurement should lead to provision
The public purpose of adult literacy statistics is improved opportunity to learn, use and sustain literacy. A measurement programme that repeatedly documents disparity without informing provision is incomplete.
Policy response should be monitored separately from the indicator.
Immediate priority 1: classify existing indicators
Countries should inventory literacy values by self, proxy, indirect, short direct and multi-domain direct method. Each record should include population, year, language and question or framework.
This inventory can immediately prevent unsupported trend and rank claims.
Immediate priority 2: publish exact metadata
Current releases should place essential method information beside principal values. Historical metadata should be recovered where feasible and uncertainty acknowledged where it cannot be found.
Unknown method is itself a reason to restrict comparison.
Immediate priority 3: analyse participation
Existing surveys should publish contact, cooperation and assessment completion by available group. Language and disability-related non-completion should not remain inside a general missing category.
Field operations should use the findings in the next cycle.
Immediate priority 4: stop automatic schooling substitution
Where education is used to assign literacy, the assumed component should be separately counted and tested in a subsample. No-schooling adults should also receive an opportunity to demonstrate capability.
This protects both statistical validity and dignity.
Immediate priority 5: preserve trend breaks
Statistical publications should mark every material change of method and refrain from subtracting unlike values. Bridge designs should be planned before the old method is discontinued.
Improved measurement should not be discouraged by fear of an apparent setback.
Medium-term priority 1: direct-assessment capacity
Countries requiring proficiency evidence should develop national sampling, language, field, scoring and analysis capacity. LAMP provides a contemporary international direction for this work.[REF-03]
Capacity development should include dissemination and policy use, not collection alone.
Medium-term priority 2: language inclusion
Written-language evidence should guide which versions are developed and how samples support analysis. Exclusions and limitations should be public.
Language inclusion should be planned with communities and specialists from the beginning.
Medium-term priority 3: method-linking studies
Validation and bridge samples should examine self, proxy and direct evidence in the same population. They should test differential relationships by relevant group.
Results should explain disagreement rather than promise a universal conversion.
Medium-term priority 4: recurring household context
Stable household modules can monitor practices, perceived difficulty, language and access between assessment cycles. They should be concise and linked to clear decisions.
Changes in wording or mode require control.
Medium-term priority 5: research access
Protected access to microdata and documentation can strengthen replication and methodological research. Disclosure controls should preserve confidentiality while avoiding unnecessary barriers.
Independent analysis should be answered with evidence, not institutional defensiveness.
Longer-term priority 1: comparable proficiency trends
Repeated assessments should retain common construct coverage and linking evidence while allowing controlled renewal of tasks. Trend should be reported with sampling and linking uncertainty.
Continuity should not preserve obsolete content that no longer represents literacy practices.
Longer-term priority 2: environments for literacy
Measurement should increasingly examine opportunities to use literacy in work, family, public services and civic life.[REF-02] [REF-04]
Policy can then address both adult learning and institutional demands that exclude people through unnecessary complexity.
Longer-term priority 3: evaluation of literacy policy
Population indicators should be integrated with implementation and evaluation evidence so that policy contribution can be judged without confusing association and causation.
Equity, sustainability and unintended effects should remain part of evaluation.
Table 21: strategic action schedule
| Horizon | Priority | Responsible institutions | Evidence of progress |
|---|---|---|---|
| immediate | indicator method inventory | statistical and education authorities | complete metadata register |
| immediate | participation profiles | survey and assessment bodies | staged response tables |
| immediate | trend-break publication | statistical publishers | annotated historical series |
| immediate | schooling-assumption disclosure | census and survey bodies | separate assumed counts |
| medium term | direct-assessment capacity | national statistical and assessment bodies | valid national proficiency study |
| medium term | language inclusion | language communities and assessment body | tested versions and coverage record |
| medium term | bridge studies | statistical and research institutions | method-effect estimates |
| medium term | household context module | survey and policy bodies | stable recurring indicators |
| longer term | comparable proficiency trend | national and international partners | linked repeated assessments |
| longer term | policy evaluation | responsible ministries and independent evaluators | contribution findings and response |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Final conclusion
Adult literacy measurement is a public judgement about what capability is observed, whose experience is represented and which comparison is justified. The apparent simplicity of a national rate conceals choices about language, schooling, respondent, task and threshold. Those choices should be visible.
Self-reported data remain useful for broad coverage, perceived difficulty and literacy practices when exact wording, proxy status and missingness are known. They cannot establish the same result as a direct assessment. Direct assessment provides richer proficiency evidence but must answer for the adults, languages and settings it fails to include.
The responsible course is not to select one method as universally authoritative. It is to construct a layered system, assign each method a defined role, bridge changes where possible and classify comparisons by the strength of equivalence. Statistical uncertainty, participation and dignity belong inside the indicator, not outside it.
The resulting evidence can support the commitments of education for all and the international Literacy Decade only when it leads to learning opportunities and environments in which adults can use literacy. Measurement earns authority through clarity, restraint and public use.[REF-01] [REF-15]
Part XIV
Indicator Specifications for Practical Use
Purpose of the specifications
The specifications below define a limited family of indicators that may coexist without being treated as interchangeable. Each identifies the evidence, denominator and permitted decision. National authorities should adapt language and population details while preserving method identity.
Declared reading-and-writing status
The numerator is adults answering that they can read and write under the exact question; the denominator is all eligible adults, with unknown status separately reported. Self and proxy responses should be distinguished.
Use is limited to declared status and broad coverage analysis. It does not establish proficiency or separate reading from writing where the question combines them.
Known-response declaration rate
This rate divides declared yes by yes plus no, excluding unknown. It is useful for comparison with the all-eligible rate and for examining the numerical effect of missing status.
It should never replace the primary coverage account without evidence that unknown responses are ignorable.
Graded self-assessed reading difficulty
Adults report whether selected reading tasks are performed easily, with difficulty or not at all. Exact materials and reference period should be stated.
This supports service-access and perceived-difficulty analysis. It is not a proficiency scale unless validated for that use.
Graded self-assessed writing difficulty
Writing is measured separately through defined activities such as completing a form or composing a short message. Response categories and assistance should be recorded.
The measure can reveal domain differences hidden by a combined literacy question.
Literacy-practice frequency
The indicator reports how often adults engage with specified texts in work, household, learning or community settings during a defined period. “Never” can reflect absence of opportunity, preference or difficulty.
Practice is relevant to maintenance and policy context but does not demonstrate task performance.
Whole-sentence reading performance
The numerator is adults who read the specified sentence fully under the administration and language rule; the denominator is all eligible adults offered valid access. Partial reading, non-performance and non-completion remain separate.
The claim is restricted to the sentence task and cannot include writing.
Partial-sentence performance
This indicator retains adults able to read part but not all of the sentence under the scoring rule. It prevents a potentially important capability range from disappearing into a binary rate.
Interpretation requires consistent probing and scorer guidance.
Assessment participation rate
Valid assessment completers are divided by all selected eligible adults. Background-only completion, refusal, language exclusion, accessibility exclusion and interruption are component measures.
The rate informs potential bias and should accompany every proficiency result.
Assessment-language coverage
The numerator is target-population adults for whom an authorised assessment language is available; the denominator is the defined target population. Estimates may use frame or survey evidence.
Availability does not prove version equivalence or actual participation.
Mean proficiency score
The weighted mean summarises the score distribution for a covered population under the scaling model. It should be reported with standard error and scale interpretation.
The mean should not stand alone where lower-tail need or distributional inequality drives policy.
Proficiency-level distribution
The indicator reports weighted shares within approved score ranges. Counts, standard errors and level descriptions accompany the distribution.
Levels are reporting categories and should not be converted into fixed identities for adults.
Lower-tail percentile
A lower percentile score can show the position beneath which a stated share of the covered population falls. It avoids some dependence on an arbitrary level boundary.
Precision can be weak in small samples, and task interpretation remains necessary.
Literacy distribution by sex
The indicator reports the selected measure separately for women and men under identical method conditions, with uncertainty. Survey participation and proxy status should also be compared.
The difference describes inequality; explanation requires evidence about education, language, work and social conditions.
Literacy distribution by age
Age groups should be defined consistently and reported with observation date. Results describe cohort and age differences at one time unless longitudinal evidence exists.
They should not be labelled skill loss solely from cross-sectional contrast.
Literacy distribution by educational attainment
Assessment or declaration results are tabulated by harmonised education level. The analysis shows variation within and between attainment categories.[REF-10]
It should demonstrate why education cannot be used as an automatic literacy classification.
Literacy distribution by language
Results may be grouped by home language, language of literacy practice or assessment language. These variables should not be conflated.
Assessment-language groups are self-selected or assigned populations and should not be ranked causally without an adequate design.
Literacy distribution by location
Urban–rural or regional estimates require stable definitions, sample support and appropriate variance. Geographic categories should reflect policy and remain comparable over time.
Local publication should avoid disclosure and unstable rank.
Self–direct agreement rate
The indicator reports the proportion for which a defined self-report category agrees with a defined direct-assessment category in a linked sample. Both categories and the assessment threshold must be stated.
Agreement does not establish that the constructs are identical or that disagreement is dishonesty.
Proxy–self agreement rate
Adults’ own responses are compared with independent proxy responses under a controlled design. Results may vary by relationship and household absence.
The study informs proxy quality but should not expose disagreement within named households.
Method discontinuity estimate
In a bridge sample, the difference between old and new methods is estimated under aligned population and period. Subgroup variation and uncertainty should be shown.
The estimate may support historical interpretation but is not necessarily stable across future cohorts or settings.
Coverage ratio
The estimated frame population is divided by the target population under aligned definitions. Known omissions and duplicates are reported separately.
A high overall ratio can conceal complete omission of a small policy-relevant group.
Contact rate
Selected eligible units successfully contacted are divided by all units known or estimated eligible. Unknown eligibility requires a stated treatment.
The indicator separates frame access from cooperation and assessment performance.
Cooperation rate
Adults agreeing to the relevant interview or assessment are divided by contacted eligible adults. The outcome should distinguish informed refusal, gatekeeper refusal and other non-cooperation.
Pressure to improve the rate must not compromise voluntary participation or respect.
Assessment break-off rate
Adults who begin but do not provide valid completion are divided by assessment starters. Reason and point of break-off should be analysed.
The rate can identify burden, difficulty, language or administration problems but does not alone assign cause.
Weight-adjustment effect
The difference between base-weighted and final-weighted principal estimates shows the effect of adjustment. It should be reported with the variables and weight stages used.
A large effect prompts review; a small effect does not establish absence of non-response bias.
Effective sample size
The effective sample size expresses loss of precision from clustering and unequal weights relative to a simple design. The exact definition should be stated.
It supports decisions about subgroup publication and future design, not a replacement for design-based variance.
Relative standard error
The standard error divided by the estimate provides one indicator of relative precision. It can be unstable near zero and should be interpreted with absolute uncertainty.
Publication rules may use it alongside unweighted count and disclosure controls.
Programme reach
Eligible adults participating in a literacy programme are divided by the defined eligible or intended population. Estimating that denominator may require local evidence.
Reach should be disaggregated and should not be confused with completion or learning.
Programme retention
Participants remaining through a defined instructional point are divided by entrants. Transfers, planned exits and missing records should be classified.
Retention is a participation result and does not establish proficiency gain.
Matched learning change
The difference between comparable baseline and follow-up measures is reported for participants with both observations. The matched count and attrition profile accompany it.
The result applies to matched participants unless an adjustment with defensible assumptions broadens inference.
Public-service text access indicator
Adults may report or demonstrate access to defined essential written information, while services record availability of accessible formats and assistance. The indicator should specify the service event.
It connects literacy policy to institutional responsibility rather than locating every barrier in the adult.
Table 22: indicator-use register
| Indicator family | Evidence class | Primary use | Mandatory companion |
|---|---|---|---|
| declared status | self or proxy report | broad coverage and outreach | respondent and unknown shares |
| graded difficulty | self-assessment | service access and perceived need | exact task wording |
| literacy practice | reported behaviour | context and opportunity | period and text type |
| sentence performance | short direct task | narrow verification | partial and non-completion categories |
| proficiency distribution | multi-item direct assessment | population skill distribution | participation, language and uncertainty |
| subgroup distribution | any stable measure | equity analysis | sample, method and confidentiality |
| method agreement | linked measures | validation and transition | threshold and overlap design |
| field participation | operational records | bias analysis | target and selected populations |
| programme reach | programme and population data | service access | eligibility definition |
| matched change | repeated participant assessment | learning among observed participants | attrition and comparison limits |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Specification conclusion
The indicator family provides complementary evidence without a false common scale. Selection should follow the decision, and every value should retain the companion information necessary to understand its population and method.
Scenario 1: an unexplained high rate
A newly submitted census rate is substantially higher than the previous observation. Before publication, the authority should examine question wording, schooling assumptions, proxy share, unknown treatment, age limits and geographic coverage.
The value should remain pending if the metadata cannot explain its method. Plausibility against expectation is not sufficient verification.
Scenario 2: a lower direct-assessment result
A direct assessment produces a lower share above a selected threshold than the existing self-declaration rate. The authority should publish the two as different measures and explain the constructs.
It should not retrospectively revise the earlier declaration rate or call the direct result a decline.
Scenario 3: missing language version
Field preparation reveals that an intended language version cannot meet validity and staffing requirements by the start date. The authority should consider staged collection, adjusted population scope or delay.
Using another language by convenience and coding non-performance as low skill is not acceptable.
Scenario 4: severe non-response in one region
National response is adequate, but one region has high contact failure and assessment break-off. Additional fieldwork should preserve selected cases, and the regional estimate should be withheld or qualified if selectivity remains.
The national estimate also requires sensitivity analysis according to the region’s population weight.
Scenario 5: item security breach
Evidence shows that some assessment tasks circulated before administration in selected areas. The authority should identify exposure, test performance anomalies and determine whether affected items or cases can support inference.
Replacement or exclusion requires a pre-authorised technical decision and an account of scale comparability.
Scenario 6: scoring disagreement
Monitoring finds low scorer agreement on constructed writing responses. Release of the affected domain should pause while rules, retraining, moderation and rescoring are completed.
Other independent domains may proceed if their validity and publication meaning are unaffected.
Scenario 7: political request to change the threshold
A ministry requests a lower threshold so the reported target is attained. A threshold may be reviewed for substantive policy reasons, but the original result and target remain on their approved basis.
Any new threshold produces a separate indicator and cannot rewrite historical performance.
Scenario 8: unstable district rankings
Local leaders request an ordered table of district rates. Where confidence intervals overlap substantially and sample counts are low, the authority should publish appropriate estimates and uncertainty without an ordinal league table.
Planning categories may be formed from multiple evidence sources under a transparent rule.
Scenario 9: proxy responses dominate
A census review finds that most working-age adult records were supplied by proxies. The data remain census evidence but should be labelled accordingly. A validation subsample can estimate disagreement.
The next operation should consider timing and respondent procedures that increase self-response without reducing coverage.
Scenario 10: programme data exceed population need
Provider records show more annual enrolments than the estimated local adult population below a literacy threshold. The figures may include repeat enrolments, commuters, unlike age ranges or an inappropriate prevalence estimate.
Reconciliation should precede any claim of universal programme reach.
Scenario 11: preliminary results change after weighting
Initial unweighted results differ materially from final weighted estimates. Public communication should use the final design-consistent values and explain that preliminary tabulations did not represent the population.
The difference is a reason for correct weighting, not evidence of manipulation.
Scenario 12: new census question
Cognitive testing supports a clearer graded question, but it differs from the previous binary item. The authority should adopt the better question where warranted and plan an overlap or bridge.
Trend continuity is valuable but should not preserve an inadequate measure indefinitely.
Scenario 13: adult refuses an assessment
The assessor should record informed refusal and conclude contact respectfully. No lowest score is assigned. Available background data may be retained only under the consent and confidentiality arrangements.
Weighting and bias analysis address population inference; coercion does not.
Scenario 14: accommodation changes mode
An adult requires a presentation or response method not covered by the standard design. The responsible specialist should determine whether the construct is preserved and whether the result belongs on the common scale.
If not, the adult’s exclusion from the standard estimate is recorded and a suitable descriptive assessment may still inform provision.
Scenario 15: correction after release
A coding error changes one subgroup estimate and a national total by 0.3 percentage points. Materiality depends on affected conclusions, not size alone. All tables, narrative and policy uses should be traced.
A dated correction should state which findings remain and which change.
Scenario 16: conflicting official values
Two public agencies release different adult literacy rates for the same nominal year. A reconciliation statement should compare observation year, population and method and assign each value an authorised use.
The goal is not necessarily to choose one number, but to prevent incompatible meanings from being presented as a factual dispute.
Decision-scenario conclusion
The scenarios demonstrate that quality governance depends on decisions made before and after fieldwork, not only on calculation. A predetermined authority, evidence requirement and release consequence shorten response and protect the public record.
Synthesis of measurement trade-offs
Adult literacy measurement presents recurring trade-offs that should be decided openly. Breadth can be increased through short census questions, but a shorter instrument offers less evidence about domains and proficiency. Depth can be increased through direct assessment, but time, language development, sampling and participation become more demanding. Frequency can be increased through recurring modules, but repeated design change can weaken the series more than an additional observation strengthens it.
There is also a trade-off between standardisation and relevant access. A common instrument supports comparison, yet literal uniformity can introduce irrelevant difficulty where language, script or document conventions differ. Adaptation improves access only if the intended construct and task demand remain sufficiently equivalent. The decision requires linguistic, substantive and empirical evidence; neither universal sameness nor unrestricted local variation is adequate.
Small-area demand creates another tension. National assessment samples can provide reliable distributions but insufficient local precision. Censuses can provide local declaration data without direct proficiency. Model-assisted estimates can fill part of the geographic gap but depend on assumptions. A responsible system presents these sources as complementary layers and does not grant the strongest methodological label to the estimate with the smallest geographic unit.
Continuity and improvement also compete. Stable wording protects trend, while an inadequate question can perpetuate weak evidence. Method renewal should proceed through testing, overlap and a visible break. A country should not be penalised analytically for improving measurement, but it should not convert the new result into an artificial historical trend.
Finally, measurement ambition must be balanced with respondent dignity and public trust. Longer assessment and repeated contact can add information, but burden and anxiety may selectively reduce participation. Confidentiality encourages cooperation but does not justify secrecy about methods. Accessible communication expands public use but must retain qualifications.
These trade-offs do not imply that every design is equally defensible. They require a reasoned choice linked to purpose, resources and consequence. The acceptable design is the one that preserves the essential population and construct, measures uncertainty, protects participants and states where interpretation must stop.
The standard for a defensible national estimate
A national estimate is defensible when its population can be reconstructed, its evidence method is named, its language conditions are visible, its participation losses are examined and its numerical uncertainty is reported. These conditions are cumulative. A precise calculation from a partial frame does not become national through weighting alone; an inclusive frame does not establish proficiency through an undefined question.
The standard is proportionate to use. Broad programme outreach may proceed from a stable census declaration with known limitations. A claim about the distribution of adult proficiency requires direct evidence and a probability design. A judgement about programme effect requires evidence of exposure and a comparison suited to causation. The same literacy label does not equalise these thresholds.
Defensibility also requires institutional independence and public explanation. Technical decisions should be made by the competent statistical authority, recorded before results are known where feasible, and open to methodological scrutiny. Policy bodies may set priorities and respond to findings, but they should not select the value or threshold that best satisfies an existing commitment.
Where the evidence falls short, the correct response is a restricted claim, a qualified estimate or further measurement. Withholding a precise but unsupported rank is not a failure of statistics. It is evidence that the system distinguishes what is known from what remains uncertain.
Closing public-interest test
Before an adult literacy indicator is adopted, the responsible authority should ask whether the measure will improve understanding of educational need, whether the population most affected can participate, whether language and disability are treated as measurement conditions, and whether the published result can be used without unjustified stigma. A method that is efficient but systematically excludes the intended beneficiaries fails this test.
The authority should also identify the decision that follows. If no institution is responsible for provision, communication or further investigation, repeated measurement may create visibility without remedy. Conversely, an urgent local response may be justified by credible evidence even when a nationally comparable proficiency estimate is not yet available.
The public-interest test therefore joins statistical fitness with institutional use. It does not allow policy urgency to weaken evidence, nor does it allow technical uncertainty to become a reason for inaction. It requires an honest account of what is known, a proportionate decision and a plan to close material evidence gaps.
Final evidentiary boundary
This report supports decisions about the design, interpretation and institutional use of adult literacy indicators. It does not supply a universal conversion between declaration and assessment, a single threshold valid for every policy, or a substitute for national linguistic and population evidence. Application should preserve these boundaries and record any departure required by local conditions.
The central requirement remains stable: a public result should identify whose literacy was considered, how evidence was obtained, which language and task conditions applied, what uncertainty remains and which decision the evidence can reasonably support.
Part XVI
Measurement Conditions across Diverse Settings
Purpose of contextual differentiation
Global comparison should recognise that national statistical capacity, linguistic plurality, settlement patterns and adult-learning systems differ. Common principles can govern definitions and disclosure while operational designs respond to context.
Context should explain design choices; it should not excuse an unsupported population claim.
High census coverage with limited assessment capacity
Where a census is the principal source, priority should be given to exact wording, proxy identification, language rules, unknown status and a small validation study. This can improve interpretation without immediately requiring a large specialised assessment.
The census result remains declared or reported literacy.
Established household-survey systems
Countries with recurring household surveys can add a concise literacy module, rotate practice questions and periodically administer direct tasks to a probability subsample.[REF-11] [REF-13]
Stable core questions protect trend, while controlled modules can address new policy needs.
Strong assessment capacity
Where technical and financial capacity support a multi-domain assessment, the design should still examine frame coverage, language access and selective completion. High psychometric quality within the achieved sample does not settle population representation.
Investment should include analysis and policy use, not fieldwork alone.
Multilingual national settings
Language selection should follow written-language evidence, population size, policy purpose and consequences of exclusion. Several versions may require larger samples and longer development.
A national result can combine valid versions where the scale supports it, while language-specific interpretations remain conditional on population differences.
Low-density rural populations
Sampling remote communities may be expensive but substantively necessary. Oversampling, extended field periods and local-language assessors may be required.
Excluding remote areas and labelling the result national can misdirect precisely the policies concerned with unequal access.
Highly mobile populations
Seasonal work, migration and unstable residence weaken household contact and usual-residence classification. Field timing and repeated contact should reflect known mobility.
The report should distinguish frame omission, temporary absence and movement outside the target territory.
Post-conflict and disrupted settings
Population frames, infrastructure, language relations and trust may be weak. A staged design may begin with area mapping, focused surveys and qualitative evidence before a national estimate is credible.
Measurement should not delay urgent adult-learning provision supported by existing evidence.
Small island and small-population settings
A census or near-census assessment may be feasible, but confidentiality and respondent burden become acute. Sampling error may be small while non-response and disclosure risk remain material.
Regional pooling requires attention to language, population and method equivalence.
Urban informal settlements
Rapid growth and irregular addresses can cause frame undercoverage. Updated area listing, local mapping and varied contact hours may improve inclusion.
Service-based samples can inform barriers but do not replace probability population estimates without a defensible frame.
Populations with limited schooling
Routing adults away from assessment because they lack schooling reproduces an untested assumption. Instruments should include accessible entry tasks that permit capability to be demonstrated.
Task difficulty should extend low enough to describe emerging proficiency without humiliation.
Populations with high formal attainment
High school completion does not remove the need for assessment where policy concerns adult proficiency and practice.[REF-06] [REF-07]
Ceiling coverage should be adequate to distinguish advanced performance, and assumed literacy should not replace observation.
Ageing populations
Upper-age exclusions may remove a growing share of adults and weaken service planning. Assessment burden, sensory access and cohort education require careful treatment.
Age differences should be separated from longitudinal change and reported with the actual age range.
Youth transition to adulthood
Results for ages 15–24 connect school quality, early work and continuing learning. School attendance and assessment language may affect interpretation.
Youth literacy should not be used as a direct substitute for adult rates or school-learning measures.
Gender-restricted participation environments
Interview timing, interviewer sex, privacy and household permission can affect women’s or men’s participation. Field protocols should protect direct informed response and record gatekeeper refusal.
An apparently small sex gap may reflect differential proxy reporting or exclusion.
Disability-inclusive population estimates
Frame and instrument design should include adults with disabilities through accessible contact and administration. Exclusions should be counted by reason and should qualify the population estimate.
Alternative modes belong on the common scale only where construct and measurement evidence support that use.
Limited statistical resources
A smaller probability sample with sound selection, tested questions and complete metadata is preferable to a larger uncontrolled collection. Existing survey infrastructure can reduce cost if its population and field methods are fit.
Technical assistance should transfer reproducible capability and documentation.
Decentralised statistical systems
Regional bodies may collect literacy data under different languages and procedures. A national framework should establish common core definitions, version control and quality evidence while allowing justified local modules.
Central aggregation should occur only after method reconciliation.
Administrative dependence on programme data
Where programme records dominate the evidence base, authorities should distinguish participation, completion and assessed learning from population prevalence. Unserved adults are absent by definition.
A household or community sampling component is required to examine unmet need.
Transition from binary to continuum measures
Countries moving toward direct proficiency assessment should preserve the binary series as a historical declared indicator and conduct an overlap study. The new distribution can answer stronger questions without being forced into the old categories.
Public targets may require formal rebasing.
Table 37: contextual design choices
| Setting condition | Priority design response | Evidence retained | Claim requiring restraint |
|---|---|---|---|
| census is principal source | improve wording, proxy record and validation | local declared status | direct proficiency |
| recurring household surveys | stable module and direct subsample | context and linked method | universal conversion |
| multilingual population | planned versions and language sampling | language-specific access | causal language ranking |
| remote population | oversample and extend fieldwork | territorial coverage | national claim after exclusion |
| mobile population | updated frame and varied contact | absence and movement status | stable-resident assumption |
| small population | intensive coverage and disclosure control | detailed national evidence | unsafe local cells |
| limited schooling | accessible low-demand tasks | emerging proficiency | schooling-based assignment |
| ageing population | inclusive age and accommodation | older-adult distribution | age difference as skill loss |
| limited resources | focused probability design | defensible bounded estimate | large uncontrolled rate |
| decentralised collection | common core and reconciliation | regional method record | aggregation before equivalence |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Contextual decision rule
Contextual adaptation is justified when it improves representation or construct relevance without obscuring method. The authority should record the condition, chosen response, evidence of adequacy and effect on comparison.
An adaptation that changes the observed capability may still be useful for local policy, but it should receive a separate indicator identity.
Contextual conclusion
Internationally credible measurement does not require every country to use one instrument. It requires common discipline in population definition, method identity, language access, uncertainty and public reporting.
These controls allow diversity of setting to enter the evidence without making every national value incomparable by default.
Part XVII
Interpretations That the Evidence Does Not Support
Literacy as a fixed personal identity
No indicator in this report supports treating literacy as an immutable identity. Capabilities can develop, be sustained or weaken according to learning and opportunities for use.[REF-02] [REF-06]
Categories describe evidence under defined conditions and should not determine an adult’s social worth or capacity to learn.
One task as complete capability
Successful reading of a sentence does not establish writing, document use, numeracy or engagement with extended text. Failure does not establish absence of every literacy capability.
The task statement should remain as narrow as the evidence.
Self-report as deliberate misstatement
Disagreement with direct assessment can arise from different constructs, thresholds, language, context and reference standards. It should not be interpreted automatically as deception.
Validation should explain patterns and improve questions rather than assign blame.
School completion as demonstrated proficiency
Formal attainment documents educational participation and completion under a system. It does not directly observe current adult literacy.[REF-06] [REF-10]
Using it as a convenient assumption changes the method and should remain visible.
Group difference as inherent deficiency
Differences by language, sex, location, age or poverty do not establish an inherent characteristic of the group. Historical educational access, opportunity, institutional demand, migration and measurement conditions may contribute.
Interpretation should direct attention to remediable barriers and provision.
Assessment-language score as language quality
Average performance among adults assessed in different languages does not establish that one language is more capable of supporting literacy. Populations, versions and opportunities differ.
Version validity and social explanation are separate analytical questions.
Association as individual destiny
Population associations between literacy and employment, income or health do not determine the outcome of a particular adult. They also do not show that raising one score alone will produce the associated outcome.
Public communication should avoid deterministic language.
Statistical significance as policy importance
A precise small difference may be statistically detectable and educationally minor. A substantively important difference may be estimated imprecisely in a small group.
Policy judgement should consider magnitude, uncertainty, affected population, equity and consequence together.
Absence of significance as equivalence
Failure to reject a statistical null does not demonstrate that two groups or periods are equivalent. The sample may lack power or the interval may include important differences.
Equivalence requires a defined margin and suitable design where that claim is needed.
National average as universal experience
A national mean or rate can conceal a lower tail, local exclusion and language disparity. It cannot describe every adult or community.
Distributional and participation evidence should accompany aggregate policy use.
Better measurement as declining literacy
Introduction of direct assessment, improved coverage or removal of schooling assumptions may lower a reported value. The difference may reflect correction of the evidence base rather than deterioration.
The series should mark the method change and avoid assigning it to population time.
Higher response as absence of bias
A high response rate reduces but does not eliminate non-response bias. A small non-responding group can be highly distinctive, and frame undercoverage exists before response is measured.
Response level, pattern and relationship to literacy all matter.
Weighting as full representation
Weights align respondents with known population variables under assumptions. They cannot create observations for an omitted language, inaccessible mode or absent frame group.
Residual bias should remain part of the conclusion.
Proficiency level as individual diagnosis
Population assessments may provide limited precision for one adult, especially under item-sampling designs. A level estimate should not be reused for individual placement, employment or benefit decisions without separate validity and due process.
Population monitoring and individual assessment are different purposes.
Programme completion as learning
Completion establishes participation through a defined point. It does not establish that intended capability changed or was sustained.
Learning evidence, assessment coverage and attrition remain necessary.
Population change as programme impact
A national indicator can change through cohorts, migration, schooling, social conditions and multiple programmes. Coincidence with one intervention does not establish attribution.
Programme evaluation requires evidence closer to exposure and a credible alternative account.
Comparability label as permission for every use
A pair of values may be comparable for broad distribution and not for precise ranking, subgroup analysis or causal inference. The authorised use must accompany the classification.
Comparability is always tied to the question.
International standardisation as uniform language
Common concepts and methods do not require the same language or culturally irrelevant documents. Adaptation can be necessary for equivalent access and task demand.
The evidence standard is disciplined equivalence, not identical appearance.
Technical detail as public inaccessibility
Complex survey and assessment methods require full technical records, but principal public meaning can still be communicated clearly. Method, population and uncertainty should not disappear from accessible accounts.
Clarity and rigour are compatible responsibilities.
Uncertainty as a reason for no action
Policy often proceeds under bounded uncertainty. Credible evidence of severe exclusion may justify immediate provision while stronger measurement is developed.
The action should be proportionate, its evidence basis explicit and its result monitored.
Final interpretive safeguard
The recurring safeguard is to keep the public claim at the level of the evidence. When a result concerns declared status, it should say so; when it concerns performance on selected tasks, it should name them; when the population is restricted, it should identify the restriction.
This precision is the foundation of internationally credible adult literacy reporting.
Implications for the next measurement cycle
The next cycle should begin from the weaknesses identified in the current evidence rather than from a general ambition to collect more data. Where proxy response is extensive, field timing and respondent selection should be improved. Where language coverage is incomplete, version development and sampling should begin early. Where direct assessment completion is selective, burden, access and trust require operational attention before another estimate is commissioned.
Continuity should be protected through stable core definitions and planned bridge evidence. Renewal remains necessary where a binary question, schooling assumption or narrow task no longer answers the policy question. The old and new measures can coexist during transition, with each assigned a clear status and use.
Institutional arrangements also require continuity. Sampling, language, scoring, analysis and public communication should not depend on knowledge held by temporary personnel. Controlled documentation, national technical teams and transparent review preserve capacity between cycles.
The next measurement cycle will be successful only if its evidence is used. Authorities should state which findings led to changes in adult-learning provision, accessible public communication, language materials or outreach, and which questions remain unresolved. This closes the relationship between statistical observation and the public commitment to literacy.
Results from the next cycle should be compared with this evidence only after the population, construct, language and method have been reconciled. A more recent value is not necessarily a more comparable value. The integrity of the series depends on preserving the observation that actually occurred and explaining the relationship between successive designs.
This discipline permits measurement to improve without losing institutional memory.
It also ensures that policy progress is judged against evidence of the same meaning, not against a numerical resemblance created by incompatible methods. Where equivalence is absent, the responsible conclusion is a documented break and a new baseline. That conclusion should remain visible in every subsequent public comparison using the series.
Literacy indicator metadata dictionary
Each indicator receives a unique identifier, title, version, responsible body, approval date, observation period and release status. Superseded versions remain available. The identifier follows the value into every table and analytical extract.
The record states reading, writing, numeracy or other domain; declared, practised or demonstrated capability; text and context; response or task demand; classification; and permitted interpretation. An undefined entry is not replaced by the general word literacy.
Age, upper-age limit, usual-residence rule, territory, household or institutional scope and material exclusions are specified. The target, frame, selected, responding and represented populations receive separate counts or estimates.
Method is classified as self-declaration, proxy declaration, indirect assignment, short direct verification or multi-item assessment. Mixed indicators identify each component and routing rule. Question text, item framework and administration mode are linked.
The dictionary records languages permitted by the concept, languages offered, language selected, translation version, script and interpreter or accommodation rules. The population unable to access an authorised version remains visible.
Sampling design, base weight, non-response adjustment, calibration, imputation, variance method, rounding and suppression rules are stated. Modelled estimates identify covariates, model version and validation.
Coverage, contact, cooperation, completion, reliability, validity, standard error, non-sampling limitations and comparability class are recorded. The quality statement includes the authorised use and any restricted use.
| Metadata group | Required fields | Release test | Failure response |
|---|---|---|---|
| identity | identifier, title, version and owner | one controlled definition | resolve duplicate or conflicting record |
| construct | domain, evidence and interpretation | claim matches observed condition | narrow title and use |
| population | age, residence, coverage and exclusions | inference population reproducible | restrict population statement |
| method | question, respondent, task and routing | method class explicit | withhold method-neutral rate |
| language | concept, offered versions and access | language condition visible | qualify or redesign |
| estimation | weights, variance and missing-data rules | estimate reproducible | correct calculation |
| quality | response, uncertainty and limitations | fitness for use stated | assign restricted status |
| history | observation, release and change dates | trend breaks traceable | restore version record |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Every change states old and new field, reason, evidence, effective date and effect on comparison. Editorial correction that does not alter meaning is distinguished from methodological change. Historical outputs retain the version under which they were released.
Essential population, method, language, period and uncertainty fields should accompany the public value. More detailed records may be provided separately, but complexity cannot excuse omission of the indicator’s basic meaning.
Sampling, response and weighting protocol
The sampling plan states the target population and quantifies frame coverage from the most reliable demographic evidence. Areas, institutions or persons omitted by design are listed with estimated size where possible. Duplicates and out-of-date units receive correction procedures.
Strata, stages, primary units, households and adult selection rules are documented with probabilities. Interviewer discretion in selection is prohibited. Reserve or replacement units are used only under an approved probability design.
Every selected unit receives a final disposition: ineligible, unknown eligibility, not located, no contact, temporary absence, refusal, language barrier, accessibility barrier, partial, complete or other defined status. Attempts, dates and modes are retained.
The background interview and direct assessment have separate outcomes. Assessment disposition identifies not offered, refused, unable to access language, accommodation unavailable, started, interrupted, partially scoreable and valid completion.
Contact, cooperation, background completion and valid-assessment rates are calculated from documented denominators. Unknown eligibility receives a stated estimation rule. Rates are presented overall and for planned strata and relevant groups.
The base weight is the inverse final probability of selecting the adult. Non-response adjustment uses variables observed for both respondents and non-respondents. Calibration aligns with reliable controls under consistent population definitions.
The distribution, minimum, maximum, percentiles, design effect and effective sample size are examined. Extreme weights prompt source review before trimming. The effect of every adjustment stage on principal estimates is recorded.
| Control stage | Required evidence | Diagnostic | Decision |
|---|---|---|---|
| frame | coverage and exclusion map | target–frame ratio by group | supplement or restrict scope |
| selection | probabilities at every stage | weights reproduce selected counts | correct selection record |
| contact | attempt and disposition history | patterned non-contact | extend or vary fieldwork |
| cooperation | informed response outcome | refusal by group or interviewer | improve information and supervision |
| assessment | start, break-off and valid completion | literacy-related loss | analyse bias and redesign |
| base weight | inverse probability | extreme or missing values | reconcile sample selection |
| adjustment | class or model variables | large estimate movement | examine assumptions |
| calibration | population controls | inconsistent definitions | replace control or restrict use |
| variance | strata, cluster and weight method | implausible precision | correct design estimation |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Where material bias is plausible, the authority should use field observations, frame variables, shortened follow-up or linked administrative evidence under proper authority. Follow-up participation is itself selective and should not be treated as a complete correction.
The final record reconciles selected adults through valid completion and weighted population. Unexplained count loss, unidentified selection probabilities or a critical unrepresented group prevents the intended whole-population release.
Language adaptation and accessibility protocol
The authority defines whether the assessment concerns any written language, selected national languages or a specified language. Population evidence, policy purpose, resources and consequences guide version selection. Exclusion is documented as a limitation, not treated as absence of literacy.
Each version uses translators with command of the source and target language, literacy and measurement specialists, and reviewers familiar with written use in the target population. Conflicts and decisions are recorded.
Before translation, the team identifies the capability, textual feature, cognitive operation and difficulty driver in each task. Elements essential to comparability are separated from surface features that may be adapted.
Independent drafts are reconciled through evidence and recorded decisions. The review examines vocabulary, syntax, script, layout, document convention, number format, names and contextual familiarity. Back translation is supporting evidence only.
Cognitive interviews and field trials include adults across relevant proficiency, age and language-practice groups. Review examines interpretation of instructions, engagement with text, response process, time and distress.
Item difficulty, discrimination, missingness and differential functioning are examined across versions with adequate samples. A flagged item receives linguistic and substantive review before retention, modification or exclusion.
The protocol defines accessible consent, communication, presentation and response methods. Each adjustment is assessed against the construct. Standard-scale inclusion and descriptive alternative assessment are distinguished.
| Stage | Evidence | Acceptance condition | Unresolved consequence |
|---|---|---|---|
| language selection | population and policy map | scope serves intended population | restrict claim or add version |
| construct analysis | task-demand specification | essential demand identified | task not translated |
| translation | independent drafts and reconciliation | meaning and demand preserved | revise wording |
| contextual adaptation | documented substitution | no irrelevant advantage | test or exclude item |
| cognitive testing | participant response evidence | intended process observed | redesign task |
| field trial | timing, missingness and performance | operational and measurement adequacy | delay main fieldwork |
| statistical analysis | cross-version item evidence | no material unexplained difference | qualify, rescale or remove |
| accessibility | construct-preservation review | barrier removed without changed demand | separate result or exclusion record |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Reports state offered languages, choice or assignment rule, version sample sizes, exclusions and known equivalence limits. Language-group results are not interpreted as effects of language without an appropriate causal design.
Later orthographic, terminology or layout changes receive controlled version identifiers and testing proportionate to their likely effect. A change that alters task demand creates a comparison break unless linking evidence supports continuity.
Self-report and proxy module protocol
The module should identify whether it supports broad declared status, perceived difficulty, literacy practices, service access or validation. Questions not connected to an analysis or policy decision should be removed.
Self-response is preferred for personal capability and practice. Proxy response may be accepted for coverage under a defined rule and is flagged at item level. The proxy’s relationship and reason for substitution are recorded.
Questions specify activity, material, language where relevant, difficulty and reference period. Reading and writing are separated where domain-specific evidence is needed. Response categories include unknown or unable to judge for proxies.
Testing examines how adults understand literacy, difficulty categories, “simple” material, language and social consequence. It includes adults with limited schooling and varied language backgrounds.
Interviewers use exact wording and neutral clarification, protect privacy and do not infer answers from education or occupation. Assistance provided to understand the interview is recorded where it affects the question.
Attendance, highest level and completion are collected under standard education classifications. They remain explanatory variables unless a published indicator explicitly and transparently uses an assumption rule.[REF-10]
Reading and writing frequency is collected across work, household, learning and community contexts with locally relevant examples. Non-use is followed by a reason where burden permits.
| Field | Response design | Quality flag | Permitted use |
|---|---|---|---|
| respondent status | self or identified proxy | proxy knowledge unknown | method profile |
| reading declaration | yes, no, unknown under exact wording | combined with writing | broad declared status |
| writing declaration | separate status | proxy uncertainty | domain description |
| reading difficulty | graded defined tasks | unstated reference standard | service need |
| writing difficulty | graded defined tasks | assistance not recorded | service need |
| language | languages read and written | spoken language substituted | provision planning |
| practices | frequency by text and setting | opportunity confused with skill | context and maintenance |
| education | standard level and completion | used as automatic literacy | association only |
| access to learning | availability, participation and barrier | enrolment treated as proficiency | programme planning |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
A probability subsample can complete direct tasks. Analysis compares self and proxy responses with assessment under the same population and period, reporting disagreement and uncertainty by relevant group.
The publication reproduces exact core wording and respondent shares. Declared status, perceived difficulty and reported practice are labelled separately. No result is described as demonstrated proficiency unless a direct task supports it.
Direct-assessment administration and scoring protocol
The assessor confirms selected adult, private setting, authorised language, accessible information, materials and version. Another household member does not replace the selected adult or answer assessment tasks.
The adult receives purpose, expected duration, confidentiality, voluntary status where applicable and contact information in an understandable form. The assessor explains that tasks are not a judgement of personal worth and that results are reported statistically.
Practice items establish understanding of response procedures without teaching assessed content. Difficulty with instructions prompts authorised clarification; inability to access the mode triggers the accommodation protocol.
Instructions, order, timing where relevant and permissible prompts are defined. Assessors do not translate spontaneously, explain vocabulary in assessed text or signal correctness. Interruptions and deviations are recorded.
The adult may pause or discontinue. The assessor records the point and reason without assigning unattempted items as incorrect unless the scoring framework expressly and validly requires it.
Objective responses follow controlled keys. Constructed responses use criteria, examples, double scoring and adjudication at a risk-based rate. Scorers are monitored for drift and unusual severity.
Optical or manual capture includes version and disposition controls. Critical fields receive verification. Edits retain original value, rule, authorisation and date.
| Stage | Required control | Deviation record | Decision consequence |
|---|---|---|---|
| identity and selection | selected eligible adult confirmed | wrong or substituted adult | invalidate case |
| language | authorised version and rule | unauthorised translation | review affected tasks |
| setting | privacy and adequate conditions | interruption or third-party help | flag administration |
| instructions | standard delivery | additional substantive explanation | validity review |
| access | approved accommodation | unavailable or altered construct | separate or exclude with count |
| completion | start, break and final status | reason for missing tasks | apply scoring rule and bias analysis |
| scoring | controlled key and agreement | scorer discrepancy | moderate or rescore |
| capture | version and response verification | missing or impossible code | correct from source |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Supervisors observe sessions, recontact a protected sample for procedural verification and examine duration, completion and result patterns by assessor. Investigation separates assignment differences from performance concerns.
The adult receives appropriate closing information without an improvised individual proficiency judgement. Materials and identifiers are secured, and any welfare or access concern is referred only under the disclosed procedure.
Estimation, variance and sensitivity protocol
Every calculation begins with the population quantity sought: proportion, mean, percentile, distribution, difference or association. The eligible population, domain and observation period are fixed before the estimator is selected.
For a binary or category indicator, the weighted proportion is the sum of final weights for adults in the category divided by the sum of final weights for all adults in the defined denominator. Unknown and inapplicable records follow the approved rule.
The weighted mean is the sum of each final weight multiplied by the relevant score, divided by the sum of weights for adults with valid score evidence. The scored population and assessment non-completion are reported.
Variance estimation reflects stratification, clustering and weighting through an approved linearisation or replication method. The technical record identifies strata, primary units, replicate weights and any certainty selections.
Differences between groups or periods require covariance where samples or linking are related. A significance test should not replace consideration of educational magnitude and method comparability.
The interval communicates sampling uncertainty under the design and estimator. It does not include every coverage or measurement error. This boundary should accompany principal intervals.
Predefined analyses may vary missing-data treatment, weight trimming, flagged items, proficiency threshold and uncertain population eligibility. Results are compared for magnitude and decision stability.
| Estimate | Required inputs | Independent check | Limitation retained |
|---|---|---|---|
| declared rate | category, denominator and final weight | component totals | proxy and missing status |
| proficiency mean | scores or population estimates and weights | replicate calculation | assessment participation |
| level share | cut points, values and weights | shares total within rounding | boundary uncertainty |
| subgroup difference | comparable estimates and covariance | reversed subtraction and interval | non-causal relationship |
| trend difference | linked construct and populations | bridge and method-break check | observation interval |
| percentile | weighted score distribution | alternative computation | tail precision |
| model-assisted local value | direct data, covariates and model | validation and residuals | predicted status |
| sensitivity range | defensible alternative assumptions | documented reruns | not a probability interval |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Unrounded calculations are retained; public rounding is consistent across narrative and tables. Percentages reconcile within stated rounding. Counts derived from weighted estimates are labelled estimates rather than achieved sample counts.
The statistician confirms arithmetic and design, while the responsible analyst confirms that narrative claims match population, method and uncertainty. Both are required because a correct calculation can support an incorrect interpretation.
Comparability and trend protocol
The analyst records the exact proposed comparison and decision: magnitude, rank, distribution, trend or association. This prevents a general declaration of comparability from being extended to a stronger use.
For each value, the record contains construct, population, age, coverage, period, method, question or framework, respondent, language, assumption, sampling, response and uncertainty.
Each dimension is classified equivalent, bounded difference, material difference or unknown. Unknown is not treated as equivalent. The analyst explains whether bounded differences could reverse the proposed conclusion.
Common items, parallel methods, overlap samples or statistical linking may strengthen comparison. The design, population and uncertainty of the link are recorded. Correlation alone does not establish interchangeable scale units.
The final comparison is strong, qualified, descriptive only or not supportable. A strong classification requires no material unresolved difference for the intended use. Qualified comparison states the exact limitation.
The series lists observation date, indicator version and breaks. Changes in question, response, schooling assumption, language, mode, frame, assessment or scaling receive markers and notes.
Strong and qualified results may appear in analytical comparisons with appropriate notes. Descriptive-only values are juxtaposed without subtraction or rank. Unsupported comparisons should not be displayed in a way that invites the prohibited inference.
| Review outcome | Evidence condition | Permitted presentation | Prohibited presentation |
|---|---|---|---|
| strong | construct, population and method aligned | magnitude, distribution and trend with uncertainty | causal claim without design |
| qualified | bounded known differences | broad contrast with adjacent qualification | precise rank or small change claim |
| descriptive only | related but materially different measures | separate panels and contextual account | subtraction or common scale |
| not supportable | fundamental or unknown difference | method description only | comparative conclusion |
| linked break | overlap evidence estimates discontinuity | marked series and bridge result | silent splicing |
| unlinked break | method changed without bridge | separate series | progress or decline across break |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
The technical comparison decision should precede policy interpretation. Disagreement is recorded and resolved by the competent statistical authority under published principles, not by selecting a convenient source.
New methods may clarify earlier differences but do not alter the contemporaneous status of historical observations. Reanalysis should be dated and should preserve the original publication and method.
Publication, confidentiality and correction protocol
The release inventory lists report, tables, metadata, methodological documentation, accessible summary and authorised data products. Each carries document number, version, date and observation cutoff.
Every principal claim is traced to a table or cited source. The reviewer tests population, method, time, modality and causal language. Qualifications appear beside the proposition they limit.
Headers state measure and population; notes state method, language, observation year and uncertainty. Rows and columns reconcile, suppressed cells remain protected and totals are not misleading after suppression.
Direct identifiers are removed from public outputs. Small cells, rare languages, geography and combined characteristics are assessed for re-identification and community harm. Restricted analytical detail remains available only to authorised users.
Public summaries use clear language, defined terms, readable tables and suitable oral or translated forms. They preserve denominators, uncertainty and method distinctions. Adults represented by the findings should be able to understand the public meaning.
Principal results are released under an announced procedure providing equal public access. Pre-release access, where authorised for operational preparation, should be limited, recorded and unable to alter results.
A reported error is logged, assessed for materiality, corrected across every affected format and announced with date and consequence. The correction authority acts independently of whether the revised result is favourable.
| Assurance | Evidence | Release condition | Correction if failed |
|---|---|---|---|
| document identity | version, date and observation period | one approved release | withdraw conflicting edition |
| claim alignment | claim-to-table and source trace | wording within evidence | revise claim |
| numerical parity | independent calculations and format comparison | all values agree | correct all representations |
| method visibility | population, method and language notes | essential meaning beside value | restore metadata |
| uncertainty | design and non-sampling account | precision not overstated | add interval and limitation |
| confidentiality | disclosure-risk review | risk controlled | suppress, combine or restrict |
| accessibility | clear and appropriate public forms | qualifications preserved | revise communication |
| equal access | release log | authorised schedule followed | disclose and correct process |
| correction | issue, decision and change notice | material error traceable | issue revised controlled version |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
The archive retains source definitions, instruments, translations, field manuals, scoring, weights, calculations, review and released versions under protection. Preservation enables later interpretation without disclosing respondent information.
The competent authority signs the release only when statistical, interpretive and confidentiality reviews pass. Policy bodies may respond to the findings after release, but their response should remain distinct from the statistical account.
Fieldwork monitoring and incident protocol
Field supervisors receive selected, contacted, completed and outstanding counts by stratum, interviewer and authorised language. Reports distinguish background completion from valid assessment. Rates are used to identify operational problems, not to impose coercive quotas.
Duration, refusal, proxy use, missing items, break-off and score distributions are examined against assignment. Outliers trigger review of workload, geography, language and practice before conduct is inferred.
Incidents include wrong respondent, unauthorised substitution, language mismatch, loss of materials, confidentiality exposure, item exposure, unsafe setting, substantive third-party assistance, system failure and participant distress. Severity and immediate protection are recorded.
The supervisor protects the adult and information, stops affected work where necessary, secures materials and reports through the defined route. A serious confidentiality or safety incident is not deferred to routine quality review.
The technical authority determines whether an incident affects the whole case, selected tasks or only operational metadata. The original responses and decision record are retained. Field personnel do not delete a case to improve completion figures.
Action may include retraining, observation, reassignment, repeat contact with consent, replacement materials, added language capacity or suspension. Re-interview is authorised only where it does not create undue pressure or compromised item exposure.
| Incident | Immediate control | Technical decision | Population implication |
|---|---|---|---|
| wrong adult | stop session and secure record | invalidate or restart authorised selection | selection integrity |
| language mismatch | do not score as low performance | reschedule or record inaccessible | language undercoverage |
| third-party help | record tasks and circumstances | flag or invalidate affected evidence | possible score bias |
| item exposure | secure material and identify extent | item or area analysis | comparability and security |
| confidentiality loss | protect and notify authority | incident response and access review | trust and participation risk |
| break-off cluster | examine burden and administration | revise field support | selective completion |
| system failure | preserve source and version | recover under controlled rule | missingness and delay |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
At fieldwork close, every selected unit has a disposition and every incident has a decision or continuing owner. The incident profile is included in the quality report and informs the next design.
Proficiency-level description protocol
Level descriptions translate score ranges into the kinds of tasks that adults at different positions on the scale are likely to perform. They support interpretation and should not become labels of intelligence, worth or fixed capacity.
Descriptions use tasks located within the scale, their content and the probability criterion adopted by the assessment. The procedure and criterion are documented. A few released examples should illustrate rather than define the entire level.
Specialists examine text form, operation, inference, competing information, length and context across tasks. The description identifies increasing demand without claiming that every task within a range shares one feature.
Adults near a cut point have similar estimated performance despite different level labels. Reports should state this and provide standard errors for level shares. “At Level 1” means classified within the approved range under the model.
The lowest category should describe the limited evidence available rather than state that adults have no literacy. Assessment may contain too few very easy items to characterise performance precisely below the scale.
The top category is also limited by task coverage. It should not be described as mastery of all literacy demands. Sparse samples at the upper tail require suitable precision notes.
| Element | Required basis | Wording control | Misuse prevented |
|---|---|---|---|
| score range | approved scale cut points | numerical range stated | hidden threshold |
| task demand | multiple located items | likely performance under criterion | single item defines level |
| boundary | measurement uncertainty | adjacent adults may be similar | categorical discontinuity |
| lowest category | available easy-task evidence | below assessed threshold | no literacy identity |
| highest category | upper task coverage | performance on demanding assessed tasks | universal mastery |
| population share | weights and variance | estimate with standard error | exact population count |
| policy use | defined task relevance | selected-level rationale | universal deficit threshold |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Level descriptions change only through controlled review of scale and task evidence. Editorial simplification should not broaden the claim. A changed cut point creates a new indicator version and comparison record.
Programme-learning measurement protocol
Programme assessment begins with the learning objective and intended participants. It should identify reading, writing, numeracy, language or applied practice expected to change and the instructional period.
All entrants are counted with eligibility, start date and baseline status. Adults unable to complete the standard baseline remain in the participation record and receive an accessible route where valid.
Tasks should reflect intended learning without reproducing only practised examples. A national population assessment may provide a reference framework but is not automatically sensitive to programme progress.
Completion, withdrawal, transfer, absence and assessment non-response are separated. Matched change is reported with the number and characteristics lost to follow-up.
Where policy requires an effect estimate, the comparison should represent what would plausibly have occurred without the programme. Assignment, matching or phased implementation may assist, subject to ethics and design.
Observed change may reflect instruction, practice, maturation, selection and measurement. The conclusion should distinguish change among completers, change among entrants after adjustment and causal contribution.
| Evidence stage | Required count or measure | Interpretation | Limitation |
|---|---|---|---|
| eligible population | defined intended adults | potential reach | denominator may be estimated |
| entrants | enrolled and started | actual reach | self-selection |
| valid baseline | comparable entry evidence | starting distribution | baseline non-completion |
| exposure | attendance, duration and instructional content | treatment received | quantity not quality alone |
| valid follow-up | comparable later evidence | observed follow-up distribution | selective retention |
| matched change | same adults and scale | change among matched adults | attrition |
| comparison result | aligned non-participant or phased group | evidence on contribution | residual confounding |
| sustained result | later comparable evidence | maintenance | later loss to follow-up |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Providers should publish reach, retention, assessment coverage and learning findings without presenting participant results as community prevalence. Negative and mixed findings remain part of programme accountability.
National measurement cycle schedule
The responsible authorities approve policy questions, governance, target population, component roles, languages, budget and release intention. The observation date and principal comparisons are fixed.
The statistical and assessment bodies prepare frames, instruments, translations, sampling, protection and analysis. Cognitive testing, pilot and field trial have distinct decision gates.
Readiness confirms valid instruments, language versions, trained field staff, verified selection, functioning capture, incident response, scoring and secure logistics. Critical failure delays or stages fieldwork.
Fieldwork follows varied contact schedules and continuous quality monitoring. Design changes during collection require central authority and a record of affected cases and comparability.
The sequence includes disposition reconciliation, capture verification, scoring, edit, weighting, variance, item and language review, disclosure and replication. Provisional results remain restricted until the relevant gate.
Technical review, public tables, accessible communication and policy briefing are prepared from the same approved estimates. Observation and release dates are distinguished.
Education and other responsible bodies issue action responses. Programme design, language provision and resource decisions refer to the relevant indicator and population. Statistical findings are not rewritten as commitments.
| Stage | Controlled product | Decision gate | Principal risk |
|---|---|---|---|
| initiation | policy and measurement brief | purpose and governance approved | instrument without decision use |
| construct | framework and indicator map | interpretations approved | broad claim from narrow evidence |
| language | tested versions and coverage | access adequate for scope | excluded communities |
| sample | frame and selection plan | population inference supported | unrepresented target group |
| field trial | operational and measurement evidence | main collection ready | unresolved task or burden |
| collection | dispositions, responses and incidents | coverage and participation acceptable | selective completion |
| processing | scores, weights and replicated estimates | numerical quality passed | hidden processing error |
| interpretation | comparability and limitation record | claims supportable | false rank or trend |
| release | controlled statistical publication | confidentiality and parity passed | inconsistent public values |
| policy response | actions, resources and evaluation plan | responsible authority acts | measurement without provision |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
Stable household questions and programme indicators may continue between direct assessments. They should not be used to estimate unobserved proficiency annually. Their role is context, access and early warning.
The next cycle begins with an evaluation of participation, language, burden, method use and policy consequence. Change is introduced through testing and linking so that improvement in relevance does not silently destroy comparability.
Indicator reconciliation and source-status protocol
Where several official literacy values exist, reconciliation establishes why they differ and which use each can support. It does not force unlike methods into one preferred number. The process should begin before values enter a common table or public target.
For every value, the record identifies producing body, publication, observation date, release date, population, question or assessment, respondent, language, assumption, weighting, uncertainty and revision status. Secondary reproductions are traced to the primary source.
A value is classified as observed declaration, observed direct task, indirect classification, mixed-method estimate, modelled estimate or projection. Preliminary, final, corrected and superseded status are also recorded.
Age ranges are recalculated to a common range where source data permit. Territory, household status and exclusions are compared. If harmonisation requires removing a material population, both original and harmonised scopes remain visible.
Exact questions, routing and tasks are compared. Shared words such as read, write, simple and understand are not treated as equivalent without operational evidence. Proxy and schooling assumptions are quantified.
Published numerators, denominators, weights and rounding are reproduced. Where only a rounded rate exists, an exact population count should not be inferred. Revisions are tied to their source notice.
| Reconciliation field | Evidence | Compatible outcome | Incompatible outcome |
|---|---|---|---|
| source status | primary release and revision history | same controlled version | preliminary mixed with final |
| observation date | fieldwork period | aligned or policy-relevant interval | unknown or distant year labelled current |
| population | age, territory and residence | common recalculated scope | unresolved material difference |
| construct | domain and interpretation | same defined capability | declaration equated with proficiency |
| respondent | self, proxy or performed task | comparable response process | proxy share unknown and material |
| language | permitted and administered forms | equivalent access | one-language result treated as any-language |
| estimation | weight, model and uncertainty | reproducible common statistic | modelled value called observation |
| trend | stable version or bridge | linked series | silent method break |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
If two values remain incompatible, the public reconciliation note states each source and use. One may support local declared need and another national proficiency distribution. Both can be valid within their boundaries.
An error correction changes the affected historical record through a dated notice. A new method does not correct the earlier value; it creates a new series or a linked series where bridge evidence exists.
Values submitted for international monitoring should include the full method class and observation year. If the requested indicator does not match the national source, the difference should be declared rather than concealed through relabelling.
A national target identifies the indicator version against which progress will be judged. If measurement improves, the target may require formal rebasing, with old and new evidence shown. Performance is not recalculated by simply applying the old numerical threshold to a different method.
The statistical authority approves the source-status map and permitted comparisons. Policy bodies acknowledge the evidence basis of targets and actions. The completed record allows later users to understand why more than one official literacy value may properly exist.
An aggregate should use observed or approved estimated values, a population weight consistent with the indicator and a rule for missing countries. The covered population share and observation-year range should accompany the result. A regional value should not be produced when missing populations are large enough to make the result materially unrepresentative under the approved rule.
Mixed-method aggregation requires particular restraint. Weighting self-declaration and direct-assessment rates into one regional number does not make their constructs equivalent. Where the monitoring purpose requires both, results should be grouped by method or presented as a qualified series with the limitation prominent.
Female, male and total values should reconcile under the population weights and categories used. A total derived from one source should not be combined with sex-specific values from another observation without disclosure. Unknown sex records and non-binary national categories, where collected, require a stated treatment consistent with the source and protection requirements.
Gender disparity measures should preserve direction, denominator and method. A difference in percentage points is not the same as a ratio, and neither alone explains the institutional barriers producing inequality.
The reconciliation is closed only when each value has an authorised status, comparisons have classifications and unresolved conflicts are disclosed. Closure does not require numerical agreement. It requires that no user be invited to treat unlike evidence as one observation.
The final map should be reviewed whenever a correction, new survey or method change enters the series. This continuing custody is essential to the integrity of long-term adult literacy monitoring.
Where publication space is limited, the minimum statement should still identify the observation year, covered population, evidence method, respondent rule, language scope and whether the value is observed, assumed or modelled. It should state any material break from the preceding value and provide the location of the full source record.
A short statement must not compress incompatible evidence into an apparently continuous rate. If comparison is descriptive only, that classification should appear with the figures. If comparison is not supportable, the values should be separated and no difference calculated.
The responsible statistical officer should confirm that the abbreviated statement preserves the conclusion of the full reconciliation. Brevity is acceptable only when it does not change the public meaning.
The statement should also retain the principal uncertainty measure and the share of the intended population not represented. Neither element may be omitted merely because the result is presented in a summary table. A comparison without these controls should be returned for qualification before publication or policy use.
Public table specification and interpretive notes
Every table has a number, substantive title, indicator version, observation period and covered population. The title states whether results are declared, assessed, modelled or programme-based.
Column headers identify count, weighted estimate, percentage, score, standard error or interval. Units are not left to surrounding prose. Age, sex, language and geography categories retain their source definitions.
The first note states evidence method, respondent rule, language scope and material assumption. A mixed-method value identifies its components. Direct-assessment tables state the domains and target population.
Sample estimates carry standard errors or confidence intervals and a marker for estimates failing the approved precision rule. Unweighted sample count is provided where it assists assessment of stability and does not create disclosure risk.
Excluded territories, institutions, ages and language groups are stated. Assessment completion and any residual non-response limitation appear with proficiency tables.
Trend and cross-national tables state comparability class. Method breaks are marked at the relevant value. Descriptive-only values are not joined by lines, differences or ranking.
| Table element | Mandatory content | Interpretation protected |
|---|---|---|
| number and title | subject, method and population | table cannot be detached from meaning |
| observation field | reference date or fieldwork period | release year not mistaken for data year |
| value header | unit and indicator version | rate, score and count not conflated |
| population note | age, residence and exclusions | scope not overstated |
| respondent note | self, proxy or assessed adult | declaration not called direct evidence |
| language note | permitted and administered languages | language-limited result not universalised |
| uncertainty | standard error and non-sampling limitation | point estimate not treated as exact |
| comparability | strong, qualified, descriptive or unsupported | rank and trend constrained |
| source | primary producing body and publication | secondary value traceable |
Source and methodological notes are stated immediately below the table in the authoritative Markdown text.
The text identifies the principal distribution, material disparity and uncertainty without repeating every cell. It distinguishes finding from explanation. Any causal interpretation cites evidence beyond the table.
Layout, font, contrast and reading order should support access. A concise explanation may use ordinary language and examples while preserving denominator, method and uncertainty. Symbols and abbreviations are defined.
A correction updates the table, structured data and every narrative claim using it. The notice identifies the cells changed and whether the substantive conclusion is altered. The superseded table remains archived under control.
Summary release statement
The summary release should state in one place the observation period, population, method, languages, achieved participation, principal estimate, sampling uncertainty and material non-sampling limitation. It should identify whether comparison with the preceding value is strong, qualified, descriptive only or unsupported.
The statement should describe the educational meaning of the result without calling adults below a threshold incapable or treating a group difference as inherent. It should identify the policy question and the responsible public response.
The primary statistical source, methodological documentation and accessible public explanation should be identified. A corrected release should point to the correction notice and retain the original observation date.
The statistical authority approves the numerical and methodological account. The responsible education authority may provide a separate policy response. Keeping the two functions distinct protects both impartial evidence and accountable action.
The summary statement should retain any limitation capable of changing the principal interpretation, even where the full technical account is available elsewhere. Coverage, language and method breaks are substantive conditions and should not be reduced to an unreferenced footnote.
Approval confirms that the shortened account remains faithful to the complete evidence record.
It also confirms that every cited value can be traced to an authorised source and that no unsupported comparative or causal conclusion has been introduced through abbreviation.
The approved statement remains part of the permanent public statistical record.
References
- REF-01
World Education Forum. The Dakar Framework for Action: Education for All — Meeting Our Collective Commitments. 2000.
The global commitment to halve adult illiteracy, improve equitable access to continuing education and monitor measurable progress.
https://unesdoc.unesco.org/ark:/48223/pf0000121147 - REF-02
UNESCO. Education for All Global Monitoring Report 2006: Literacy for Life. 2005.
The principal contemporary synthesis of literacy concepts, global estimates, measurement limitations, policy conditions and adult literacy priorities.
https://unesdoc.unesco.org/ark:/48223/pf0000141639 - REF-03
UNESCO Institute for Statistics. Literacy Assessment and Monitoring Programme (LAMP): Information Brochure. 2006.
The contemporary development of direct literacy assessment, background information, proficiency distributions and improved national measurement capacity.
https://uis.unesco.org/sites/default/files/documents/literacy-assessment-and-monitoring-programme-lamp-information-brochure-en.pdf - REF-04
UNESCO. The Plurality of Literacy and Its Implications for Policies and Programmes. 2004.
The treatment of literacy as plural, socially situated and related to language, purpose, culture and practice.
https://unesdoc.unesco.org/ark:/48223/pf0000136246 - REF-05
UNESCO Institute for Statistics. Aspects of Literacy Assessment: Topics and Issues from the UNESCO Expert Meeting, 10–12 June 2003. 2005.
Technical and conceptual issues in literacy assessment, including construct, language, context, administration and interpretation.
https://uis.unesco.org/sites/default/files/documents/aspects-of-literacy-assessment-topics-and-issues-from-the-unesco-expert-meeting-2005-en_0.pdf - REF-06
OECD and Statistics Canada. Literacy in the Information Age: Final Report of the International Adult Literacy Survey. 2000.
Comparative direct assessment of prose, document and quantitative literacy and the distribution of adult proficiency.
https://www.oecd.org/en/publications/literacy-in-the-information-age_9789264181762-en.html - REF-07
OECD and Statistics Canada. Learning a Living: First Results of the Adult Literacy and Life Skills Survey. 2005.
The contemporary extension of adult literacy assessment to prose, document, numeracy and problem-solving domains, with background and participation analysis.
https://www.oecd.org/en/publications/learning-a-living_9789264010390-en.html - REF-08
United Nations Statistics Division. Principles and Recommendations for Population and Housing Censuses, Revision 1. 1998.
The contemporaneous census framework for population characteristics, literacy enquiry, concepts, classifications and census operations.
https://unstats.un.org/unsd/demographic-social/standards-and-methods/ - REF-09
United Nations Statistical Commission. Fundamental Principles of Official Statistics. 1994.
Professional independence, scientific methods, transparency, confidentiality and public access in official statistics.
https://unstats.un.org/unsd/dnss/gp/fundprinciples.aspx - REF-10
UNESCO. International Standard Classification of Education: ISCED 1997. 1997.
Comparability controls for educational attainment and the separation of educational level from directly observed literacy proficiency.
https://uis.unesco.org/sites/default/files/documents/international-standard-classification-of-education-1997-en_0.pdf - REF-11
World Bank. Designing Household Survey Questionnaires for Developing Countries: Lessons from 15 Years of the Living Standards Measurement Study. 2000.
Questionnaire design, respondent selection, field operations, non-response, household reporting and data-quality controls.
https://documents.worldbank.org/en/publication/documents-reports/documentdetail/452741468778781879/ - REF-12
United Nations Department of Economic and Social Affairs, Statistics Division. Household Sample Surveys in Developing and Transition Countries. 2005.
Sampling, coverage, field implementation, non-response, weighting, error and dissemination in household surveys.
https://unstats.un.org/unsd/hhsurveys/pdf/Household_surveys.pdf - REF-13
UNESCO Institute for Statistics. Guide to the Analysis and Use of Household Survey and Census Education Data. 2004.
Analysis of education variables from household and census sources, including definitions, denominators, disaggregation and comparability.
https://uis.unesco.org/sites/default/files/documents/guide-to-the-analysis-and-use-of-household-survey-and-census-education-data-en_0.pdf - REF-14
United Nations. The Millennium Development Goals Report 2006. 2006.
The contemporary global development-monitoring context, including education and gender disparities relevant to literacy participation and reporting.
https://unstats.un.org/unsd/mdg/resources/static/products/progress2006/mdgreport2006.pdf - REF-15
United Nations General Assembly. International Plan of Action for the United Nations Literacy Decade. 2002.
The international policy context for literacy as a continuum, strengthened monitoring, partnerships and attention to diverse learning needs.
https://undocs.org/A/57/218