KTP/
github ↗
intervention A · A — Measurement

Measurement Reform.

Stop measuring what's cheap. Start measuring what's load-bearing.

Measurement Reform · v1.0Pre-political · 0 jurisdictions
Primary author: Chris Perkins · Last updated: 2026-05-06
Trust target: integrityThe core trust problem in K-12 measurement is integrity — the system claims to measure learning but measures compliance with accountability metrics. Restoring integrity means making the measurement honest about what it captures.
§01Public card
Fight
Stop measuring what's cheap. Start measuring what's load-bearing.
Who Wins
Students, teachers, and parents who can see what learning actually accumulated, not what the test rewards.
Who Loses
Pearson, ETS, College Board ($2B+ testing industry); accountability-hawk politicians whose careers run on a single test score; ed-tech vendors selling test-prep adaptive software.
What You Do Monday
200-word talking point for state board comment. 3-page model resolution: Recommend the Board direct staff to develop oral defense and original synthesis assessments under controlled conditions. 1-page objection responder: 'You're hiding achievement gaps' — the response is BETTER measurement, not less.
Status
Vehicle: Pre-political · 0 jurisdictions adopting · scoreboard: A — Measurement
§02Narrative
Lived experience

A third-grade teacher spends six weeks of the year on test prep. She knows the test measures something narrower than what her students need. She does it anyway. Her evaluation depends on it.

The system

State accountability systems built around standardized tests select for a narrow band of cognitive performance: rapid recognition under time pressure, multiple-choice elimination, surface-form fluency. The tests are cheap to score, defensible in court, and produce a single number that travels well politically. Every other capacity — argument, original synthesis, oral defense, primary-source reading — is unmeasured and therefore, in budget terms, undefended.

The failure mode

What the system can measure becomes what teachers teach. What teachers teach becomes what students learn. What students learn becomes what they can do as adults. The narrow measure becomes the wide reality. AI now does most of what these tests measure better than the median student. The capacities the tests do not measure — interpretation, defense, judgment — are the only ones that remain load-bearing.

The leverage point

Meadows ranks goals of the system above information flows above rules. The leverage point is not 'better tests'; it is changing the goal of measurement from 'sortable number' to 'evidence the student can defend their own thinking.' Under that goal, the tests change shape.

The intervention

State boards of education direct staff to develop oral defense, Socratic dialogue, and original synthesis assessments under controlled conditions. These are conducted by trained observers, not auto-scored. They are slower, more expensive per student, and politically harder to defend in court — and they measure what AI cannot replace.

The strongest objection

You're hiding achievement gaps. Standardized tests are how civil-rights organizations surface inequity. Removing the measurement is the favorite move of districts that don't want to be accountable for their poorest students.

Why the intervention survives

The reform demands BETTER measurement, not less. Oral defense and original synthesis can be normed and reported by demographic subgroup just as standardized tests can. Achievement gaps will still be visible — and they will be visible on capacities that matter, not capacities that have been automated away. The civil-rights commitment is preserved; the instrument changes.

§03Adversaries
  • Pearson PLC — $2.1B revenue from K-12 testing and adaptive products
  • Educational Testing Service (ETS) — administers SAT, AP, GRE; prepares state assessments
  • College Board — $1.6B nonprofit revenue from SAT and AP; vertical integration into AP curriculum
  • Accountability-hawk legislators (both parties) whose policy identity is built on 'measurable results'
  • Civil-rights organizations rightly worried that removing measurement will hide gaps (NAACP LDF, Education Trust, Civil Rights Project at UCLA)
§04Attack surface
  • 'You're hiding achievement gaps.'
  • 'This is unmeasurable, therefore unaccountable.'
  • 'Subjective grading is biased grading.'
  • 'Districts will game oral defense the way they game everything else.'
  • 'Federal ESEA accountability requires standardized measures.'
§05Failure modes (capture / watered-down / weaponized)
Captured
Pearson and ETS develop 'AI-scored oral defense' products. The new measure becomes a more elaborate version of the old one, with the same vendors. The pipeline is preserved; the substrate is not.
Watered-down
States adopt 'portfolio' assessments that are reduced to checklists scored by clerks. The form survives; the function does not.
Weaponized
Districts under pressure use oral defense to remove students from accountability cohorts (claiming subjective adjustments). Achievement gaps are hidden under the new measure.
§06Falsifier

If three or more states adopt oral defense and original synthesis assessments and, after five years, demographic achievement gaps on those assessments are not narrower than gaps on the prior standardized tests — and student capacity to defend an argument under cross-examination has not measurably risen — the intervention has failed in the form deployed. Revision required.

§07Coalition map
Allies
  • FairTest (National Center for Fair & Open Testing)30-year position against high-stakes standardized testing
  • Knowledge Matters CampaignContent-rich curriculum advocacy; aligned on measure-the-substrate framing
Wobbly
  • American Federation of TeachersAnti-test reform but cautious about oral-defense workload
  • National Education AssociationSame posture; depends on local affiliate
Opponents
  • Pearson PLC, ETS, College BoardDirect revenue threat
Should support, doesn't
  • Civil-rights organizations focused on gap-surfacingThe intervention preserves their core commitment; the framing has not yet reached them
§08Equity vector & exclusion risk
Equity vector

Oral defense and original synthesis assessments require trained observer time and small-group settings. Without paired investment in school staffing (see Intervention D), the assessments will be high-quality in well-resourced districts and degraded in under-resourced ones — reproducing the gap the reform is meant to surface.

Exclusion risk

Tier-blindness. Wealthy districts already do oral defense (it is the structure of private-school humanities programs). Reform without provisioning hardens the existing tier structure. PAIR WITH INTERVENTION D.

§09Mechanism of luxury

Oral defense is a luxury good today: private schools, classical schools, and elite public magnets do it routinely; comprehensive public schools do not. The mechanism is staffing density (one observer per six students versus one teacher per twenty-eight) and credentialing (humanities-degreed staff who can run Socratic dialogue).

§10Constitutional & substrate anchors
Constitutional anchor

State constitutions establish public education as a state responsibility (every state but Hawaii has explicit education clauses). State boards of education have rulemaking authority over assessment systems under most state administrative procedure acts. ESEA Title I federal accountability provisions require state assessments but do not specify multiple-choice format.

Substrate precondition

Trained observers (humanities-degreed school staff). Small-group settings. Time on the school day for assessment. Norming process to surface demographic differences without hiding them. None of these exist at scale in comprehensive public schools today.

Amendability

State boards review assessment frameworks on 5-7 year cycles. Amendments can be proposed by board members, the state superintendent, or via formal stakeholder petition. Federal alignment changes require ESEA reauthorization (typically 7-10 year cycles).

§11Impact
First-order impact

Teachers reclaim 4-6 weeks per year previously spent on test prep. Students develop sustained interpretive practice. Assessment data carries more signal about what students can actually do.

Second-order impact

Curriculum shifts to support the new assessment. Teacher training programs add Socratic-method coursework. Parents who used to track test scores start asking about defense scores, which carry richer information. The political vocabulary of accountability shifts.

§12Vehicle & timing
Vehicle

State board of education regulatory action (faster than legislation; 1-3 year cycles versus 5-10). Pair with district-level pilots through model school board resolution.

Preemption risk

Low. Federal ESEA does not specify assessment format; states retain authority. Risk is not preemption but federal accountability score-card incentives that reward the old measure.

Timing window

2027-2029. The next ESEA reauthorization opens federal flexibility on assessment format. State board cycles reset 2027-2030 in 28 states. AI-driven test obsolescence accelerates the political case.

§13Implementation status
(none yet)
Phase 1 launch — no jurisdictions have adopted this framing yet. Track here as resolutions and bills move.
Pending
§14Citations
  1. Koretz, D. (2008). Measuring Up: What Educational Testing Really Tells Us. Harvard University Press.
  2. Ravitch, D. (2010). The Death and Life of the Great American School System. Basic Books.
  3. Meadows, D. (1999). Leverage Points: Places to Intervene in a System. Sustainability Institute. donellameadows.org/archives/leverage-points-places-to-intervene-in-a-system
  4. U.S. Department of Education (2015). Every Student Succeeds Act (ESEA reauthorization), Title I assessment requirements. Public Law 114-95. www.congress.gov/bill/114th-congress/senate-bill/1177
§15Change log
v1.02026-05-06Initial publication. Phase 1 launch. Tier-blindness flagged; pair with Intervention D.
§16Corrections

Corrections via /advocacy/contribute. Public response SLA: 14 days for substantive corrections, 30 days for elaborations. No comments section by design.

Submit a correction →
§17Enforcement architecture
Who has authority

State Departments of Education (assessment policy compliance); State Boards of Education (standard-setting authority); AERA/APA/NCME (voluntary standards, no legal authority).

Trigger condition

Formal complaint to state DOE citing specific misuse of diagnostic data for accountability purposes; or annual district compliance certification showing violation of protected-diagnostic-space requirements.

Recourse on failure

State DOE can condition discretionary funding on compliance; federal Title I monitoring can flag measurement-integrity violations. No private right of action currently exists in most states.

Known gap

No federal private right of action for measurement-integrity violations. Enforcement is administratively discretionary, creating highly variable implementation across districts.

§19Commons design principles
Ostrom principles instantiated
  • Principle 1 (Clearly defined boundaries): Protected diagnostic space is defined by use restriction — data collected for instruction improvement cannot be repurposed for accountability.
  • Principle 3 (Collective choice arrangements): AERA/APA/NCME standards are set through a collective process; the Metric Passport requires that standard-setting committees exclude commercial testing vendors.
  • Principle 6 (Conflict resolution mechanisms): Correction surface on each intervention page is the public dispute mechanism; amendment threshold (10+ submissions) is the collective-choice trigger.
Standing to challenge

Any teacher, student, family, or district official can file a complaint through the state DOE's assessment office. The correction surface on this page is the public-facing channel.

Anti-enclosure clause

Diagnostic data protected under this framework cannot be used for commercial profiling, vendor product development, or any purpose other than improving instruction for the assessed student. Violation converts the data from protected diagnostic to impermissible commercial use, triggering state data privacy statute remedies.

§20Repair protocol
Competence violation

When a district misuses diagnostic data due to misunderstanding (not bad faith): apology from district leadership, remediation plan showing new data-governance training, third-party audit of data use within 60 days.

Integrity violation

When a district knowingly repurposes diagnostic data for accountability after committing not to: formal state DOE complaint, independent audit of all assessment practices, public disclosure of the violation and remediation timeline. Apology alone is insufficient — requires demonstrated dispositional change in data governance.

How to cite
APA 7th ed.

Perkins, C. (2026, May 6). Measurement Reform. Kinetic Trust Protocol Advocacy. https://kinetic-trust-protocol.net/advocacy/k12/measurement-reform
Plain text

Chris Perkins, "Measurement Reform," KTP Advocacy, https://kinetic-trust-protocol.net/advocacy/k12/measurement-reform (accessed May 6, 2026).
reference
pending — source not yet published
RFC tree on GitHub ↗
created 24 August 2026 · last modified 24 August 2026
cross-cutting framing

For the cross-cutting framing — the four-step pedagogy, the three required components, the five-tier separation, and the twelve-field Metric Passport pattern as they apply across all domains — see /advocacy/kpis. This page is the K-12 instantiation: the canonical implementation of the pattern named there.

The substrate-level reform agenda

Fix the map before fighting over the route.

What follows is the long form of the intervention above. Its ambition is structural and pre-political: a measurement constitution that makes the data legible to constituencies that do not yet agree on curriculum, governance, funding formulas, unions, vouchers, charters, discipline, or accountability. The claim is that the first-order reform of public education is not any of those. It is the epistemic infrastructure beneath them.

§01Definition and frame

Substrate-level, pre-political measurement reform means reforming the shared measurement infrastructure beneath education politics — the constructs, instruments, audit standards, and use licenses — before arguing over ideology, governance, funding formula, curriculum, unions, vouchers, charters, discipline, or accountability. None of those second-order fights can be settled honestly when the underlying numbers do not measure what the public, the schools, and the researchers think they measure.

Pre-political does not mean value-free. It means politically prior. The agenda is built so that factions who disagree about what schools should do can still trust the measurement of what schools are doing, and act on that measurement before agreeing on next moves. It is the prerequisite for a real fight, not a substitute for one.

Public education needs a measurement constitution before it needs another reform war.

§02The core problem — measurement becomes governance

The structural failure of US K-12 measurement is that the same number gets used for diagnosis, formative feedback, public reporting, accountability sanctions, and research — all from one instrument, all at once. The same signal becomes both microscope and hammer. A microscope is meant to make something visible. A hammer is meant to land a consequence. When the same instrument is asked to do both, the microscope warps to protect against the hammer, and the hammer lands on whatever the microscope can still see.

The Every Student Succeeds Act preserved the high-stakes annual testing architecture inherited from No Child Left Behind even as it loosened federal sanctions. The collapse of measurement tiers into a single accountability-flavored signal predates ESSA and survived it. RAND's monograph-length evaluation of test-based accountability documents the consequences in detail: narrowed curriculum, score inflation, gaming of student classifications, and a measurable gap between scores on consequential tests and scores on lower-stakes audit tests measuring the same construct.

Sources
§03Campbell's Law — the first law of education measurement reform

"The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor."

— Donald T. Campbell, 1976

Campbell's Law is the first design constraint of any honest measurement reform. It is not a warning to be acknowledged in a footnote and then ignored; it is the reason the architecture exists. A measurement reform that does not begin with Campbell's Law is not a reform — it is a new generation of the same trap. The Institute of Education Sciences has documented how publicly reporting safe-schools data corrupts the reporting itself, in exactly the way Campbell predicted half a century ago.

Sources
§04The five-tier separation

A defensible measurement system has at least five distinct tiers, each with its own purpose, audience, stakes, and validity-evidence requirements. Current US K-12 measurement collapses them into one scoreboard.

  1. Diagnostic. Identify what an individual student does and does not yet know, so a teacher can adjust instruction. Audience: Teacher, student, parent. Stakes: Low-stakes by construction. Results are formative input, never used for school grades or staff personnel actions.
  2. Formative. Track progress within a unit or term to inform pacing, grouping, and re-teaching decisions. Audience: Classroom teacher, instructional coach, department lead. Stakes: Internal to the school. No public reporting role.
  3. Public reporting. Give parents, journalists, and the public a legible, comparable picture of how schools are doing across populations. Audience: Parents, journalists, taxpayers, researchers. Stakes: Reputational. Should be lower-stakes than accountability — a thermometer, not a sentence.
  4. Accountability. Trigger consequential action — interventions, restructuring, sanctions — when defined floors are crossed. Audience: State education agency, federal compliance officers, courts. Stakes: High. Must use validity evidence specific to the consequential decision being made.
  5. Research. Build cumulative evidence about what works for whom, under what conditions, and why. Audience: Researchers, evaluators, policy analysts. Stakes: None for individual students or schools. Findings inform future policy design.
The collapse

Current systems run a single annual summative state assessment and route it into all five uses — diagnostic, formative, public reporting, accountability, and research. Each tier's specific gaming pressures arrive together; each tier's specific validity-evidence requirements get ignored together. The resulting number is asked to be everything to everyone, and ends up being trustworthy to no one.

Tier
Specific failure mode under collapse
Diagnostic
Diagnostic results are pulled into accountability pipelines, so teachers stop trusting them and instruction drifts toward the high-stakes test rather than the underlying construct.
Formative
Vendors market formative platforms whose item banks are also feed-forward into state summative tests, collapsing formative and accountability into one revenue stream.
Public reporting
The reported number is the same number used to close schools or fire principals, so the reported number gets gamed, audited away, or quietly changed in definition.
Accountability
Accountability inherits whatever instrument is cheapest, which is usually the public-reporting test. The same signal becomes both microscope and hammer.
Research
Research access is gated by accountability politics, so researchers cannot get the data they need without agreeing to use it for accountability purposes that compromise the research.
§05The seven-construct map

What is a school actually doing? Seven construct families cover the legitimate territory: learning achievement, learning growth, opportunity to learn, instructional quality, school climate and safety, system capacity, and post-school outcomes. A defensible public picture surfaces all seven, with confidence intervals, rather than collapsing them into a single composite score that conceals more than it reveals.

Construct family
What it tries to observe
Why current systems struggle
Learning achievement
The level of skill or knowledge a student demonstrates at a moment in time, against an explicit standard.
Single end-of-year tests are point estimates with non-trivial measurement error, especially at the individual level. Confidence intervals are rarely shown to readers, and a 4-point swing inside the error band is treated as a real change.
Learning growth
The change in a student's skill or knowledge over time, ideally with the student as their own control.
Growth models depend on vertical scales, two reliable test points, and stable populations. Mobility, opt-outs, and instrument changes between years all contaminate the signal.
Opportunity to learn
Whether students were exposed to the content, instruction time, materials, and qualified teachers needed to learn what is being tested.
OTL data is collected unevenly — vacancies and curriculum access are reported by the same agency that benefits from underreporting them. The measurement of inputs lags far behind the measurement of outputs.
Instructional quality
What actually happens in classrooms: task complexity, feedback, discussion, materials in use.
Direct observation is expensive and labor-intensive. Proxies (like classroom video samples or principal walkthroughs) are scarce; instead the system substitutes student test scores as a proxy for teacher quality, which is psychometrically unjustified.
School climate and safety
Whether students experience the school as physically safe, socially supported, and engaging.
When the same survey instrument is tied to school accountability or safety reporting, schools learn to coach the responses. IES has documented how publicly reporting safety data warps reporting itself.
System capacity
Teacher retention, leadership stability, facility condition, counselor and special-education staffing levels.
Capacity data is reported by districts to the same agencies funding them, with no independent audit, and is rarely surfaced alongside outcome data — so the public sees test scores without the production conditions that generated them.
Post-school outcomes
What graduates do — college enrollment, persistence, labor-market earnings, civic participation.
Linkages between K-12 records, postsecondary enrollment, and earnings data are partial and inconsistent across states. Privacy regimes and data-sharing agreements are not standardized.
§06The Metric Passport

Every public education metric ships with a structured passport: twelve fields documenting what the metric measures, what it is licensed for, what it is forbidden for, and how it can be gamed. The passport is published on the same surface as the metric. A number without its passport is not a measurement — it is a rumor with a citation.

The passport pattern follows the AERA/APA/NCME Standards for Educational and Psychological Testing, which require that validity claims be specific to the use being made of the score. AERA's position on high-stakes testing is unambiguous on this point: a single test should not be used as the sole determinant of consequential decisions about individuals, schools, or systems.

Sources
Sample passport · Chronic absenteeism rate (school level, K-12)

Chronic absenteeism is the rare metric that already has strong validity evidence as a leading indicator of school disengagement, that maps cleanly onto a recognizable construct, and that nonetheless gets gamed in predictable ways the moment it acquires consequences. It is a fair test for the Metric Passport pattern.

Construct
Habitual school engagement, operationalized as the proportion of enrolled students missing 10% or more of school days.
Instrument
Administrative attendance records aggregated to the school year, drawn from district student information systems, with a 10%-of-enrolled-days threshold.
Population
All students enrolled for at least 10 cumulative days in the school year. Excludes pre-K and homebound special-education placements. Includes part-time and continuously-enrolled.
Reliability
Year-to-year stability r ≈ 0.65 at school level (administrative replication studies). Individual-day attendance is a count rather than an estimate, so the reliability question is procedural — whether the count is correctly recorded — rather than statistical.
Validity evidence
Predictive validity established: chronic absenteeism, particularly in 9th grade, is among the strongest non-cognitive predictors of high-school non-completion (Allensworth & Easton, University of Chicago Consortium; Balfanz & Byrnes, Johns Hopkins). Construct validity grounded in school-engagement literature; convergent with course-failure and grade-retention indicators.
Use license
Public reporting: yes, at school level with confidence intervals. Research: yes, with appropriate de-identification. Accountability sanctions on individual schools: only when paired with opportunity-to-learn data and audited absence-cause coding (illness vs unstable housing vs safety vs disengagement). Diagnostic and formative use: yes, at student level for early-warning systems.
Forbidden uses
Not licensed for: ranking individual teachers; triggering automatic truancy fines on families; calculating staff or principal bonuses; school-closure decisions in isolation from the paired OTL panel.
Known gaming risks
Schools may (a) reclassify chronically absent students as withdrawn, depressing the denominator-numerator math in their favor; (b) mark partial-day attendance as full-day; (c) selectively excuse absences in tested grades; (d) push enrollment-eligibility decisions to game the 10-day threshold; (e) under-report in years immediately preceding consequential reporting cycles.
Equity risks
Chronic absenteeism correlates with housing instability, untreated asthma, lack of safe transit, family caregiving burdens, and immigration enforcement fear. Attaching consequences to schools without adjusting for these production conditions penalizes schools serving the highest-need communities for serving them. Required audit: paired reporting with the opportunity-to-learn panel.
Update cadence
Refreshed annually after the school-year close. Methodology reviewed every three years by an external psychometric panel. Revision triggered by (a) reliability drop below 0.55 at school level; (b) published evidence of systematic gaming above an established threshold; (c) any change in the underlying enrollment-counting rules.
Uncertainty display
School-level rate reported with 95% confidence interval. Small-school suppression at n < 20. Year-over-year change reported as a flag only when the confidence interval excludes zero. Public dashboard renders the confidence band visually beside the point estimate.
Audit trail
Computed by the state education agency from district student information system extracts using published calculation rules. Source code for the aggregation pipeline is open. Random-sample audit of district attendance records by the state inspector general on a two-year cycle.
Passport template · the 12 required fields

Every public metric ships with these twelve fields documented. Empty fields fail the publication standard. The fields exist so a reader, a journalist, or an opposing party can see what the metric is, what it is licensed to do, and how it can be gamed.

Construct
The latent thing the metric is trying to observe, named in plain language, mapped to one of the seven construct families.
e.g. Habitual school engagement (subset of school climate / engagement).
Instrument
The actual measurement procedure — the survey, test, log, or count that produces the number.
e.g. Administrative attendance records aggregated to school year, with a missed-day threshold of 10% of enrolled days.
Population
Who is included, who is excluded, and how denominators are defined.
e.g. All students enrolled for at least 10 cumulative days in the school year. Excludes pre-K and homebound special-education placements.
Reliability
How stable the measurement is across reasonable variations (raters, occasions, items). Reported as a coefficient with its method.
e.g. Year-to-year stability r ≈ 0.65 at school level (administrative replication). Individual-day attendance is a count, not an estimate, so reliability is a definitional rather than statistical question.
Validity evidence
Documented evidence that the metric reflects the construct, not just an artifact of the instrument. Per AERA/APA/NCME Standards.
e.g. Predictive validity: 9th-grade chronic absenteeism is among the strongest non-cognitive predictors of high-school non-completion (Allensworth & Easton; Balfanz & Byrnes).
Use license
What this metric is licensed to be used for. Each tier (diagnostic / formative / public reporting / accountability / research) is a separate license.
e.g. Public reporting: yes. Research: yes. Accountability sanctions on individual schools: only with paired opportunity-to-learn data and audited absence-cause coding.
Forbidden uses
What this metric is explicitly NOT licensed for, with reasoning. This is the anti-Goodhart move encoded into the metric itself.
e.g. Not licensed for: ranking individual teachers; triggering automatic truancy fines on families; calculating staff bonuses.
Known gaming risks
How the metric can be manipulated without changing the underlying construct. Published with the metric, not in a separate audit document.
e.g. Schools may reclassify chronically absent students as withdrawn; mark partial-day attendance as full-day; selectively excuse absences for tested vs untested grades.
Equity risks
How this metric, even when correctly computed, can compound existing inequities — and what audit is required.
e.g. Chronic absenteeism is correlated with housing instability, asthma, and unsafe transit corridors. Attaching consequences to schools without adjusting for these production conditions punishes schools for serving harder-hit communities.
Update cadence
How often the metric is refreshed, how often the instrument or methodology is reviewed, and what triggers a methodological revision.
e.g. Refreshed annually. Methodology reviewed every three years. Revision triggered by reliability drop below 0.55, or by published evidence of systematic gaming above a threshold.
Uncertainty display
How error and confidence are surfaced to readers. Not a footnote; on the same panel as the point estimate.
e.g. School-level rate reported with 95% confidence interval; small-school suppression at n < 20; year-over-year change reported only when the interval excludes zero.
Audit trail
Who computed the metric, from what data, with what code, and how a third party could replicate the calculation.
e.g. Computed by state education agency from district student information system extracts. Calculation rules published. Random-sample audit of district records by the inspector general every other year.
§07NAEP — yardstick, not weapon

The National Assessment of Educational Progress is the closest thing the United States has to a working model of disciplined measurement. It is sampled rather than universal, group-level rather than individual, and lower-stakes by design. No student, teacher, or school is sanctioned on the basis of a NAEP score. The result is that NAEP can do what state summative tests cannot: report what students actually know without being warped by the consequences attached to the report.

NAEP is a thermometer. State summative tests are increasingly asked to be a sentencing instrument. A defensible measurement architecture keeps the thermometer separate from the sentence, and uses the thermometer for the public's macroscopic understanding of system performance.

Sources
§08Opportunity-to-learn data — symmetric input measurement

A school is a learning production system. Asking only what its outputs are without measuring what it is being given to work with is not measurement; it is moral theater. Opportunity-to- learn data brings input measurement to the same resolution as output measurement.

The question shifts. Instead of asking why are scores low at this school? a defensible system asks what learning production system generated these outcomes? and reports both panels side by side.

The fourteen input categories
  • Teacher vacancies and long-term substitute deployment
  • Class sizes by grade and subject
  • Instructional minutes per subject area
  • Curriculum access (which materials the school actually has)
  • Advanced course access (AP, IB, dual enrollment, advanced math sequences)
  • Special-education service delivery against IEP minutes owed
  • English-learner support staffing and program model fidelity
  • Counselor-to-student ratios
  • Facility condition (HVAC, water quality, seismic, lead exposure)
  • Connectivity (broadband at school, broadband at home)
  • School safety conditions (real, not survey-coached)
  • Chronic absenteeism drivers (housing instability, transit, health access)
  • Leadership turnover (principal years-of-service)
  • Teacher retention (year-over-year, by school)
§09Anti-Goodhart — per-metric gaming analysis

For every public metric, the gaming model is published with the metric. Not in an audit document a few researchers will read. Not in a footnote. On the same surface as the number. Each metric, the predictable gaming behavior it produces under consequential pressure, and the structural countermeasure required to make the metric usable.

Metric
Likely gaming behavior
Countermeasure
Standardized test scores (state assessment, summative)
Narrowing the curriculum to tested subjects; coaching to item formats; selectively classifying lower-scoring students into special-education or English-learner exemption pools; teaching to released items; concentrating instructional time in tested grades.
Sampled audit testing on a NAEP-style design; matrix-sampling so no student takes the full assessment; rotating item banks; an opportunity-to-learn panel reported alongside scores; cross-validation against curriculum-embedded performance tasks.
Graduation rate (4-year cohort)
Reclassifying near-failing students as transfers, withdrawn, or homeschool; awarding credit for courses below standard; pushing struggling students to alternative-school tracks excluded from the cohort; loosening course-completion definitions in 11th and 12th grade.
Independent audit of withdrawal and transfer codes; require receiving-school confirmation for all transfers out; track and publish a parallel rate that includes alternative-school destinations; spot-audit course-completion records.
Suspension and exclusionary-discipline rate
Reclassifying suspensions as in-school assignments not coded as suspensions; sending students home informally without paperwork ('go home for the day'); pushing repeat offenders into expulsion or alternative placement to clear the count; underreporting incidents.
Audit of disciplinary actions against parent communications and attendance records; require coded reporting of all out-of-class time exceeding a threshold; cross-reference suspension data with student-level absence flags.
Attendance rate (average daily attendance)
Marking partial-day attendance as full-day; recoding absences as excused once consequential thresholds approach; under-counting students who arrive late or leave early; excluding absent students from the denominator via creative enrollment classification.
Sampled in-school audit by independent observers; reconcile attendance against instructional-time records; publish the full distribution (chronic absenteeism band) alongside the average so a tail-flattening trick changes the visual.
College enrollment rate (post-graduation)
Counting any postsecondary registration regardless of persistence or selectivity; including students enrolled briefly then withdrawn; treating non-credit certificate programs as equivalent to degree-track enrollment; under-tracking students who go directly to skilled trades or the military, which makes the denominator look better.
Report enrollment, persistence, and completion as a connected sequence rather than a single rate; include trades and military pathways in a separate panel rather than excluding them; use National Student Clearinghouse data with full denominator transparency.
Teacher evaluation (composite or value-added)
Teaching to the rubric the evaluator uses; choosing students or sections that maximize value-added growth scores; cooperating with peers to share favorable observation timing; avoiding assignments to schools or students where measured growth is structurally harder.
Use multiple measures with explicitly bounded weight on test-based components; include classroom-observation video samples scored independently; publish the standard error of teacher value-added estimates alongside the point value; never publish individual teacher value-added scores without those error bands.
§10Three protected data zones

Honest improvement requires data that is firewalled from punitive use. Honest research requires data access that is firewalled from accountability politics. Honest accountability requires data that has survived the validity tests for the specific decision being made. Three zones, each with its own rules.

Public accountability

The data the public legitimately gets to see, with consequences attached. Designed to be defensible against gaming and survivable under adversarial reading.

Validity evidence specific to the consequential decision; published gaming model; uncertainty visible on the same panel as the point estimate; opportunity-to-learn data published alongside outcome data.

Improvement

Data educators use to actually fix instruction. Honest visibility requires protection from punitive reuse.

Firewalled from accountability use. Cannot be subpoenaed into personnel actions. Surfaced inside schools by educators; surfaced outside schools only in aggregate forms that cannot be back-traced to individual classrooms.

Research

Data researchers need to build cumulative evidence about what works, for whom, under what conditions.

Strong de-identification; institutional review; data-use agreements with publication requirements; access independent of accountability politics so findings cannot be suppressed when inconvenient.

§11The pluralist signature — why factions can sign

The agenda is structured so that constituencies who disagree about what schools should do can still agree on how schools should be measured. What each faction gets:

Civil-rights advocates
Honest, audited data on opportunity-to-learn gaps and disciplinary disparities — with consequences attached only when the production conditions are also visible.
Teachers and teacher unions
Protection from being judged on a single test score; firewall between formative diagnostic tools and personnel decisions; recognition that working conditions are part of what is being measured.
Parents
A clear, plain-English picture of how their child's school is doing on what — including the inputs the school is being given to work with, not just the outputs.
Fiscal conservatives
Documented metric definitions, audit trails, and gaming analyses — so public dollars chase real performance rather than gamed appearances.
Progressives
Symmetric input measurement that makes resource gaps visible at the same resolution as outcome gaps, so equity arguments rest on data the public can see.
Choice advocates
Comparable data across district, charter, and private settings — same Metric Passport, same audit standards, same gaming analyses, so cross-sector comparisons mean what they appear to mean.
District leaders
Diagnostic and formative data that is actually trustworthy because it is firewalled from punitive reuse; a stable measurement environment that does not collapse every reform cycle.
Researchers
Access to high-quality data with clear use licenses, reliable longitudinal linkages, and protection from being weaponized into accountability fights they did not design for.
Students
Measurement that distinguishes what they have learned from what they have been allowed to learn — and a system that treats them as more than a data point in a school's ranking.
§12The compressed thesis

Public education needs a measurement constitution before it needs another reform war.

The deepest design principle of the agenda is one sentence:

Metrics should make reality more legible without making the institution more gameable.

The first-order reform is not curriculum, governance, funding, or accountability. It is epistemic infrastructure: valid, auditable, anti-gaming, context-aware measurement that separates learning from proxies, diagnosis from punishment, and public transparency from political theater.

§17
Measurement Discipline

Measurement reform is itself unmeasured. These KPIs track whether the epistemic infrastructure of public education is improving — not just whether individual scores are moving. Progress here is a precondition for trusting any other indicator.

KPI A1Annual
States with mandatory school climate surveys

The proportion of U.S. states requiring all public school districts to administer a validated school climate survey (e.g., ED School Climate Survey / CSCL or equivalent) as part of their state data systems. School climate data is the canonical non-academic substrate indicator; its absence from mandatory reporting systems is the clearest signal of a measurement constitution gap.

Baseline

Fewer than 20 states require a validated, statewide school climate survey; most use voluntary or locally-designed instruments with no comparability guarantee. No federal mandate exists post-ESSA.

Target

All 50 states require at least one validated, annually-reported school climate instrument with public district-level results by 2030.

Method

Annual review of state ESSA consolidated plans and state education code; cross-referenced against ED's Civil Rights Data Collection coverage.

Source

ED ESSA Consolidated State Plans; National School Climate Center state-by-state survey; Education Commission of the States policy tracker.

KPI A2Annual
Districts reporting wellbeing indicators in state accountability plans

The percentage of school districts whose state accountability system includes at least one student wellbeing indicator (chronic absenteeism, school connectedness, or student self-reported wellbeing) in its public-facing performance dashboard. ESSA's 'school quality or student success' (SQSS) indicator slot was designed for this; tracking uptake reveals whether states are using it.

Baseline

~38 states include chronic absenteeism in their SQSS indicator as of 2024; fewer than 10 include any validated wellbeing or climate measure beyond attendance.

Target

All 50 states include at least two non-academic wellbeing indicators (including one student-reported measure) in SQSS by 2028.

Method

Annual review of CCSSO state ESSA plan tracker; comparison of each state's SQSS indicator definition against CASEL and CDC validated construct lists.

Source

Council of Chief State School Officers (CCSSO) ESSA Tracker; Attendance Works chronic absenteeism database; ED state plan submissions.

KPI A3Biennial
States using CASEL-aligned SEL assessment in accountability

The number of states that have adopted at least one CASEL-vetted social-emotional learning assessment instrument — with published validity and reliability evidence — in any official accountability or reporting context. This is the operational test of whether 'whole child' language in policy is backed by instrumentation.

Baseline

As of 2025, fewer than 8 states use CASEL Guide-listed instruments in any official state reporting role; most SEL measurement occurs at district level with inconsistent instruments.

Target

At least 25 states use a CASEL-vetted SEL measure in official state or district reporting by 2030, with published comparability standards.

Method

CASEL state policy scan; annual survey of state accountability systems; cross-reference against CASEL's Assessments Guide publication.

Source

CASEL State Scan (casel.org/state-scan); CASEL Assessments Guide; Education Commission of the States SEL policy tracker.

KPI A4Annual
Equity gap in wellbeing indicator adoption

The differential in wellbeing indicator adoption between high-income and low-income (Title I) districts within the same state. When substrate measurement is voluntary, wealthy districts adopt it first; mandatory systems close the gap. This KPI operationalizes the luxury-good failure mode of the intervention.

Baseline

In states with voluntary school climate surveys, adoption rates in non-Title I districts run approximately 2–3x the rates of Title I districts (AIR 2023 state-by-state analysis).

Target

Within-state adoption gap between Title I and non-Title I districts below 10 percentage points for any mandatory wellbeing indicator.

Method

State-level analysis of survey participation data by Title I status; American Institutes for Research (AIR) annual state profile reports.

Source

American Institutes for Research (AIR) School Climate State Policy Reports; ED Common Core Data Title I designation file; state DOE participation reports.

KPI A5Annual
ESSA state plans including a non-academic SQSS indicator

The percentage of ESSA state plans whose 'school quality or student success' (SQSS) indicator is defined as something other than a test-score proxy or graduation-rate proxy. The SQSS slot was the legislative foothold for substrate measurement; tracking what states actually put there reveals whether the accountability architecture is improving or calcifying around academic proxies alone.

Baseline

As of 2024 CCSSO review, approximately 44 of 50 states use SQSS indicators that are primarily attendance, graduation, or test-rate-based; ~6 states include validated climate or wellbeing constructs with published validity evidence.

Target

At least 30 states include a non-academic, validated SQSS construct with published reliability data in their ESSA plan by 2028.

Method

Annual review of approved ESSA state plans; systematic coding of SQSS indicator definitions against CASEL, CDC, and NCES construct taxonomies.

Source

ED ESSA State Plan Repository (ed.gov/essa); CCSSO ESSA Tracker; AIR ESSA SQSS analysis series.

On baselines

Where no national baseline exists, the KPI is still listed. The absence of a baseline is itself the measurement problem — and the first advocacy deliverable is establishing one. Targets marked as contingent on baseline establishment should be read as agenda items for the measurement reform intervention (A), not deferrals.

Substrates engaged · 21 of 64 · 5 bottlenecks
open the map →

The substrates this intervention engages, organized by family. Bottleneck substrates (★) are the upstream layers that determine whether every later argument is sane or distorted.

AEpistemic / Measurement
  • Construct substrate#01
  • Assessment substrate#02
  • Validity substrate#03
  • Reliability substrate#04
  • Comparability substrate#05
  • Growth-measurement substrate#06
  • Opportunity-to-learn substrate#07
  • Data-quality substrate#08
  • Data-governance substrate#09
  • Metric-use substrate#10
  • Gaming-resistance substrate#11
  • Public-reporting substrate#12
  • Uncertainty-communication substrate#13
  • Longitudinal-data substrate#14
  • Causal-inference substrate#15
BGovernance
  • Accountability substrate#19
CFiscal
  • Budget-transparency substrate#32
  • ROI/effectiveness substrate#34
DLabor / Human-Capital
  • Evaluation substrate#40
EInstructional
  • Standards substrate#47
FStudent / Family / Community
  • Attendance substrate#57