Measurement Reform.
Stop measuring what's cheap. Start measuring what's load-bearing.
A third-grade teacher spends six weeks of the year on test prep. She knows the test measures something narrower than what her students need. She does it anyway. Her evaluation depends on it.
State accountability systems built around standardized tests select for a narrow band of cognitive performance: rapid recognition under time pressure, multiple-choice elimination, surface-form fluency. The tests are cheap to score, defensible in court, and produce a single number that travels well politically. Every other capacity — argument, original synthesis, oral defense, primary-source reading — is unmeasured and therefore, in budget terms, undefended.
What the system can measure becomes what teachers teach. What teachers teach becomes what students learn. What students learn becomes what they can do as adults. The narrow measure becomes the wide reality. AI now does most of what these tests measure better than the median student. The capacities the tests do not measure — interpretation, defense, judgment — are the only ones that remain load-bearing.
Meadows ranks goals of the system above information flows above rules. The leverage point is not 'better tests'; it is changing the goal of measurement from 'sortable number' to 'evidence the student can defend their own thinking.' Under that goal, the tests change shape.
State boards of education direct staff to develop oral defense, Socratic dialogue, and original synthesis assessments under controlled conditions. These are conducted by trained observers, not auto-scored. They are slower, more expensive per student, and politically harder to defend in court — and they measure what AI cannot replace.
You're hiding achievement gaps. Standardized tests are how civil-rights organizations surface inequity. Removing the measurement is the favorite move of districts that don't want to be accountable for their poorest students.
The reform demands BETTER measurement, not less. Oral defense and original synthesis can be normed and reported by demographic subgroup just as standardized tests can. Achievement gaps will still be visible — and they will be visible on capacities that matter, not capacities that have been automated away. The civil-rights commitment is preserved; the instrument changes.
- Pearson PLC — $2.1B revenue from K-12 testing and adaptive products
- Educational Testing Service (ETS) — administers SAT, AP, GRE; prepares state assessments
- College Board — $1.6B nonprofit revenue from SAT and AP; vertical integration into AP curriculum
- Accountability-hawk legislators (both parties) whose policy identity is built on 'measurable results'
- Civil-rights organizations rightly worried that removing measurement will hide gaps (NAACP LDF, Education Trust, Civil Rights Project at UCLA)
- 'You're hiding achievement gaps.'
- 'This is unmeasurable, therefore unaccountable.'
- 'Subjective grading is biased grading.'
- 'Districts will game oral defense the way they game everything else.'
- 'Federal ESEA accountability requires standardized measures.'
If three or more states adopt oral defense and original synthesis assessments and, after five years, demographic achievement gaps on those assessments are not narrower than gaps on the prior standardized tests — and student capacity to defend an argument under cross-examination has not measurably risen — the intervention has failed in the form deployed. Revision required.
- FairTest (National Center for Fair & Open Testing) — 30-year position against high-stakes standardized testing
- Knowledge Matters Campaign — Content-rich curriculum advocacy; aligned on measure-the-substrate framing
- American Federation of Teachers — Anti-test reform but cautious about oral-defense workload
- National Education Association — Same posture; depends on local affiliate
- Pearson PLC, ETS, College Board — Direct revenue threat
- Civil-rights organizations focused on gap-surfacing — The intervention preserves their core commitment; the framing has not yet reached them
Oral defense and original synthesis assessments require trained observer time and small-group settings. Without paired investment in school staffing (see Intervention D), the assessments will be high-quality in well-resourced districts and degraded in under-resourced ones — reproducing the gap the reform is meant to surface.
Tier-blindness. Wealthy districts already do oral defense (it is the structure of private-school humanities programs). Reform without provisioning hardens the existing tier structure. PAIR WITH INTERVENTION D.
Oral defense is a luxury good today: private schools, classical schools, and elite public magnets do it routinely; comprehensive public schools do not. The mechanism is staffing density (one observer per six students versus one teacher per twenty-eight) and credentialing (humanities-degreed staff who can run Socratic dialogue).
State constitutions establish public education as a state responsibility (every state but Hawaii has explicit education clauses). State boards of education have rulemaking authority over assessment systems under most state administrative procedure acts. ESEA Title I federal accountability provisions require state assessments but do not specify multiple-choice format.
Trained observers (humanities-degreed school staff). Small-group settings. Time on the school day for assessment. Norming process to surface demographic differences without hiding them. None of these exist at scale in comprehensive public schools today.
State boards review assessment frameworks on 5-7 year cycles. Amendments can be proposed by board members, the state superintendent, or via formal stakeholder petition. Federal alignment changes require ESEA reauthorization (typically 7-10 year cycles).
Teachers reclaim 4-6 weeks per year previously spent on test prep. Students develop sustained interpretive practice. Assessment data carries more signal about what students can actually do.
Curriculum shifts to support the new assessment. Teacher training programs add Socratic-method coursework. Parents who used to track test scores start asking about defense scores, which carry richer information. The political vocabulary of accountability shifts.
State board of education regulatory action (faster than legislation; 1-3 year cycles versus 5-10). Pair with district-level pilots through model school board resolution.
Low. Federal ESEA does not specify assessment format; states retain authority. Risk is not preemption but federal accountability score-card incentives that reward the old measure.
2027-2029. The next ESEA reauthorization opens federal flexibility on assessment format. State board cycles reset 2027-2030 in 28 states. AI-driven test obsolescence accelerates the political case.
- Koretz, D. (2008). Measuring Up: What Educational Testing Really Tells Us. Harvard University Press.
- Ravitch, D. (2010). The Death and Life of the Great American School System. Basic Books.
- Meadows, D. (1999). Leverage Points: Places to Intervene in a System. Sustainability Institute. donellameadows.org/archives/leverage-points-places-to-intervene-in-a-system
- U.S. Department of Education (2015). Every Student Succeeds Act (ESEA reauthorization), Title I assessment requirements. Public Law 114-95. www.congress.gov/bill/114th-congress/senate-bill/1177
Corrections via /advocacy/contribute. Public response SLA: 14 days for substantive corrections, 30 days for elaborations. No comments section by design.
Submit a correction →State Departments of Education (assessment policy compliance); State Boards of Education (standard-setting authority); AERA/APA/NCME (voluntary standards, no legal authority).
Formal complaint to state DOE citing specific misuse of diagnostic data for accountability purposes; or annual district compliance certification showing violation of protected-diagnostic-space requirements.
State DOE can condition discretionary funding on compliance; federal Title I monitoring can flag measurement-integrity violations. No private right of action currently exists in most states.
No federal private right of action for measurement-integrity violations. Enforcement is administratively discretionary, creating highly variable implementation across districts.
- Principle 1 (Clearly defined boundaries): Protected diagnostic space is defined by use restriction — data collected for instruction improvement cannot be repurposed for accountability.
- Principle 3 (Collective choice arrangements): AERA/APA/NCME standards are set through a collective process; the Metric Passport requires that standard-setting committees exclude commercial testing vendors.
- Principle 6 (Conflict resolution mechanisms): Correction surface on each intervention page is the public dispute mechanism; amendment threshold (10+ submissions) is the collective-choice trigger.
Any teacher, student, family, or district official can file a complaint through the state DOE's assessment office. The correction surface on this page is the public-facing channel.
Diagnostic data protected under this framework cannot be used for commercial profiling, vendor product development, or any purpose other than improving instruction for the assessed student. Violation converts the data from protected diagnostic to impermissible commercial use, triggering state data privacy statute remedies.
When a district misuses diagnostic data due to misunderstanding (not bad faith): apology from district leadership, remediation plan showing new data-governance training, third-party audit of data use within 60 days.
When a district knowingly repurposes diagnostic data for accountability after committing not to: formal state DOE complaint, independent audit of all assessment practices, public disclosure of the violation and remediation timeline. Apology alone is insufficient — requires demonstrated dispositional change in data governance.
For the cross-cutting framing — the four-step pedagogy, the three required components, the five-tier separation, and the twelve-field Metric Passport pattern as they apply across all domains — see /advocacy/kpis. This page is the K-12 instantiation: the canonical implementation of the pattern named there.
Fix the map before fighting over the route.
What follows is the long form of the intervention above. Its ambition is structural and pre-political: a measurement constitution that makes the data legible to constituencies that do not yet agree on curriculum, governance, funding formulas, unions, vouchers, charters, discipline, or accountability. The claim is that the first-order reform of public education is not any of those. It is the epistemic infrastructure beneath them.
Substrate-level, pre-political measurement reform means reforming the shared measurement infrastructure beneath education politics — the constructs, instruments, audit standards, and use licenses — before arguing over ideology, governance, funding formula, curriculum, unions, vouchers, charters, discipline, or accountability. None of those second-order fights can be settled honestly when the underlying numbers do not measure what the public, the schools, and the researchers think they measure.
Pre-political does not mean value-free. It means politically prior. The agenda is built so that factions who disagree about what schools should do can still trust the measurement of what schools are doing, and act on that measurement before agreeing on next moves. It is the prerequisite for a real fight, not a substitute for one.
Public education needs a measurement constitution before it needs another reform war.
The structural failure of US K-12 measurement is that the same number gets used for diagnosis, formative feedback, public reporting, accountability sanctions, and research — all from one instrument, all at once. The same signal becomes both microscope and hammer. A microscope is meant to make something visible. A hammer is meant to land a consequence. When the same instrument is asked to do both, the microscope warps to protect against the hammer, and the hammer lands on whatever the microscope can still see.
The Every Student Succeeds Act preserved the high-stakes annual testing architecture inherited from No Child Left Behind even as it loosened federal sanctions. The collapse of measurement tiers into a single accountability-flavored signal predates ESSA and survived it. RAND's monograph-length evaluation of test-based accountability documents the consequences in detail: narrowed curriculum, score inflation, gaming of student classifications, and a measurable gap between scores on consequential tests and scores on lower-stakes audit tests measuring the same construct.
- U.S. Department of Education. Every Student Succeeds Act (ESSA). www.ed.gov/laws-and-policy/laws-preschool-grade-12-education/every-student-succeeds-act-essa
- Hamilton, Stecher, & Klein (RAND). Making Sense of Test-Based Accountability in Education (MR-1554). www.rand.org/content/dam/rand/pubs/monograph_reports/2002/MR1554.pdf
"The more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor."
— Donald T. Campbell, 1976
Campbell's Law is the first design constraint of any honest measurement reform. It is not a warning to be acknowledged in a footnote and then ignored; it is the reason the architecture exists. A measurement reform that does not begin with Campbell's Law is not a reform — it is a new generation of the same trap. The Institute of Education Sciences has documented how publicly reporting safe-schools data corrupts the reporting itself, in exactly the way Campbell predicted half a century ago.
- Institute of Education Sciences (REL Mid-Atlantic). Ask an Expert: How publicly reporting safe-schools data affects accuracy. ies.ed.gov/rel-mid-atlantic/2025/01/ask-expert-summarize-behavioral-research-how-publicly-reporting-safe-schools-data-affects-accuracy
A defensible measurement system has at least five distinct tiers, each with its own purpose, audience, stakes, and validity-evidence requirements. Current US K-12 measurement collapses them into one scoreboard.
- Diagnostic. Identify what an individual student does and does not yet know, so a teacher can adjust instruction. Audience: Teacher, student, parent. Stakes: Low-stakes by construction. Results are formative input, never used for school grades or staff personnel actions.
- Formative. Track progress within a unit or term to inform pacing, grouping, and re-teaching decisions. Audience: Classroom teacher, instructional coach, department lead. Stakes: Internal to the school. No public reporting role.
- Public reporting. Give parents, journalists, and the public a legible, comparable picture of how schools are doing across populations. Audience: Parents, journalists, taxpayers, researchers. Stakes: Reputational. Should be lower-stakes than accountability — a thermometer, not a sentence.
- Accountability. Trigger consequential action — interventions, restructuring, sanctions — when defined floors are crossed. Audience: State education agency, federal compliance officers, courts. Stakes: High. Must use validity evidence specific to the consequential decision being made.
- Research. Build cumulative evidence about what works for whom, under what conditions, and why. Audience: Researchers, evaluators, policy analysts. Stakes: None for individual students or schools. Findings inform future policy design.
Current systems run a single annual summative state assessment and route it into all five uses — diagnostic, formative, public reporting, accountability, and research. Each tier's specific gaming pressures arrive together; each tier's specific validity-evidence requirements get ignored together. The resulting number is asked to be everything to everyone, and ends up being trustworthy to no one.
What is a school actually doing? Seven construct families cover the legitimate territory: learning achievement, learning growth, opportunity to learn, instructional quality, school climate and safety, system capacity, and post-school outcomes. A defensible public picture surfaces all seven, with confidence intervals, rather than collapsing them into a single composite score that conceals more than it reveals.
Every public education metric ships with a structured passport: twelve fields documenting what the metric measures, what it is licensed for, what it is forbidden for, and how it can be gamed. The passport is published on the same surface as the metric. A number without its passport is not a measurement — it is a rumor with a citation.
The passport pattern follows the AERA/APA/NCME Standards for Educational and Psychological Testing, which require that validity claims be specific to the use being made of the score. AERA's position on high-stakes testing is unambiguous on this point: a single test should not be used as the sole determinant of consequential decisions about individuals, schools, or systems.
- American Educational Research Association, American Psychological Association, National Council on Measurement in Education. Standards for Educational and Psychological Testing. ncme.org/resources/books/testing-standards
- American Educational Research Association. Position Statement on High-Stakes Testing. www.aera.net/About-AERA/AERA-Rules-Policies/Association-Policies/Position-Statement-on-High-Stakes-Testing
Chronic absenteeism is the rare metric that already has strong validity evidence as a leading indicator of school disengagement, that maps cleanly onto a recognizable construct, and that nonetheless gets gamed in predictable ways the moment it acquires consequences. It is a fair test for the Metric Passport pattern.
Every public metric ships with these twelve fields documented. Empty fields fail the publication standard. The fields exist so a reader, a journalist, or an opposing party can see what the metric is, what it is licensed to do, and how it can be gamed.
The National Assessment of Educational Progress is the closest thing the United States has to a working model of disciplined measurement. It is sampled rather than universal, group-level rather than individual, and lower-stakes by design. No student, teacher, or school is sanctioned on the basis of a NAEP score. The result is that NAEP can do what state summative tests cannot: report what students actually know without being warped by the consequences attached to the report.
NAEP is a thermometer. State summative tests are increasingly asked to be a sentencing instrument. A defensible measurement architecture keeps the thermometer separate from the sentence, and uses the thermometer for the public's macroscopic understanding of system performance.
- National Center for Education Statistics. About NAEP — The Nation's Report Card. nces.ed.gov/nationsreportcard/about
A school is a learning production system. Asking only what its outputs are without measuring what it is being given to work with is not measurement; it is moral theater. Opportunity-to- learn data brings input measurement to the same resolution as output measurement.
The question shifts. Instead of asking why are scores low at this school? a defensible system asks what learning production system generated these outcomes? and reports both panels side by side.
- Teacher vacancies and long-term substitute deployment
- Class sizes by grade and subject
- Instructional minutes per subject area
- Curriculum access (which materials the school actually has)
- Advanced course access (AP, IB, dual enrollment, advanced math sequences)
- Special-education service delivery against IEP minutes owed
- English-learner support staffing and program model fidelity
- Counselor-to-student ratios
- Facility condition (HVAC, water quality, seismic, lead exposure)
- Connectivity (broadband at school, broadband at home)
- School safety conditions (real, not survey-coached)
- Chronic absenteeism drivers (housing instability, transit, health access)
- Leadership turnover (principal years-of-service)
- Teacher retention (year-over-year, by school)
For every public metric, the gaming model is published with the metric. Not in an audit document a few researchers will read. Not in a footnote. On the same surface as the number. Each metric, the predictable gaming behavior it produces under consequential pressure, and the structural countermeasure required to make the metric usable.
Honest improvement requires data that is firewalled from punitive use. Honest research requires data access that is firewalled from accountability politics. Honest accountability requires data that has survived the validity tests for the specific decision being made. Three zones, each with its own rules.
The data the public legitimately gets to see, with consequences attached. Designed to be defensible against gaming and survivable under adversarial reading.
Validity evidence specific to the consequential decision; published gaming model; uncertainty visible on the same panel as the point estimate; opportunity-to-learn data published alongside outcome data.
Data educators use to actually fix instruction. Honest visibility requires protection from punitive reuse.
Firewalled from accountability use. Cannot be subpoenaed into personnel actions. Surfaced inside schools by educators; surfaced outside schools only in aggregate forms that cannot be back-traced to individual classrooms.
Data researchers need to build cumulative evidence about what works, for whom, under what conditions.
Strong de-identification; institutional review; data-use agreements with publication requirements; access independent of accountability politics so findings cannot be suppressed when inconvenient.
The agenda is structured so that constituencies who disagree about what schools should do can still agree on how schools should be measured. What each faction gets:
Public education needs a measurement constitution before it needs another reform war.
The deepest design principle of the agenda is one sentence:
Metrics should make reality more legible without making the institution more gameable.
The first-order reform is not curriculum, governance, funding, or accountability. It is epistemic infrastructure: valid, auditable, anti-gaming, context-aware measurement that separates learning from proxies, diagnosis from punishment, and public transparency from political theater.
Measurement reform is itself unmeasured. These KPIs track whether the epistemic infrastructure of public education is improving — not just whether individual scores are moving. Progress here is a precondition for trusting any other indicator.
The proportion of U.S. states requiring all public school districts to administer a validated school climate survey (e.g., ED School Climate Survey / CSCL or equivalent) as part of their state data systems. School climate data is the canonical non-academic substrate indicator; its absence from mandatory reporting systems is the clearest signal of a measurement constitution gap.
Fewer than 20 states require a validated, statewide school climate survey; most use voluntary or locally-designed instruments with no comparability guarantee. No federal mandate exists post-ESSA.
All 50 states require at least one validated, annually-reported school climate instrument with public district-level results by 2030.
Annual review of state ESSA consolidated plans and state education code; cross-referenced against ED's Civil Rights Data Collection coverage.
ED ESSA Consolidated State Plans; National School Climate Center state-by-state survey; Education Commission of the States policy tracker.
The percentage of school districts whose state accountability system includes at least one student wellbeing indicator (chronic absenteeism, school connectedness, or student self-reported wellbeing) in its public-facing performance dashboard. ESSA's 'school quality or student success' (SQSS) indicator slot was designed for this; tracking uptake reveals whether states are using it.
~38 states include chronic absenteeism in their SQSS indicator as of 2024; fewer than 10 include any validated wellbeing or climate measure beyond attendance.
All 50 states include at least two non-academic wellbeing indicators (including one student-reported measure) in SQSS by 2028.
Annual review of CCSSO state ESSA plan tracker; comparison of each state's SQSS indicator definition against CASEL and CDC validated construct lists.
Council of Chief State School Officers (CCSSO) ESSA Tracker; Attendance Works chronic absenteeism database; ED state plan submissions.
The number of states that have adopted at least one CASEL-vetted social-emotional learning assessment instrument — with published validity and reliability evidence — in any official accountability or reporting context. This is the operational test of whether 'whole child' language in policy is backed by instrumentation.
As of 2025, fewer than 8 states use CASEL Guide-listed instruments in any official state reporting role; most SEL measurement occurs at district level with inconsistent instruments.
At least 25 states use a CASEL-vetted SEL measure in official state or district reporting by 2030, with published comparability standards.
CASEL state policy scan; annual survey of state accountability systems; cross-reference against CASEL's Assessments Guide publication.
CASEL State Scan (casel.org/state-scan); CASEL Assessments Guide; Education Commission of the States SEL policy tracker.
The differential in wellbeing indicator adoption between high-income and low-income (Title I) districts within the same state. When substrate measurement is voluntary, wealthy districts adopt it first; mandatory systems close the gap. This KPI operationalizes the luxury-good failure mode of the intervention.
In states with voluntary school climate surveys, adoption rates in non-Title I districts run approximately 2–3x the rates of Title I districts (AIR 2023 state-by-state analysis).
Within-state adoption gap between Title I and non-Title I districts below 10 percentage points for any mandatory wellbeing indicator.
State-level analysis of survey participation data by Title I status; American Institutes for Research (AIR) annual state profile reports.
American Institutes for Research (AIR) School Climate State Policy Reports; ED Common Core Data Title I designation file; state DOE participation reports.
The percentage of ESSA state plans whose 'school quality or student success' (SQSS) indicator is defined as something other than a test-score proxy or graduation-rate proxy. The SQSS slot was the legislative foothold for substrate measurement; tracking what states actually put there reveals whether the accountability architecture is improving or calcifying around academic proxies alone.
As of 2024 CCSSO review, approximately 44 of 50 states use SQSS indicators that are primarily attendance, graduation, or test-rate-based; ~6 states include validated climate or wellbeing constructs with published validity evidence.
At least 30 states include a non-academic, validated SQSS construct with published reliability data in their ESSA plan by 2028.
Annual review of approved ESSA state plans; systematic coding of SQSS indicator definitions against CASEL, CDC, and NCES construct taxonomies.
ED ESSA State Plan Repository (ed.gov/essa); CCSSO ESSA Tracker; AIR ESSA SQSS analysis series.
Where no national baseline exists, the KPI is still listed. The absence of a baseline is itself the measurement problem — and the first advocacy deliverable is establishing one. Targets marked as contingent on baseline establishment should be read as agenda items for the measurement reform intervention (A), not deferrals.
The substrates this intervention engages, organized by family. Bottleneck substrates (★) are the upstream layers that determine whether every later argument is sane or distorted.
- ★Construct substrate#01
- Assessment substrate#02
- ★Validity substrate#03
- Reliability substrate#04
- Comparability substrate#05
- Growth-measurement substrate#06
- ★Opportunity-to-learn substrate#07
- ★Data-quality substrate#08
- Data-governance substrate#09
- ★Metric-use substrate#10
- Gaming-resistance substrate#11
- Public-reporting substrate#12
- Uncertainty-communication substrate#13
- Longitudinal-data substrate#14
- Causal-inference substrate#15
- Accountability substrate#19
- Budget-transparency substrate#32
- ROI/effectiveness substrate#34
- Evaluation substrate#40
- Standards substrate#47
- Attendance substrate#57