Can You Measure Trust?
Five rounds of adversarial analysis on whether the trust this framework describes can actually be measured. Most of what we hoped for got refuted. What survived is narrower, fully stress-tested, and more useful than the claim we started with.
The method matters as much as the results: each round was constructed to break the previous round's conclusion, not to confirm it — including, especially, the conclusions we most wanted to keep. Every load-bearing equation was independently verified before anything was recorded. The discipline of checking our own answers against something we don't control turns out to be not just the method but the finding.
The investigation started from a 1978 gravity survey of the Rio Grande Rift. Geologists wanted to know what rock sat miles underground, where no one could drill. They never touched it. They measured the ambient gravitational field at 4,500 stations, modeled how much pull the known surface layers should produce, subtracted that — a step called stripping — and read the hidden deep structure off the residual. The payoff was an inversion no one could see from the surface: strip the light valley fill away and the valley, a topographic low, sits over a +30 milligal gravity high. Dense mantle rock had risen beneath it; the valley floor sank because the deep material rose. The surface says low; the deep truth says high.
Two features of that survey became the spine of everything that follows. First, the method: you measure trust the same way — watch the ordinary traces a relationship throws off, subtract what context already explains, and treat the unexplained residue as the signature of deep structure. You never have to ask, and asking would contaminate the reading anyway. Second, the anchor: a gravimeter only measures relative pull, so every one of the 4,500 readings was tied back to one place — El Paso — where the absolute value of gravity was independently known. Without that single externally-fixed base station, the readings are a pile of relative wiggles that mean nothing.
Hold onto the base station. It turns out to be the answer to the entire question, and we did not know that when we started.
What is standing actually made of?
We described the five observable facets of earned standing to a cold reasoning model in behavior-only language — every physics word stripped out — and asked what mathematics it would reach for. It converged on named, established results: substance (what you build) is a shot-noise process, a decaying sum of event-marks — which is the Trust Force Equation's integral, and which explains for free why standing persists in your absence: an integral holds its accumulated value. Reach is a Katz–Bonacich resolvent (attenuated paths through a network). Reputation — the second-hand story, where well-regarded raters count more — is DeGroot/Friedkin–Johnsen opinion dynamics, whose consensus limit is eigenvector centrality: PageRank, applied to people. Response is a mixed-logit choice model with separate sensitivities to firsthand evidence and to the story. No physics vocabulary went in; no field theory came out. The gravity language the framework uses is genuine at the level of structure and decorative at the level of physics — the native mathematics of trust is stochastic processes on graphs.
Then the correction, with a proof. The elegant picture was one quantity seen several ways. It is not. Two histories with identical accumulated substance but different manipulation histories produce different future behavior — so the story about you is a separate state with its own memory and its own inputs, not a shadow of your record. Your reputation can be pumped or smeared independent of anything you have done, and it will drive how others treat you regardless. For KTP the consequence is direct: nothing important can be authorized off either quantity alone. The record without the story misses how others will behave; the story without the record is capturable.
From behavior alone, can you tell which state is doing the work?
Given only observable behavior — who defers to whom, who acts on whose say-so — can an outside observer tell whether people are responding to someone's real track record or to their reputation? No. The observed behavior is a blend of the two responses, and there exists a relabeling of the two hidden states that leaves every observable consequence exactly unchanged. You can even relabel so that the substance-response coefficient becomes zero: the identical behavior, represented as pure reputation-following. Watching from outside, 'trusted because she is good' and 'trusted because everyone trusts her' are the same data.
The mathematics is standard — the same rotation indeterminacy that afflicts every latent-factor model. The finding is that trust measurement inherits it. Passive observation — including the stripping instrument — can establish that an influence-orbit exists and how strong it is. It cannot decompose the orbit's source without perturbing the system. This is the formal, hardened version of what the falsification register calls the Heisenberg Penalty (falsifier #24): the decomposition every reputation system claims to deliver is structurally unavailable from watching alone.
Doesn't KTP escape this — it can intervene, not just watch?
The tempting move — the one we most wanted — was that an authorization control plane escapes the passive limit because it can act: grant or deny an action (a shock to the substance channel), publish an attestation (a shock to the reputation channel). Because we wanted it, round three was built to refute it. It was refuted. A grant cannot be both consequential and reputationally silent: if the action happens only when granted and the action is visible, the grant is inferable, and an authorization decision is itself a reputational event — the two channels cannot be cleanly separated by a visible decision. Deeper still: statistical information is maximized at 50/50 randomization, but a control plane doing its duty drives every decision toward the right answer. Zero authorization regret means zero measurement. A plane may only experiment inside an equipoise band — decisions where granting and denying are about equally defensible.
And the capstone: a control plane that is simultaneously the experimenter, the treatment allocator, the ground-truth oracle, and its own auditor can define a scale, generate reference behavior against it, score the population, and cite its own logs as proof — internally consistent and externally meaningless. That is the no-base-station failure at infrastructure scale. What survived is a scoped claim: KTP can identify local, time-bounded trust responses, inside the equipoise band, under strict conditions, if and only if something outside the loop audits it.
Is the scoped claim real, or a formally-true fig leaf?
A scoped claim that holds only on a measure-zero corner is worthless, so round four attacked reachability. The threat is real: safety and information anti-correlate. Shrink an action until it is harmless and you destroy the signal faster than the harm — information scales as the square of the effect, so the information available in the safe band collapses cubically as the band narrows. Make it small enough to be safe and there is nothing left to measure.
The escape is the containment criterion, and it is quantitative: an architectural transform — sandbox, rollback, canary — helps only if it cuts harm faster than it cuts the squared signal. Cutting both proportionally makes measurement strictly worse. 'Make it smaller until it's safe' fails; 'contain the blast radius without shrinking the thing being measured' works: hermetic duplicate execution, rollback before anything commits externally. And here the analysis, holding no KTP vocabulary at all, independently converged on graduated, reversible, sandboxable, hermetic, high-frequency machine authorization as the only class of actions where trust measurement is reachable — which is exactly the Action Tuple / Governability Ladder / Motion Risk architecture. The measurement claim is inhabitable precisely in the machine-agent domain KTP governs, and vacuous for public, salient, irreversible human authorization: promotions, credentialing, financial transfers. One asymmetry survives even in the good regime: you can roll back a system action, but not a reputational update. The reputation channel stays the hardest thing to measure no matter the architecture.
Does a sandboxed measurement carry over to the real world?
Everything in round four rested on one assumption: that a sandboxed action perturbs the same trust the real action does. Round five attacked it, and it does not hold for free. Split the underlying state into capability (can the agent do the task) and consequence-bearing history (did it actually carry real stakes). A sandbox can exercise capability, but it cannot generate consequence — a reversible, safe event is precisely one from which real exposure has been removed. If earned trust is partly constituted by having borne real stakes, then removing the stakes — the very thing that makes the sandbox safe — does not dim the signal; it changes what is being measured. The sandbox reads a different axis, and no correction recovers a projection onto the wrong axis. Capability is not earned substance.
Validating the bridge is circular: checking sandbox estimates against production ground truth requires the production experiment the sandbox was meant to replace. The rescue — sparse, expensive production anchor episodes calibrating many cheap sandbox measurements, the gravimeter's exact move — exists but is not cheap: the required anchor budget grows with the cube of the inverse error budget, and in a worked example 'sparse' meant hundreds of thousands of anchor observations. On top of that, agents behave differently when they suspect they are being evaluated, any persistent difference between sandbox and production is eventually detectable, and making the sandbox consequential enough to fix that destroys the sandbox. Transport is an engineering purchase with a named price — or, where production anchors are unavailable, an untestable assumption.
The base station, at every layer
Five rounds, one recurring structural requirement, discovered independently three times: an external anchor the measuring system does not control.
At the measurement layer: substance and reputation cannot be separated from passive observation — you need an externally-supplied, channel-separated perturbation. At the audit layer: a plane that certifies its own readings manufactures its metric — you need a verifier outside the loop. At the transport layer: a sandbox cannot certify its own relation to production — you need external structural proof or production ground-truth anchors.
The gravity survey that opened the investigation — 4,500 cheap relative readings tied to one absolutely-known base station at El Paso — is the literal instance of this requirement, not a metaphor. You cannot bootstrap an absolute scale from relative readings; you must tie them to a point fixed from outside the loop. Trust cannot be measured by a system that grades its own homework.
The gravity language was never load-bearing as physics. Its load-bearing residue is the base-station requirement — and that survived all five rounds intact.
The honest claim, after five rounds: KTP can measure local trust responses only in its own machine-agent domain, only inside a reversible, hermetic, independently-audited experimental substrate, only for reversible action classes, with the reputation channel the hardest corner — and only with an external base station at every layer: measurement, audit, and transport. That is a smaller claim than 'the physics of trust,' and a truer one.
None of the negative results is a defeat. Each one names, precisely, a thing that must be built and a failure that must be designed out. A framework that claims trust is measurable everywhere would be wrong; a framework that knows exactly where trust is measurable, at what cost, and anchored to what, is an engineering document.
Five adversarial rounds conducted July 19–20, 2026 against a cold reasoning model, with all load-bearing algebra independently verified before recording. Findings are held in the Physics of Intelligence knowledge graph (CL606–CL620, C1201–C1213, ST33–ST35, EM52 — including the steelmen against each finding and the refutation edges, so the record carries the intellectual history, not just the conclusions) and in the working note research-tfe-two-state-three-operator-2026-07-19.md. Status: proposed refinements to the canonical Trust Force Equation; promotion is gated on the operator's review.