KTP/
github ↗
created 24 August 2026 · last modified 24 August 2026
the credibility close

Methods, limits, what comes next.

KTP does not solve all AI risk. That is why it is credible. The v0.2 Scorecard is a deterministic first-pass review expansion, not human adjudication or empirical validation. This page says what that means, and what it does not.

§01provenance

run date2026-05-04
row count1,612 baseline risks
baseline inputKTP_RPT_Enriched_Master_v0_1.csv
source: MIT AI Risk Repositoryairisk.mit.edu
source: MIT AI Risk Mitigationsairisk.mit.edu/ai-risk-mitigations
source: MIT AI Incident Trackerairisk.mit.edu/ai-incident-tracker
source: Kinetic Trust Protocolkinetic-trust-protocol.net

The MIT AI Risk Repository is the external risk universe; the Scorecard is KTP's audited response. Neither was generated by the other. The match is the test.

§02limitations (verbatim)

from the manifest
“v0.2 is a deterministic first-pass review expansion, not human adjudication or empirical validation.”

That sentence is doing real work. It says: every score in the 1,612-row matrix was assigned by a deterministic ruleset against the MIT risk text, not by a panel of human reviewers and not by comparison to real incidents. The scoring is reproducible. It is not yet validated.

Read the Scorecard accordingly: as a first map of where KTP could plausibly apply, where it explicitly does not, and where its claims are still at the conceptual or mechanistic tier rather than the scenario-supported one. The 170 P0 risks flagged for human adjudication, the 819 candidates for incident backtest, and the 337 generated test cases are exactly the work the v0.3+ passes will do.

§03failure modes

Naming the failure modes is what allows defense-in-depth. These are the seven explicit ways a KTP-class architecture can fail to govern an action. Read them as what the surrounding controls must handle, not as what makes KTP a bad bet. For where the enforcement surfaces are designed, see architecture.

01
agent bypasses PEP
An agent reaches a resource through a direct API call that does not pass through the policy enforcement point.
02
unmanaged tool path
A tool is wired into the agent runtime without being registered with the gateway, so its actions are never evaluated.
03
incomplete telemetry
Decisions are made and actions are taken without being logged, breaking audit and removing the trajectory record.
04
incorrect policy classification
A risk is mapped to the wrong bucket, so the wrong rule fires; the action is allowed when it should be constrained, or vice versa.
05
over-permissive delegation
Humans grant the agent more scope than the task requires, expanding the blast radius beyond what policy can contain.
06
human override misuse
The escalation path becomes a rubber-stamp; manual approval degrades into unchecked authorization at speed.
07
pre-runtime risks
Training-data poisoning, model misalignment, and supply-chain compromise occur before any action reaches the gateway. KTP cannot see them.
08
the representation boundary
The should gate's normative half cannot be derived from telemetry. A fact that is both fast and signatureless, a layoff an hour ago, a response out of all proportion, has no signal to read and no rule to pre-set. Where that context is neither injected nor routed to a human, the should gate is blind. This is a structural limit, not a bug: appropriateness needs a model that holds meaning, and that model lives outside the physics. The should gate is also earlier-stage than the can gate, and is built second.

§04what v0.3+ will add

The next passes turn the deterministic first-pass into something adjudicated, backtested, and run.

human adjudication
Real review of the 170 P0 risks by domain experts, not deterministic scoring. Each P0 claim either survives, is downgraded, or is retracted before public use.
incident backtest
The 819 backtest candidates compared against the AI Incident Database (AIID). For each, the counterfactual question: would KTP have prevented, constrained, or only attributed this incident?
prototype runs
The 337 generated test cases executed against a working KTP reference implementation. Real veto, real allow, real escalation; not just expected behavior on paper.
governability tagging
Every risk in the scorecard tagged with a level on the governability ladder (see §06), so readers can see at a glance which risks KTP makes preventable, which it makes interruptible, and which it only makes observable.

§05the bounded public claim

the claim, narrower than the framework
“KTP directly addresses the subset of AI risks that depend on autonomous action, tool use, data movement, delegation, transaction, or boundary crossing. For other risks, KTP may provide accountability, telemetry, or governance evidence, but not direct mitigation.”
the strongest critique, and the response
critique
“KTP is useful, but the verb framing risks collapsing too much complexity into action labels. AI risk does not always present as a clean action event. Some harms are cumulative, statistical, latent, delegated, or structural. If KTP claims to make AI risk governable by controlling verbs, it may overstate its reach. To be credible, KTP must distinguish between risks it can directly constrain, risks it can only observe, risks it can route to governance, and risks outside its scope.”
response
“Correct. KTP is not a theory of all AI risk. It is a control methodology for the action layer of AI risk. Its value is not universal coverage. Its value is disciplined translation: when a risk becomes an attempted motion, KTP asks whether that motion should be allowed, constrained, delayed, logged, challenged, routed, or denied.”

§06the governability ladder

The word governable does too much work on its own. Governable can mean observable, attributable, constrainable, interruptible, reversible, or pre-authorized. These are not the same thing. The ladder makes the difference explicit.

L0not visibleThe action happens without leaving a trace the environment can see.
L1observableThe action is logged. Something happened, and the record exists.
L2attributableThe action can be traced back to a specific actor, agent, or trajectory.
L3constrainableThe action's scope, rate, or parameters can be limited at the moment of execution.
L4interruptibleThe action can be stopped mid-flight, by policy or by escalation.
L5reversible / containableThe consequence can be rolled back, or its blast radius is bounded by design.
L6pre-authorized / pre-denied by policyThe decision is made before the action is attempted; the environment refuses or permits without negotiation.

KTP does not make every risk preventable. It moves risks up the governability ladder.

next doors

Skeptical

Concession first. Nine objections, answered. The line you can use against us.

Risks

The data the methods describe. 1,612 risks, the heat map, the bimodal distribution.

Architecture

Where the controls live. PDP/PEP, seven enforcement surfaces, ten control sets.

How to cite
APA 7th ed.

Perkins, C. (2026, May 6). Methods, limits, what comes next. Kinetic Trust Protocol. https://kinetic-trust-protocol.net/enterprise/methods
Plain text

Chris Perkins, "Methods, limits, what comes next," Kinetic Trust Protocol, https://kinetic-trust-protocol.net/enterprise/methods (accessed May 6, 2026).