Code is notthe product.
Agents now write code faster than anyone can read it. Onus is the layer every coding agent works through. It reads each change for what it means, routes it by risk, and approves it on evidence.
1,450 changed lines
AddedRemoved
4 changes in meaning
1 needs a personRouted to a person: a new outside vendor will receive customers' phone numbers.
Point at a change to see the lines behind it. An illustrative report for the manifesto's example feature.
Output scales. Attention doesn't.
When one company set out to double merged pull requests per engineer, output did double. Review couldn't keep up. The share of pull requests with any human review fell from 89% to 68%, while review by AI bots rose from 19% to 84%.1
| Pull requests | 3.1× |
|---|---|
| Reviewers | 1.5× |
Tired humans can't be the check. Green tests can't be the check.
When people stop reading, agents learn to satisfy tests rather than intent. In one benchmark, an agent's “C compiler” turned out to be a 2,900-line lookup table of the expected outputs.2
The idea
Changes should be understood as shifts in meaning, scored by risk, and approved by evidence rather than by someone reading lines.
Users never see code. They see behavior. So the source of truth moves to three things that outlive any version of the code, and agents can rewrite the rest freely.
- Contracts
- What each component promises: inputs, outputs, side effects and invariants. People own these.
- Evidence
- Tests, traces and benchmarks showing the promises hold. Machines produce it and re-run it constantly.
- The map
- How components relate: who calls whom, what data flows where, who owns what. Derived from code.
How a change moves through Onus
Agents propose. Machines check the facts. People see only what deserves their attention.
A live map, derived from code
Onus rebuilds a map of components, contracts, owners and risk labels on every change. Agents ask it questions instead of reading 50 files and guessing.
What depends on UserPreferences?notifications, billingOrderShippednewordersuser-preferencesnotificationsbillingSMS providerMeaning, not lines
Every change is described by what it does to the system. A rename across 40 files becomes one quiet row. A
>that became>=in a payments check goes to the top.- const total = formatOrderTotal(order);+ const total = formatTotal(order); …and 39 more files412 lines changed. Meaning changed: none.
Internal- if (subtotal > threshold) {+ if (subtotal >= threshold) { return applyDiscount(order);1 character changed. Every order at the threshold now gets a discount.
PaymentsAccess is earned with proof
Each task runs with a short-lived token that opens only the doors it needs. To get more, an agent attaches evidence. A structured way to escalate cut test-gaming from 23.6% to 5.3% across eight frontier models.3
Token for “Text customers when their order ships”
- Write
- notifications internals
- Read
- OrderShipped, UserPreferences
- Network
- SMS provider test endpoint
- Expires
- When the task ends
Escalation
Add an optional phoneVerified field to UserPreferences
- Failing test: texts an unverified number
- Map path: notifications reads UserPreferences
Every change gets a lane
Rules set the floor, and a model can only move a change up. Auth, payments and personal data go to a person, and so does any change to tests, CI or policy.
Auto-merge
Evidence checks alone
Internal refactorNew testsJudge
The verifying judge
Additive contract changeNew internal dependencyHuman
A person reads the report
New SMS vendor receives phone numbersBlocked
Nobody, until it's fixed
Write outside scopeTests weakenedA judge that verifies
It re-runs the evidence in a fresh environment and runs checks the author never saw. It never reads the author's reasoning, since models favor work that looks like their own.4
- Re-ran the evidence in a fresh environmentTests run from scratch; the author's results were not reused
- Intent matches effectFour changes in meaning, all named in the stated intent
- Contracts and rules holdUserPreferences change is additive; no boundary crossed
- No tests were weakenedNo deleted assertions, skips or special-cased inputs
- Held-out checks passedOpt-out and verification combined, which the author never saw
- TasteNaming and structure fit the codebase. Counts, but less.
Two moments of human attention.
The SMS feature, end to end. Like air traffic controllers, people set the route and step in when something unusual happens. Agents do the rest, and machines check it with evidence at every step.
- 1 IntentPerson
- 2 PlanPlanning agent
- 3 TokenToken service
- 4 BuildCoding agent
- 5 EscalateAgent and judge
- 6 ApplyCoding agent
- 7 RouteClassifier
- 8 VerifyJudge
- 9 ReviewPerson
- 10 Roll outProduction
Step 1, intent
“Text customers when their order ships. Respect opt-outs. Only text verified numbers.”
Step 9, review
Reads a one-page summary and approves the new vendor in minutes. Not 1,400 lines.
Any agent. Same rules.
Onus isn't an agent and doesn't pick one for you. Codex, Claude Code or one you built yourself: each gets the same scopes, lanes and review, and earns its own track record.
Speak MCP or run a CLI
Any agent that can run a shell command can use Onus. Agents that support MCP get the same tools natively.
Work inside its token
And ask for more through escalation, with evidence attached.
Submit evidence with every change
The stated intent, the semantic report, and the tests and traces behind it.
Swap every agent overnight and nothing in Onus changes. Only the track records start fresh.
Where we're starting.
Five phases. Each one is useful on its own, and each moves on only when a measurable gate is passed.
Phase 1
Designing nowSemantic change reports
A bot posts a meaning-level summary on every pull request. Big agent PRs become reviewable.
Gate to the next phase: Reviewers prefer the report to the raw diff on large pull requests.
Phase 2
The map as a service
Agents and people query the map through MCP and a CLI. Agents stop guessing.
Gate to the next phase: Map answers match reality on a sampled set of questions.
Phase 3
Scoped tokens and escalation
Per-task tokens enforced at the doors. A hijacked agent reaches only what its task needed.
Gate to the next phase: Escalations don't stall work, and no writes land out of scope.
Phase 4
Risk lanes and the judge
Rules-first routing; the judge re-runs the evidence. Attention concentrates on the risky slice.
Gate to the next phase: Audits of auto-approved changes find a low miss rate.
Phase 5
Evidence factory and production loop
Disposable environments by default; outcomes retrain the classifier. It improves with use.