Nobody Marks Their Own Homework
How our agent harness is built, and what it would take to build another.
At half past three this morning a coding agent was auditing the harness it works in, looking for things to improve. It noticed that the nightly backup worked from a hand-written list of repositories, and a list like that goes stale. So it fixed it.
It wrote a replacement that finds the repositories on every run, about thirty lines of shell, and tested it on a scratch copy. Nobody asked for a second opinion. The harness's rules say an author is never its own only reviewer, so the agent handed its diff to a second agent that had none of its context and no permission to edit anything, with instructions to find a state in which the script would lose data without saying so.
The reviewer found one in about four minutes: a file that had been staged but not yet committed would have ended up in neither half of the backup. The author did not take its word for that. It built a throwaway repository, staged a file, ran the script and looked for the file in the archive, and the file was missing. The repair went in twenty minutes before the four o'clock run. Afterwards the agent compared the offsite copy with the local archive by checksum and restored two repositories from the archive.
No human wrote, reviewed or corrected any of it. An agent improved the system it runs in, checked its own work the way the system demands, and left a record. That is what we mean by a harness, and it is why we think the interesting engineering in this field has moved out of the models and into what stands around them.
// agent and harness
A model predicts text. A coding agent is a model running in a loop with tools: it reads files, runs commands, edits code, looks at what happened and decides what to do next, for minutes or hours inside one session. Off-the-shelf programs such as Codex and Claude Code run that loop, and the word harness is often used for those programs alone.
We use it for something wider. Our harness is everything around the agents that makes their work useful and accountable: how a task gets its own working copy, which rules every agent reads, where decisions are remembered, who reviews a change, how a release is undone, and where the record of all that is kept. Earlier posts called it the workshop. It runs current frontier models from multiple vendors within custom coding agents, and the agents maintain much of it themselves.
// the rule
Nobody marks their own homework. The rule applies at three levels: to a change, to a reviewer's finding, and to the harness itself. The rest of this post takes them in that order and then lists what it would take to build the same thing.
// a change
The harness has seats: an orchestrator, an independent worker, a verifier and an advisor. The seats are fixed and the models in them are not. Frontier models change so quickly that we benchmark each new one on our own tasks when it arrives and move it into whichever seat it earns. Sol has orchestrated, and so have Astra, Fable and Opus. Swapping who orchestrates, who codes and who checks is a change of configuration.
At the time of writing Opus 5.5 orchestrates, and it was the agent in the story above. Earlier this week the seat was Sol's. The orchestrator states the outcome, the acceptance criteria and what is out of scope, and opens an isolated working copy for the task. The code is written by Sol while Terra, in a separate context, examines the contract and designs tests meant to break it. Separable pieces go to Terra workers with explicit ownership.
Tests run before any model is asked for an opinion. Then the exact candidate goes to Fable, a different vendor's model, in a fresh session with no edit tools and no memory of earlier rounds. Anything non-trivial also gets a blind review from Luna. The reviewers do not see each other's findings until they are adjudicated.
A confirmed defect is repaired by a fresh worker, and the repaired candidate is reviewed again from the start, because an approval belongs to exact bytes and a repair changes them. The first seating came out of a 608-row tournament, and the method is published.
// eli5: the second signature
If you just want the idea: a bank does not let the person who writes a cheque also approve it. Our agents work the same way, with one addition. The second signatory cannot simply object. They have to show the problem happening before the cheque is stopped.
// a finding
A reviewer's finding is a hypothesis. The orchestrator has to reproduce it against the real code before anything is repaired, and the result is written down either way.
That rule is why we can say how good the reviewers are. Every review attempt is appended to a ledger with its route, whether it completed, what it claimed, and what happened when each claim was tested. On 1 October the ledger had 5,255 entries covering about 1.9 billion reported tokens of review.
| Attempts | Claims | Reproduced | Refuted | |
|---|---|---|---|---|
| Luna, verifier | 1,060 | 1,087 | 833 | 161 |
| Terra, independent worker | 707 | 560 | 435 | 79 |
| Opus 5, summer advisor | 558 | 1,641 | 1,231 | 155 |
Table of three reviewer routes against attempts, claims made, claims reproduced and claims refuted. Luna, the usual verifier: 1,060 attempts, 1,087 claims, 833 reproduced, 161 refuted. Terra, the independent worker: 707 attempts, 560 claims, 435 reproduced, 79 refuted. Opus 5, the standing advisor for most of the summer: 558 attempts, 1,641 claims, 1,231 reproduced, 155 refuted.
Somewhere between one claim in nine and one in six fails when somebody tries to make it happen. A workflow that repairs whatever its reviewer says is acting on those as well, and each such repair is a change nobody asked for.
These counts describe practice. The routes were given different work, some claims were never adjudicated, and a ledger of findings has no row for the defects every reviewer missed. We do not know of another public record of reviewer precision kept this way. If you keep one, we would like to compare notes.
Architecture decisions go to a panel of six models, voting blind. Since April its seats have been held by models from nine vendors, and the line-up changes as the models do. Its findings are held to the same standard. Of its findings on code that have been tested, 174 were confirmed by reproduction and 39 were refuted.
// what the agents share
Several custom coding agents work in the same workspace, so the things they share are files. One contract of working rules is read by every agent at the start of a task. Memory is plain Markdown, 534 pages of decisions and corrections, with a search index derived from it, so no model's private recall is the authority. A code graph answers what calls a function and what breaks if it changes.
Each task gets its own working copy and ends with a written record saying whether it is complete or retained, by whom and until when. Temporary servers and test jobs stop when their session does. Services run from versioned releases with the previous one kept beside a rollback receipt. When the memory index had to move across 43 schema versions, the migration was rehearsed forwards and back again on a restored copy before the live database was touched, and every page came through with its content hash unchanged.
Because the rules and the memory live in files and not inside any one model, a model can be replaced without the harness forgetting anything. Routine work can run on a local model on our own hardware, which keeps the frontier models for the work that needs them.
// the harness itself
The third level is the harness improving itself, and it is the one we added last. On 29 and 30 September the model then in the orchestrator's seat audited the harness, repaired what it found and scored the result 8.5 out of ten. On 1 October we gave the same question to a different vendor's model in a different coding agent and told it to collect its own evidence.
It gave eight. It agreed on the nines for coding and for review. It also found things the first auditor could not have seen from where it stood, among them that two of the coding agents had drifted onto different versions of the code graph within a day of a careful release. The agent that found them repaired them the same night, and the backup change at the top of this post was one of those repairs.
One finding has no fix yet. The rule book grows a little with every incident, and it is the first thing every agent reads.
// what it would take
None of this is secret, and none of it depends on a model that we have and you don't. The models are the easy part. We have had a new one in service two days after it existed and lost another to a government on a Friday night, and the harness carried on both times.
| What it is | What it cost us | |
|---|---|---|
| Two vendors, fixed seats | The writer and the reviewer are never the same model; who sits where changes with the models | A 608-row tournament the first time, and a benchmark for each new arrival |
| Reproduce before repair | A finding is a hypothesis until it is made to happen | Tokens, on every change |
| A ledger from the first day | Every attempt, claim and outcome, appended | Hard to fill in afterwards; ours had no outcomes until June |
| Isolation and releases | A working copy per task, versioned releases, rehearsed migrations | Ordinary engineering, done every time |
| Memory as files | Markdown is the authority; the index is derived | The discipline to write corrections down |
| An outside auditor | A model that did not do the work marks it | One evening, and the willingness to read the result |
Table of six parts of the harness with what each is and what it cost. Two vendors in fixed seats: the writer and the reviewer are never the same model, and who sits where changes with the models; it cost a 608-row tournament the first time and a benchmark for each new arrival. Reproduce before repair: a finding is a hypothesis until it is made to happen; it costs tokens on every change. A ledger from the first day: every attempt, claim and outcome is appended; it is hard to fill in afterwards, and ours had no outcomes until June. Isolation and releases: a working copy per task, versioned releases and rehearsed migrations; ordinary engineering, done every time. Memory as files: Markdown is the authority and the index is derived; it costs the discipline to write corrections down. An outside auditor: a model that did not do the work marks it; it costs one evening and the willingness to read the result.
The expensive parts are time and tokens. A ledger is hard to fill in afterwards. Our panel had been voting since April, and in June its record still held votes and not one confirmed outcome. It has 174 now because recording the outcome became part of the job. Reproduction costs tokens every day, and the harness's spend reflects that. Most of the machinery was built in answer to a specific failure, and the failures arrived on their own schedule, starting in March.
I have run the same kinds of task through a clean setup, one model with nothing around it, and it was a disaster. Inside the harness, two deliveries in three go through in one shot. Where the ledger records it, 1,372 of 2,117 reviewed deliveries needed no repair at all. The other third were caught by the steps above before they were released, which is the harness doing its job.
The models still invent things. In the ledger that is the refuted column. What the harness changes is where an invention stops: at the reproduction step, before it has become a change to anything. No hallucination has reached a code release.
The models in this post will have been replaced by the time anyone builds this. The rule about homework can stay as it is.