The Workhorse Comes Home
A local inference box and the 28-second test that found a silent harness failure
The timer had reached two hours and forty minutes. The coding harness still looked alive. Its interface was updating, the agent was supposedly working, and not one of its ten audit lanes had finished.
The transcript ended halfway between two tool results. There was no error and no retry. It was the least useful kind of failure: the kind that continues to look like work.
We sent the same prompt, at the same context size, directly to the model endpoint. It answered in 28 seconds.6 The inference box was fine. The harness was not.
fault located in the harness, not the model server
Diagram of the differential test. One prompt at one context size goes down two paths. The first runs through the coding harness, with its ten audit lanes, to the model endpoint: after two hours and forty minutes the harness still looked alive, not one lane had finished, there was no error and no retry, and the transcript ended halfway between two tool results. The second runs directly to the model endpoint and is answered in 28 seconds. On a shared time scale the 28-second bar is a sliver beside the two-hour-forty-minute one, drawn at a small minimum width so it stays visible. A closing line reads: fault located in the harness, not the model server.
That small test did more than locate a hang. It was one of the moments when the box stopped being a hardware project and became part of the workshop. By Friday it was carrying routine inference, four coding harnesses had audited the same frozen repository through it, and one of those harnesses had replaced our daily driver.
// one box, one job
In July the plan was to benchmark a Qwen coder model on an RTX 5090 before spending £2,894.99 on a unified-memory machine. That machine stayed on the wishlist until VRAM became a measured blocker rather than a suspicion.1
The benchmarks came back useful, so we built around the card we already understood. The result is an ordinary desktop tower with an RTX 5090 bought at retail. It can hold a 27-billion-parameter model and a 260K context slot at the same time. It is air cooled, starts on demand and has no other job.
That last part matters. Audit lanes and bulk reading can queue overnight. Interactive work can tolerate a ten-second cold start. The workload is bursty but patient, which makes it a better fit for owned hardware than our earlier experiments suggested.
The serving stack soon acquired a launcher and a doctor command. That is usually the point at which an experiment has become infrastructure, whether or not anybody remembers holding a ceremony.
The workshop host keeps the harnesses, memory and code intelligence. A dedicated SSH connection forwards inference requests to a llama.cpp server bound to loopback on the other machine. The endpoint is not exposed to the LAN, and it closes when the session ends. SSH is already the boundary between two single-user hosts, so adding a separate API credential would not improve it.
The risk we were designing around was slow configuration drift. A temporary tunnel acquires another consumer, then another profile, and eventually nobody remembers why it is permanent. A June audit had found a port forwarder which had survived for fifty-one days that way. This bridge has one consumer profile, starts when asked and has a doctor command which complains when reality differs from that description.
// what the card can actually do
The model is Qwen3.8-27B at Q5_K_M, with one 260K context slot. A smaller draft model guesses a short run of tokens and the larger model verifies them in one pass.2 When those guesses are good, generation gets much faster without changing the larger model's output.
Range chart of measured throughput in tokens per second on a logarithmic axis. Prefill from a cold start: roughly 680. Prefill when warm: 1,300 to 1,780. Generation during a live conversation at 67,000 tokens of context: 72 to 88. Generation at 47 percent draft acceptance: about 50. Generation at 67 percent draft acceptance: about 109. The figures come from server logs, and the 47 and 67 percent points are from one evening.
At 47 percent draft acceptance we measured about 50 generated tokens a second. At 67 percent it reached about 109. Conversation prefixes are also reused by longest-common-prefix similarity, cutting repeat prompt evaluation by roughly four times during a working session.3
I initially wrote that prefill ran at 700 tokens a second. That was yesterday morning's cold-start window. The fuller log showed warm runs between 1,300 and 1,780. While this section was being revised, a live conversation at 67,000 tokens of context was generating at 72 to 88 tokens a second. It felt ordinary, which is a useful result for a machine assembled to make local inference stop feeling like a demonstration.
The model is not good enough to hold the centre of a long engineering job. We use it for bulk reading, independent review, bounded worker tasks and ordinary conversation. Frontier models keep authorship and final judgement. That division was a working assumption until the harness audit gave us something concrete to test it against.
// eli5: the apprentice model
If you just want the idea: imagine a careful craftsperson working beside a quick apprentice. The apprentice guesses the next few words before the craftsperson reaches them. Good guesses are kept and save time. Bad ones are crossed out. The finished answer still comes from the larger model.
Why run it at home? Much of this workshop's AI work is reading and filtering, and the card would otherwise be idle. Once the machine is already there, the extra bill is mostly electricity.
// four harnesses, one endpoint
Public benchmark results had not separated the harness candidates. The same-model estimates overlapped, and each had few runs.4 We therefore pinned one repository, checked its tree hash before and after every arm, and gave the same read-only brief to native Codex, oh-my-pi, Hermes and Qwen Code.
All four used the local box, one slot and one model. Each arm had to work without reading the others' output and bank only findings its coordinator could verify against exact files and lines. The comparison still had an important limit: the harnesses chose different lane policies and concurrency, so their exposure to stalls was not controlled. We treated recovery behaviour as operational evidence, not a score.5
| Result | Supervision | |
|---|---|---|
| Codex | Completed | No stall |
| oh-my-pi | All 26 lanes completed, 52 findings banked | Coordinator corrected five citation errors before they reached the report, then recorded 45 files and about 5,800 lines outside the lane map |
| Hermes | Six system reports completed | Devised its own worker watchdog |
| Qwen Code | Found real issues | Coordinator restart after the silent 2 h 40 min run; ten lanes starved against the one-slot endpoint; resumed on two lanes with a watchdog whose first version produced two false positives |
Comparison grid of the four harnesses that audited one pinned repository through the same local endpoint, one slot and one model, from the same read-only brief. Codex: completed; no stall. oh-my-pi: completed all 26 lanes and banked 52 findings; its coordinator corrected five citation errors before they reached the report, then recorded that 45 files and about 5,800 lines had fallen outside the lane map. Hermes: completed six system reports; devised its own worker watchdog. Qwen Code: found real issues; its coordinator needed a restart after the silent two-hour-forty-minute run, ten lanes had starved against the one-slot endpoint, and the resumed run used two lanes and a watchdog whose first version produced two false positives.
oh-my-pi completed all 26 lanes and banked 52 findings. Its coordinator corrected five citation errors before they reached the report, then recorded that 45 files and about 5,800 lines had fallen outside the lane map. Hermes completed six system reports and devised its own worker watchdog. Codex completed without a stall.
Qwen Code found real issues too, but its coordinator needed a restart after the silent two-hour-forty-minute run. Ten lanes had also starved against a one-slot endpoint. The resumed run used two lanes and a watchdog, whose first version produced two false positives. That is messy evidence, but it is much more useful than pretending every harness experienced the same test.
Two other ideas failed during the same week. A viral harness offered self-refereed DeepSeek numbers which we could not reproduce. A proposal to improve review with logit arithmetic survived 504 calls and eighteen cents before failing on frozen gold data. Neither was adopted. This is normal. Most candidates are supposed to lose.
// the Friday switch
Our July note kept native Codex as the daily driver and treated oh-my-pi as a specialist pending a security preflight. The August audit changed that decision. oh-my-pi is now the daily harness.7
The swap took an afternoon because the surrounding contract did not move. Every harness reads the same instruction file. The launcher pins a release, verifies its checksum and retains the previous binary for rollback. Running sessions finish on the version they started with. The harness's self-updater is disabled because it would change the installation outside that boundary.
The model behind those daily sessions currently arrives through an anonymous router tier. There is no lab name or system card to lean on, so we treat it like any other route: its claims count only when they reproduce, and the ledger records what actually answered. A famous badge would not change that rule.
The audit chose the daily harness. The local endpoint made the comparison cheap enough to run and gave every candidate the same conditions. It now has a standing shift carrying similar audit and review traffic.
// the bill
The card draws about 450 watts under inference and the whole machine uses roughly half a kilowatt at the wall. At British electricity prices, that is about fifteen pence an hour under full load. Real sessions spend time below that.
Local inference is slower than the cheap hosted open-weight tier, but its marginal cost is lower for our duty cycle. The arithmetic depends on a box built for this workload and a card bought at retail. Anyone pricing the hardware against their own usage may quite reasonably conclude that renting remains cheaper.8
The unified-memory machine is still on the wishlist. The workload has not yet proved that it needs one. For now the older June box is waiting in the corner with a spare PCIe slot, as Linux boxes tend to do.
The purchase gate remains unchanged: prove on real work that VRAM is the blocker before spending £2,894.99. A 27B Q5 model and the full 260K slot fit on the current card, so the evidence has not passed that gate. ↩
The quant file is a community "Uncensored" build, a label for the quant-maker's modification rather than our workload. The matching Q4 draft uses a three-token horizon and eight-token blocks. The model family's licence permits this use. ↩
Acceptance is the fraction of proposed draft tokens retained by the verifier. The figures come from server logs. The 47 and 67 percent points are from one evening, the prefill range spans a cold start at roughly 680 through warm runs at 1,780, and the live figures come from the session running during the rewrite. ↩
The June Terminal-Bench snapshot reported Pi at 76.0 percent, oh-my-pi at 74.2, OpenCode at 72.7 and Codex at 71.2 on the same model. The intervals overlapped and the sample sizes were small. ↩
The endpoint was shared, but lane policies, concurrency limits and watchdog behaviour were not. Only two arms were exposed to a stall. This is why the stall history informed operations without becoming a leaderboard column. ↩
The differential test replayed one prompt at the coordinator's exact context size. It answered in 28 seconds, locating the fault in the harness rather than the model server. ↩
The managed launcher verifies each pinned release, keeps the previous one available, and prevents the tool from updating itself. New sessions receive the new pin while existing sessions finish on the old binary. ↩
Hardware amortisation depends on hosted rates and duty cycle. Our audits, review lanes and research sessions keep the card useful, but that does not make the same purchase sensible for another workshop. ↩