1. one box, one job
  2. what the card can actually do
  3. eli5: the apprentice model
  4. four harnesses, one endpoint
  5. the Friday switch
  6. the bill
engineering11 min

The Workhorse Comes Home

A local inference box and the 28-second test that found a silent harness failure

The timer had reached two hours and forty minutes. The coding harness still looked alive. Its interface was updating, the agent was supposedly working, and not one of its ten audit lanes had finished.

The transcript ended halfway between two tool results. There was no error and no retry. It was the least useful kind of failure: the kind that continues to look like work.

We sent the same prompt, at the same context size, directly to the model endpoint. It answered in 28 seconds.6 The inference box was fine. The harness was not.

Same prompt, same context size
Coding harnessten audit lanesModel endpoint
still looking alive at 2 h 40 min: no lane finished, no error, no retry
Model endpoint
28 s

fault located in the harness, not the model server

Same prompt, with and without the harness. The harness had spent two hours and forty minutes looking busy. The model, asked the same thing directly at the same context size, answered in 28 seconds, which settled where the fault lived.

Diagram of the differential test. One prompt at one context size goes down two paths. The first runs through the coding harness, with its ten audit lanes, to the model endpoint: after two hours and forty minutes the harness still looked alive, not one lane had finished, there was no error and no retry, and the transcript ended halfway between two tool results. The second runs directly to the model endpoint and is answered in 28 seconds. On a shared time scale the 28-second bar is a sliver beside the two-hour-forty-minute one, drawn at a small minimum width so it stays visible. A closing line reads: fault located in the harness, not the model server.

That small test did more than locate a hang. It was one of the moments when the box stopped being a hardware project and became part of the workshop. By Friday it was carrying routine inference, four coding harnesses had audited the same frozen repository through it, and one of those harnesses had replaced our daily driver.

// one box, one job

In July the plan was to benchmark a Qwen coder model on an RTX 5090 before spending £2,894.99 on a unified-memory machine. That machine stayed on the wishlist until VRAM became a measured blocker rather than a suspicion.1

The benchmarks came back useful, so we built around the card we already understood. The result is an ordinary desktop tower with an RTX 5090 bought at retail. It can hold a 27-billion-parameter model and a 260K context slot at the same time. It is air cooled, starts on demand and has no other job.

That last part matters. Audit lanes and bulk reading can queue overnight. Interactive work can tolerate a ten-second cold start. The workload is bursty but patient, which makes it a better fit for owned hardware than our earlier experiments suggested.

The serving stack soon acquired a launcher and a doctor command. That is usually the point at which an experiment has become infrastructure, whether or not anybody remembers holding a ceremony.

Workshop hostkeeps the harnesses, memory and codeintelligenceInference boxordinary desktop tower, air cooled, starts ondemand, no other jobHarnessesMemoryCode intelligencellama.cpp serverbound to loopback, not exposedto the LAN, closes when thesession endsQwen3.8-27B at Q5_K_Mone 260K context slot, plus asmaller draft modelRTX 5090bought at retaildedicated SSHconnection
Harnesses stay home, weights live in the box. Harnesses, memory and code intelligence stay on the workshop host, and inference crosses one SSH connection to a server bound to loopback. One consumer profile and a doctor command are there to stop the tunnel quietly becoming permanent, as a June audit found one forwarder doing for fifty-one days.

The workshop host keeps the harnesses, memory and code intelligence. A dedicated SSH connection forwards inference requests to a llama.cpp server bound to loopback on the other machine. The endpoint is not exposed to the LAN, and it closes when the session ends. SSH is already the boundary between two single-user hosts, so adding a separate API credential would not improve it.

The risk we were designing around was slow configuration drift. A temporary tunnel acquires another consumer, then another profile, and eventually nobody remembers why it is permanent. A June audit had found a port forwarder which had survived for fifty-one days that way. This bridge has one consumer profile, starts when asked and has a doctor command which complains when reality differs from that description.

// what the card can actually do

The model is Qwen3.8-27B at Q5_K_M, with one 260K context slot. A smaller draft model guesses a short run of tokens and the larger model verifies them in one pass.2 When those guesses are good, generation gets much faster without changing the larger model's output.

tokens per second, log scale
Prefill, cold start680
roughly 680
Prefill, warm1,300 - 1,780
1,300 to 1,780
Generation, live conversation at 67,000 tokens of context72 - 88
72 to 88
Generation at 47% draft acceptance50
about 50
Generation at 67% draft acceptance109
about 109
Measured throughput on one RTX 5090. The draft model's hit rate is the main throttle: generation ran at about 50 tokens a second at 47 percent acceptance and about 109 at 67. Both points come from one evening's logs, so read them as a range rather than a law.

Range chart of measured throughput in tokens per second on a logarithmic axis. Prefill from a cold start: roughly 680. Prefill when warm: 1,300 to 1,780. Generation during a live conversation at 67,000 tokens of context: 72 to 88. Generation at 47 percent draft acceptance: about 50. Generation at 67 percent draft acceptance: about 109. The figures come from server logs, and the 47 and 67 percent points are from one evening.

At 47 percent draft acceptance we measured about 50 generated tokens a second. At 67 percent it reached about 109. Conversation prefixes are also reused by longest-common-prefix similarity, cutting repeat prompt evaluation by roughly four times during a working session.3

I initially wrote that prefill ran at 700 tokens a second. That was yesterday morning's cold-start window. The fuller log showed warm runs between 1,300 and 1,780. While this section was being revised, a live conversation at 67,000 tokens of context was generating at 72 to 88 tokens a second. It felt ordinary, which is a useful result for a machine assembled to make local inference stop feeling like a demonstration.

The model is not good enough to hold the centre of a long engineering job. We use it for bulk reading, independent review, bounded worker tasks and ordinary conversation. Frontier models keep authorship and final judgement. That division was a working assumption until the harness audit gave us something concrete to test it against.

// eli5: the apprentice model

If you just want the idea: imagine a careful craftsperson working beside a quick apprentice. The apprentice guesses the next few words before the craftsperson reaches them. Good guesses are kept and save time. Bad ones are crossed out. The finished answer still comes from the larger model.

Why run it at home? Much of this workshop's AI work is reading and filtering, and the card would otherwise be idle. Once the machine is already there, the extra bill is mostly electricity.

// four harnesses, one endpoint

Public benchmark results had not separated the harness candidates. The same-model estimates overlapped, and each had few runs.4 We therefore pinned one repository, checked its tree hash before and after every arm, and gave the same read-only brief to native Codex, oh-my-pi, Hermes and Qwen Code.

All four used the local box, one slot and one model. Each arm had to work without reading the others' output and bank only findings its coordinator could verify against exact files and lines. The comparison still had an important limit: the harnesses chose different lane policies and concurrency, so their exposure to stalls was not controlled. We treated recovery behaviour as operational evidence, not a score.5

ResultSupervision
CodexCompletedNo stall
oh-my-piAll 26 lanes completed, 52 findings bankedCoordinator corrected five citation errors before they reached the report, then recorded 45 files and about 5,800 lines outside the lane map
HermesSix system reports completedDevised its own worker watchdog
Qwen CodeFound real issuesCoordinator restart after the silent 2 h 40 min run; ten lanes starved against the one-slot endpoint; resumed on two lanes with a watchdog whose first version produced two false positives
Same endpoint, different supervision loads. All four audited the same frozen repository through the same one-slot endpoint. Lane policies and concurrency were left to each harness, so stall exposure was not controlled and the stalls count as operational evidence rather than a score.

Comparison grid of the four harnesses that audited one pinned repository through the same local endpoint, one slot and one model, from the same read-only brief. Codex: completed; no stall. oh-my-pi: completed all 26 lanes and banked 52 findings; its coordinator corrected five citation errors before they reached the report, then recorded that 45 files and about 5,800 lines had fallen outside the lane map. Hermes: completed six system reports; devised its own worker watchdog. Qwen Code: found real issues; its coordinator needed a restart after the silent two-hour-forty-minute run, ten lanes had starved against the one-slot endpoint, and the resumed run used two lanes and a watchdog whose first version produced two false positives.

oh-my-pi completed all 26 lanes and banked 52 findings. Its coordinator corrected five citation errors before they reached the report, then recorded that 45 files and about 5,800 lines had fallen outside the lane map. Hermes completed six system reports and devised its own worker watchdog. Codex completed without a stall.

Qwen Code found real issues too, but its coordinator needed a restart after the silent two-hour-forty-minute run. Ten lanes had also starved against a one-slot endpoint. The resumed run used two lanes and a watchdog, whose first version produced two false positives. That is messy evidence, but it is much more useful than pretending every harness experienced the same test.

Two other ideas failed during the same week. A viral harness offered self-refereed DeepSeek numbers which we could not reproduce. A proposal to improve review with logit arithmetic survived 504 calls and eighteen cents before failing on frozen gold data. Neither was adopted. This is normal. Most candidates are supposed to lose.

// the Friday switch

Our July note kept native Codex as the daily driver and treated oh-my-pi as a specialist pending a security preflight. The August audit changed that decision. oh-my-pi is now the daily harness.7

The swap took an afternoon because the surrounding contract did not move. Every harness reads the same instruction file. The launcher pins a release, verifies its checksum and retains the previous binary for rollback. Running sessions finish on the version they started with. The harness's self-updater is disabled because it would change the installation outside that boundary.

The model behind those daily sessions currently arrives through an anonymous router tier. There is no lab name or system card to lean on, so we treat it like any other route: its claims count only when they reproduce, and the ledger records what actually answered. A famous badge would not change that rule.

The audit chose the daily harness. The local endpoint made the comparison cheap enough to run and gave every candidate the same conditions. It now has a standing shift carrying similar audit and review traffic.

// the bill

The card draws about 450 watts under inference and the whole machine uses roughly half a kilowatt at the wall. At British electricity prices, that is about fifteen pence an hour under full load. Real sessions spend time below that.

Local inference is slower than the cheap hosted open-weight tier, but its marginal cost is lower for our duty cycle. The arithmetic depends on a box built for this workload and a card bought at retail. Anyone pricing the hardware against their own usage may quite reasonably conclude that renting remains cheaper.8

The unified-memory machine is still on the wishlist. The workload has not yet proved that it needs one. For now the older June box is waiting in the corner with a spare PCIe slot, as Linux boxes tend to do.


  1. The purchase gate remains unchanged: prove on real work that VRAM is the blocker before spending £2,894.99. A 27B Q5 model and the full 260K slot fit on the current card, so the evidence has not passed that gate. ↩

  2. The quant file is a community "Uncensored" build, a label for the quant-maker's modification rather than our workload. The matching Q4 draft uses a three-token horizon and eight-token blocks. The model family's licence permits this use. ↩

  3. Acceptance is the fraction of proposed draft tokens retained by the verifier. The figures come from server logs. The 47 and 67 percent points are from one evening, the prefill range spans a cold start at roughly 680 through warm runs at 1,780, and the live figures come from the session running during the rewrite. ↩

  4. The June Terminal-Bench snapshot reported Pi at 76.0 percent, oh-my-pi at 74.2, OpenCode at 72.7 and Codex at 71.2 on the same model. The intervals overlapped and the sample sizes were small. ↩

  5. The endpoint was shared, but lane policies, concurrency limits and watchdog behaviour were not. Only two arms were exposed to a stall. This is why the stall history informed operations without becoming a leaderboard column. ↩

  6. The differential test replayed one prompt at the coordinator's exact context size. It answered in 28 seconds, locating the fault in the harness rather than the model server. ↩

  7. The managed launcher verifies each pinned release, keeps the previous one available, and prevents the tool from updating itself. New sessions receive the new pin while existing sessions finish on the old binary. ↩

  8. Hardware amortisation depends on hosted rates and duty cycle. Our audits, review lanes and research sessions keep the card useful, but that does not make the same purchase sensible for another workshop. ↩

Back to blog