1. agents read for a living
  2. eli5: why the biggest model is mostly a teacher
  3. students get cheaper
  4. the price gap
  5. where the token margins go
  6. reasons this could take longer
  7. three bets
research11 min

The Cloud Was a Phase

What a local inference box suggests about the falling cost of capable models

By Friday evening, a 27-billion-parameter model was answering routine workshop requests from a box on the other end of an SSH lead. It was slower than the cheap hosted routes but good enough for the work we gave it. At full load, the whole machine cost about fifteen pence an hour to run.

The practical story is in The Workhorse Comes Home: the hardware, the measured throughput, and the four harnesses we ran through one local endpoint.1 This article starts after that decision.

This one desk cannot establish a market trend. It did force us to look again at an assumption which had gone unchallenged for years: useful compute belongs in the cloud because utilisation makes renting cheaper than ownership.

That remains true for spiky demand, uncertain growth and machines nobody wants to maintain. We spent a decade moving servers, databases and queues upstairs for good reasons. AI arrived there by default. Then agent workloads started behaving less like ordinary infrastructure and more like a bill which grows whenever somebody has another idea.

// agents read for a living

An agent spends far more time reading than writing. Production traces have reported input-to-output ratios around a hundred to one. Instrumented academic traces range from the mid-fifties into the hundreds, with the loop still dominated by decoding whenever a prompt falls out of cache.2

Fresh sessions, child tasks and compaction all disturb the cache. A rented fleet is excellent at absorbing cold demand, while a single local slot performs best when prefixes stay warm. That makes the correct route depend on the workload. We measure it rather than treating either location as a principle.

The scale of the bill changes the conversation. Our own records include five figures of API use across a long weekend and four figures in one evening. Published enterprise telemetry puts average daily use in double-digit dollars per developer, and at least one company has capped individual engineers at $1,500 a month after exhausting its AI budget in four months.3

Frontier inference is worth its price on some tasks. Routine filtering and review consume the same metered tokens unless somebody gives them a cheaper route.

Owning the weights also removes one familiar problem. A provider can change the model, the price or the terms on Friday night. The number already written to your disk is still there on Monday.

// eli5: why the biggest model is mostly a teacher

If you just want the idea: a very large model is expensive to build and run. A smaller model can study its answers and learn to give similar ones, rather like a sous-chef learning the menu from the chef who created it. The student will not invent the next cuisine, but it can cook many familiar dishes on much cheaper equipment.

// students get cheaper

The largest models remain useful because they can do things smaller models cannot. They also teach.

Distillation trains a smaller model to imitate the outputs of a larger one. The student does not reproduce the teacher perfectly, but it can carry a large share of ordinary work at a fraction of the serving cost. Once the weights are released, that capability becomes a file which can be copied and run on owned hardware.

Students, roughly 27B to 70Bonce released, a file that can be copied andrun on owned hardwareFrontier teacherdoes what smaller models cannot;expensive to build and run27Bfits on the current cardDense 70Beasier to serve predictably inracksSparse, about 5B activereads several times faster thana dense 70B on the same memoryExpensive fleetcapacity left over once routinereading, filtering and boundedreview move away from premiumservingTrains the next model ormakes longer researchjobs affordableUnwanted depreciationdistillation:imitate itsoutputsroutine work movesthe nextcapability worthdistilling

the actual split is unknowable from one desk

One teacher, several cheaper students. Distillation carries part of the frontier's ability into students that run on owned hardware. The fleet they relieve trains the next model, makes longer research affordable or simply depreciates, in proportions nobody can see from one desk.

Diagram of distillation and its knock-on effect. At the top is the frontier teacher, which can do things smaller models cannot and is expensive to build and run. Distillation, training a smaller model to imitate its outputs, fans out to three students in the roughly 27B to 70B weight class, each of which, once released, is a file that can be copied and run on owned hardware: a 27B model that fits on the current card; a dense 70B, easier to serve predictably in racks; and a sparse model with about five billion active parameters that reads several times faster than a dense 70B on the same memory. Below, as routine reading, filtering and bounded review move to the students, the expensive fleet has capacity left over. Some of it trains the next model or makes longer research jobs affordable, and a dashed line runs from there back to the teacher, labelled the next capability worth distilling. Some becomes unwanted depreciation. A note says the actual split is unknowable from one desk.

Four agent-grade models shipped inside seventeen days this spring. The cheapest listed price was well below a dollar per million tokens, and most traffic on openly routed services already uses open weights.4 A frontier-class model released under an MIT licence in June could run on one node at perhaps a fifth of Western rental rates. In one security exercise it beat an unaided commercial agent, although the scaffolding mattered more than either model.

Our local result sits at the small end of the same curve. The 27B model held a useful lane in a four-harness audit backed by the same endpoint. Final judgement remained with a frontier model.

Mixture-of-experts models make this more practical. Their total parameter count influences what they can know, while only part of the model is active for each token. A sparse model with about five billion active parameters can read several times faster than a dense 70B on the same memory. Dense models remain easier to serve predictably in racks. For a desk, the useful range is currently a weight class rather than one winner: roughly 27B to 70B, sparse or dense depending on the hardware.

// the price gap

Fixed capability has been getting roughly five to ten times cheaper each year, while the newest frontier tasks have become more expensive.5 The estimates have wide error bars, but the direction has held throughout the period we have tracked it.

Fixed capability
Frontier task
no change10× cheaper5× cheaper3× dearer18× dearer

price change per year, log scale

Fixed capability gets cheaper, the frontier gets dearer. Fixed capability gets five to ten times cheaper a year while the frontier task gets three to eighteen times more expensive. Both are estimates with wide error bars, but the direction has held for as long as we have tracked it.

Diverging range chart on a logarithmic axis of price change per year, centred on a line marked no change. To the left, fixed capability spans five to ten times cheaper per year. To the right, the frontier task spans three to eighteen times more expensive per year. Both bars have soft ends because both rates are estimates with error bars. The chart shows annual rates only, with no absolute prices and no time axis.

Suppose routine reading, filtering and bounded review move away from premium serving. The expensive fleet has capacity left over. Some of it trains the next model or makes longer research jobs affordable. Some of it becomes unwanted depreciation. The actual split is unknowable from one desk, but the pressure is easy to see: cheap hardware can now perform work that recently required scarce hardware.

The frontier premium survives wherever the quality gap is worth paying for. In this workshop that means authorship, orchestration and decisions where another round can prevent an expensive mistake. Local models handle more of the surrounding volume. The boundary moves each time a student improves.

// where the token margins go

Reconstructed inference margins at the major labs have risen from below forty percent to about seventy percent while unit costs fell. Two large labs filed to go public this summer at annualised revenue measured in tens of billions, around the same time reporting said they were considering price cuts.6

Those facts can coexist while frontier access is scarce. They become harder to reconcile when last year's quality arrives as downloadable weights on a regular schedule. Serving then looks more like hosting: cards, electricity, reliability and geography.

The businesses around the model are harder to copy. Distribution matters. So do integration, trust, compliance and the systems that turn a model response into checked work. The teacher pipeline matters most of all: data, evaluations and training environments which produce the next capability worth distilling.

The labs may become foundries, paid to produce the frontier and to sell outcomes which smaller models cannot yet deliver. That can be a good business. It is different from assuming today's premium token volume continues indefinitely.

// reasons this could take longer

Memory is the immediate hardware constraint. DRAM contract prices rose by more than half in one quarter, and relief is not expected until late 2027. A 27B model fits on the current card. Most 70B models do not.

The remaining capability gap is concentrated in long-running orchestration and judgement. Those are expensive workloads, and they are precisely where frontier models still earn the premium here.

Self-hosting also returns the pager. Quantisation, context slots, drivers and thermals all become somebody's problem. Homelabs and small technical teams can accept that sooner than a large organisation. Many will continue renting because avoiding that work is the product they wanted.

Compliance may keep sensitive or heavily certified work with large providers even when the raw inference is cheaper elsewhere.

These constraints affect how quickly work moves and how much of it moves. They are also why our current setup is hybrid rather than a declaration of independence from the cloud.

// three bets

Three dated bets, so the future can grade us: within twelve months, a 27-to-70B open-weight student passes a blind worker-selection round here on quality, not just cost - a full seat, not a spare-room curiosity. Within twenty-four, most of this workshop's token volume originates on metal we control or open routes we could control. And within the same window, at least one major lab sells outcomes - verified work, orchestrated and reviewed - rather than tokens, because the token business will no longer clear the multiple. If the third bet is wrong, it will be because the second one stayed small. If both come true, the phrase "cloud AI" will read the way "cloud storage" reads now: accurate, historical, and describing where most of the value isn't.

  1. An open-weight student earns a full seat

    a 27-to-70B open-weight student passes a blind worker-selection round here on quality, not just cost

  2. Most token volume on metal we control

    or on open routes we could control

  3. A major lab sells outcomes, not tokens

    verified work, orchestrated and reviewed; if this bet is wrong, it will be because the second stayed small

Three bets, with deadlines. Dated so the future can grade them, and not independent: if the third bet fails, it will be because the second stayed small.

Timeline of three dated bets, measured from publication. Within twelve months: a 27-to-70B open-weight student passes a blind worker-selection round in this workshop on quality, not just cost, earning a full seat rather than a spare-room curiosity. Within twenty-four months: most of this workshop's token volume originates on metal it controls or open routes it could control. Within the same twenty-four-month window: at least one major lab sells outcomes, meaning verified, orchestrated and reviewed work, rather than tokens. A note on the third bet says that if it is wrong, it will be because the second one stayed small.


  1. The Workhorse Comes Home covers the bridge architecture, measured throughput and the four-harness audit from the same week. ↩

  2. Sources carried forward from our June workload analysis include Manus's production figure of roughly 100:1 input to output and a May 2026 trace of coding agents across open 27B to 31B models. The latter reported mean ratios from about 54:1 upward, warm-cache hit rates of 84.6 to 99.5 percent and decode-bound wall clock. ↩

  3. Our billing figures were reconstructed from session transcripts. The enterprise averages came from Anthropic's published telemetry. The capped company said on the record that the connection between spend and return remained unproven while about seventy percent of its committed code was machine-written. ↩

  4. The April cluster contained four agent-grade models in seventeen days. The cheapest listed price was $0.435 per million input tokens and $0.87 per million output tokens, with cached input below a tenth of a cent. Individual vendor claims remain thinly reproduced, so the argument does not depend on one leaderboard row. ↩

  5. The fixed-capability estimate is five to ten times cheaper per year, while the frontier task becomes three to eighteen times more expensive. Both are estimates with error bars. ↩

  6. SemiAnalysis supplied the margin reconstruction. The confidential filings and reported pre-IPO price discussions were covered in The Most Expensive Houseguest. ↩

Back to blog