1. memory before exploration
  2. recursive exploration when exploration is needed
  3. sequential dispatch when work has dependencies
  4. mission, not role
  5. verification before completion claim
  6. html for read-once human artefacts
  7. what the negative space is showing
ai11 min

What I Stopped Doing With Claude

Two months ago none of the six reflexes below felt wrong. They were defaults for reasons that used to be good. Someone influential said them in 2024. The demos all showed them. The model that existed when you first wrote your workflow rewarded them. Then evidence arrived in batches, and the defaults turned into mistakes. Five of them stopped happening this spring because of evidence: papers, audits, sessions that failed loudly, or whole workflows that we ran in production and then retired. The sixth had never been a reflex here at all. It was the local default from the first weeks of this work last winter, and only acquired public reasons in the last fortnight.

The way I work with Claude across one task has about six steps. Eight weeks ago, at every one of them, I was reaching for a default without thinking. Today, at every one of them, I'm not. Here is the walk, in order.

The shape of those last two categories, the local-then-public and the simultaneous-arrival, is going to keep happening. In this corner of work, the space of working answers about how to use these tools is small enough that arrivals at the same answer from different starting lines will become routine. When something quietly works on your stack, expect company in the literature within months. The good answers are too few in number for everyone to have a different one.

// memory before exploration

The default, two months ago: drop into a session, get a question, start grepping. Pull files. Build the mental model from scratch. Once enough context was in the window to answer, do the answering.

The cost of that default added up before anyone noticed. One codebase alone burned roughly half a billion subagent tokens in ten days during a phase of heavy work, most of it spent re-deriving knowledge that was already written down. In the CLAUDE.md file. In the planning directory. In the memory wiki. One particularly enthusiastic exploration of the same codebase ran to 11.2 million tokens in a single session. The agent had been told nothing it could not have looked up. It went and read the codebase from the leaves up anyway, because that was the reflex.

The replacement is small. Check memory first. The memory wiki is markdown. There is a search tool in front of it that ranks results by salience. Anything I have learned about this project, this user, this dependency, this past incident, has a high chance of being in there. The Read tool is for what the memory does not cover. The grep is for what neither the memory nor the read can answer.

The numbers shifted overnight. The same project, post-rule, on a typical session-resume primer: about five thousand tokens of compressed context, generated by a script and pulled in by name. Three orders of magnitude less than the flat-exploration cost. The answers don't get worse, because the substrate was carrying the knowledge whether anyone read it or not.

1k10k100k1M10M100Mthree orders of magnitudeabout 5,000session-resume primer, memory first11.2 millionone flat exploration, single session
Tokens spent building context. The same codebase, read from the leaves up in one session, then resumed later from a primer pulled in by name. The 11.2 million was one particularly enthusiastic session, so treat the gap as a ceiling rather than an average.

Dot plot on a logarithmic axis running from 1,000 to 100 million tokens. On the same codebase, one flat exploration session used 11.2 million tokens, marked in amber towards the right. A typical session-resume primer under the memory-first rule used about 5,000 tokens of compressed context, marked in cyan towards the left. A bracket between the two dots is labelled 'three orders of magnitude'.

The reflex this prunes is the one that says "to be useful, I should re-understand." A capable model with a good memory layer does not need to re-understand. It needs to know where the existing understanding is and how to refresh only the parts that moved.

// recursive exploration when exploration is needed

Some sessions still have no memory to start from. The codebase has to be walked from scratch. The default for those, two months ago, was grep and read and grep and read, until the picture cohered or the context window ran out, whichever came first.

In January, a paper called Recursive Language Models arrived. Zhang, Kraska, and Khattab from MIT CSAIL, arXiv 2512.24601. The technique is one of those that sounds obvious once stated. Treat the long prompt as an external environment, not as raw text in the context window. Have the model write code that chunks, filters, and queries itself recursively, via an lm_query function it can call on itself. The model carries pointers to the data, not the data. Inputs a hundred times larger than the context window become navigable. Eight-billion-parameter models fine-tuned on a thousand examples approach GPT-5 on long-context benchmarks. Up to twice the accuracy at the same scale.

A skill called /rlm-explore shipped on the local stack in April with the paper's pattern adapted to Claude Code. Audits and cross-cutting analyses now decompose into focused sub-agents that return compact briefs to the main session. The main context window stays clean. The token cost stays bounded. The answers get better, because each sub-agent is reasoning over a tighter slice than a flat exploration would have given it.

What this prunes is the reflex that says "I will simply look at all of it." Most of the time that reflex was wasting tokens, blowing out context, and producing worse answers than a structured decomposition would have. Sometimes it was the right move, for genuinely tiny problems. For everything above that scale, the recursive approach now wins by such a wide margin that the old default is hard to defend.

// sequential dispatch when work has dependencies

For the last eighteen months, the Claude Code marketing story has been parallelism. Spawn agents in parallel. Run independent lookups concurrently. Fan out, gather results, synthesise. The Task tool's parallel-call pattern is everywhere in the documentation. Every demo deck mentions it.

It turns out to be wrong, often, when the work has dependencies.

In March, a paper from MIPT (Dochkina, arXiv 2603.28990) ran 25,000 tasks across eight models and four coordination protocols. The protocols were fully parallel, central coordinator, debate, and what the paper called Sequential. The Sequential protocol gave each agent a fixed ordering and, crucially, passed each agent the completed outputs of all predecessors. Not summaries of intent. Not plans. The actual outputs of work already done.

Sequential beat fully-parallel by forty-four percent, Cohen's d of 1.86. It beat the central coordinator by fourteen percent. The spread between the best and worst protocol on identical models was larger than the spread between the best and worst models on identical protocols. The way you coordinate the agents matters more than which agents you have.

Fully parallelthe parallel-by-default reflexSequentialbeat fully parallel by 44% (Cohen's d 1.86)and a central coordinator by 14%taskagentagentagenttaskagent 1agent 2agent 3completed outputcompleted outputsall predecessors
Fan out, or pass the work forward. For eighteen months the demos said fan out; across 25,000 tasks, passing finished work forward beat it by 44 percent. The agent counts drawn here are schematic.

The paper's explanation is one of those analogies that does not feel like an analogy. A sports draft. Each pick is informed by every previous pick. The team naturally fills complementary roles without anyone coordinating it. Parallel selection produces a roster of nine point guards.

The change locally: any task that decomposes into four or more dependent steps now goes sequential, with each step's completed output passed forward to the next. Parallel survives where it actually fits, which is independent lookups across unrelated files, multi-grep across orthogonal topics, anything where step N would not benefit from seeing step N minus one. That is a narrower set of cases than the parallel-by-default reflex assumed. Maybe a third of multi-agent work, no more.

// mission, not role

Mission-not-role is the example on this list with the tightest parallel-discovery story. The local change happened on March 25th, when we retired Anthill, our ten-agent orchestration framework, because specialised agents had turned out to be one of the things wrong with it. The MIPT paper that made the case for mission-not-role across 25,000 tasks appeared in the same March, from a research group studying coordination protocols. Same month. Two completely different starting points.

A great deal of prompt engineering advice from 2023 and 2024 told you to open by telling the model what role it was playing. "You are a senior security engineer." "Act as a refactoring specialist." "You are an expert Python developer." The shape was that you assigned identity, and identity assigned approach.

It never sat right. The thing about giving a capable model a role is that you are constraining the very judgement you are paying it for. If the model is smart enough to plan its own approach to a problem, telling it that it is a refactoring specialist when the task might benefit from being looked at as a testing problem is taking value off the table. So locally, since the day after Anthill was retired, the pattern has been mission plus facts. State what needs to be true at the end. State the file paths, the observed behavior, the constraints, the existing patterns. Then leave the angle to the model. The model picks the angle. Sometimes it picks one you would not have.

Role2023 and 2024 adviceMission plus factshere since Anthill was retiredyouan identity"You are a senior securityengineer."the approachyoumission plus factswhat needs to be true at theend; file paths, observedbehaviour, constraints, existingpatternsthe modelthe anglesometimes one you would not haveassignassignsstateleave the angle topicks
Who picks the angle. The switch here came with Anthill's retirement in March. In the MIPT study, agents left to it invented 5,006 role descriptions of their own across 25,000 tasks.

The MIPT paper, the same one as above, included a finding about exactly this. The best-performing protocols gave agents a mission, an ordering, and zero role assignments. Across the 25,000 tasks, eight agents collectively invented 5,006 unique role descriptions, voluntarily, one per task per agent on average. The paper computed a Role Stability Index that measured how much each agent reverted to a fixed self-description from task to task. In the best configurations it converged to roughly zero. The agents were reinventing their specialisation every task.

This was confirmation arriving at the same moment as the practice on this stack, not afterwards. Anthill went away on March 25th. The MIPT paper appeared in March too. The practice and the research that backed it showed up from completely independent starting points within weeks of each other, which is the rate this post is going to suggest is becoming the norm in this corner of work.

// verification before completion claim

The default, two months ago: implement the change, look at the code, conclude it should work, claim done. Move on.

The failure mode was specific. The code would type-check. The build would pass. The test suite would go green. The actual feature would be broken in some way that the existing tests did not catch. Once a week or so this would be discovered hours later, after the next branch had been started, by which point untangling what changed required a small archaeological expedition.

The new rule is small and dull and works. Every completion claim now requires running the verification command and reading the output before any "done" is said out loud. If the change is "the calendar now shows blocked times in red", the verification is opening the calendar, looking at it, and confirming that yes, blocked times now appear in red. If the change is "the migration adds a column with a backfill default", the verification is the migration log line confirming the backfill ran on the expected row count. Not "should". Not "looks right". Read the output. Then claim.

It sounds trivial. It is trivial. It is also, by some margin, the thing that has most reduced rework on this stack in the last two months. There is a one-line feedback rule in the memory wiki that exists only to keep this discipline in front of every session. It fires before every claim of completion. It has not produced a false positive yet.

The reflex this prunes is the one that says "the code compiles and the type-checker is happy, therefore the feature works." Type-checkers and compilers verify code correctness. They do not verify feature correctness. The only thing that verifies feature correctness is using the feature.

// html for read-once human artefacts

This is the one reflex on the list that was never an issue here. The default was set correctly from the first weeks of this work last winter, in private, and only acquired public reasons in the last fortnight.

The default for AI-generated reports for the last several years has been markdown. Markdown is the lingua franca of LLM output. Every chat interface renders it natively. Every documentation toolchain ingests it. Every model produces it by default. Markdown is correct, often, for the things the model wants to produce.

It is not correct, often, for the things a human is going to read once and then either link, share, or close. A long markdown report becomes a vertical wall. Tables in markdown lose information density. Diagrams in markdown become ASCII boxes. There is no styling. No tabs. No interactive widgets. No SVG. Reading a hundred-and-fifty-line markdown report on a phone is its own particular punishment.

The local pattern that emerged, without anyone deciding it, was that anything a human was going to read once got generated as HTML. The reports directory is full of HTML files. Equity research. SEO audits. Postmortems. The strategy section of the nullhex site is HTML. The ops dashboard is HTML. The council dashboard is HTML. The decision-record directories for the substrate overhaul have fifteen or twenty HTML iterations apiece. The pattern had been running quietly for months. We had never thought to share it. It was an in-house thing to do, not the sort of practice that felt worth a write-up.

On May 8th, Thariq Shihipar wrote a tweet that started with the sentence "HTML is the new markdown" and made the case for the pattern in public. Three days later Karpathy quote-tweeted it. The discussion that followed had the slightly disorienting quality of reading what is currently happening in your own filesystem being described as a hypothetical workflow that someone might one day adopt. The pattern landed. The explicit rule got written down. Nothing about the actual practice changed.

The rule, in its final form, is one line. Agent reads it back, markdown. Human reads it once and links or shares, HTML. The memory wiki is markdown. The briefing is markdown. The planning files are markdown. The investigation reports, audit dumps, comparisons, dashboards, and anything-else-shareable: HTML.

// what the negative space is showing

Five of the six things above stopped happening in the last eight weeks because evidence arrived in some form. A paper. A token bill. A session that failed loudly. An audit. The retirement of an entire workflow. The sixth had been correct since the first weeks of this work last winter, and only got reasons attached to it recently, from people who arrived at the same answer from a different starting point.

Parallel arrivals like these are going to keep happening, and the gap is already short. Mission-not-role and the MIPT paper that backed it arrived in the same March, from completely different starting points. HTML-for-read-once took about six months to show up in a tweet. The next one will be faster still.

PublicThis stacklast winter//JanuaryMarchAprilMaythe paper's pattern, adaptedsame monthlocal-then-public, about six monthsHTML for read-once artefactsthe local defaultRecursive Language ModelsZhang, Kraska, Khattab (MITCSAIL), arXiv 2512.24601MIPT paperDochkina, arXiv 2603.28990Anthill retiredmission, not role/rlm-explore shipson the local stack"HTML is the new markdown"Thariq Shihipar; Karpathyquote-tweets it three dayslaterThis stackPubliclast winter//JanuaryMarchAprilMaythe paper'spattern, adaptedsame monthlocal-then-public,about six monthsHTML forread-onceartefactsthe localdefaultRecursiveLanguage ModelsZhang, Kraska,Khattab (MITCSAIL), arXiv2512.24601MIPT paperDochkina, arXiv2603.28990Anthill retiredmission, notrole/rlm-exploreshipson the localstack"HTML is thenew markdown"Thariq Shihipar;Karpathyquote-tweets itthree days later
Three orders of arrival. Paper first, same month, practice first: the three orders in which an answer reached this stack and the field. Three arrivals are too few to call a trend, so read the shrinking gap as a bet rather than a measurement.

Two-lane timeline comparing public evidence with practice on this stack. Public lane: in January, the Recursive Language Models paper by Zhang, Kraska and Khattab of MIT CSAIL, arXiv 2512.24601; in March, the MIPT paper by Dochkina, arXiv 2603.28990; on May 8th, Thariq Shihipar's tweet 'HTML is the new markdown', quote-tweeted by Karpathy three days later. This-stack lane: last winter, HTML for read-once artefacts as the local default; on March 25th, Anthill retired and mission, not role adopted; in April, the /rlm-explore skill shipped on the local stack. Connectors pair them: the January paper to the April skill, labelled 'the paper's pattern, adapted'; the March paper to the March 25th retirement, labelled 'same month'; and the last-winter HTML default to the May 8th tweet, labelled 'local-then-public, about six months', highlighted in amber as the longest gap.

This is the most reassuring property of the current moment in AI work. The space of working answers about how to use these tools well is small enough that good answers converge from many directions at once. When you reach one, expect company, and expect it within months. The thing that distinguishes the work is not which answer you found, because most working answers are going to be found by multiple people on different stacks. The thing that distinguishes the work is the lag between finding the answer and acting on it.

Eight weeks was enough time to retire five defaults and find that one more had been right since well before the field had a name for it. The next eight will produce another batch in some mixture, by which point the field will probably have published a couple of them and the rest will still be running quietly in someone else's planning directory.

That is the post. Don't re-derive what is already in memory. Don't grep flat through a codebase when you can decompose it. Don't fan out parallel agents when the work has dependencies. Don't tell subagents who they are. Don't claim done without running the verification command. Don't write markdown for things a human will read once. None of these are clever. All of them are, today, the right answer. By August at least two of them will be other people's posts.

Back to blog