1. the council audit
  2. the audit that lied, gently
  3. the two ghosts
  4. the instruction files
  5. what changed about the machine
  6. the thing one is left with
engineering8 min

Checks That Run By Themselves

The question I started with was a small one. I'd accumulated a few rules over the previous weeks - a multi-model review practice for the harder decisions, a memory wiki with confidence levels on each entry, a discipline of running the verification command before claiming the work was done - and I wanted to know what else was worth doing. The kind of question one asks on a Wednesday morning expecting a five-minute conversation and a list of three things.1

The first item on the list that came back was: audit the review practice and see if it's actually working. The other two read as downstream of the first. The afternoon I'd planned turned into two days, and the audit turned up rather more than the audit had been asked to.

// the council audit

Council is multi-model blind review on decisions that are hard to take back. Architectural commitments. Security boundaries. Money. Pay a small amount of compute to surface what a single mind would otherwise miss. The audit pulled the twenty-nine most recent sessions, each carrying its own ledger of where I'd stood going in and where I'd stood coming out.

I read every changed-decision session individually rather than sampling. Twenty-three sessions of prose about my own opinions: tedious, but the sort of work that needs doing honestly or not at all.

The headline finding was kinder than I had any right to expect. Of the twenty-two sessions in which my decision had moved, seven were genuine verdict-flips - I'd come in pointing one way and walked out pointing somewhere else entirely. Six were scope changes, where the direction held but the boundary moved. Five were what the audit ended up calling test-upgrades, where the course remained but the council had inserted a falsifying probe I hadn't designed in advance. Zero were cosmetic. No rubber-stamping. Council was paying for itself.

sessions
Verdict-flips7
came in pointing one way, walked out pointing another
Scope changes6
direction held, boundary moved
Test-upgrades5
course held, a falsifying probe inserted
Cosmetic0
no rubber-stamping
How decisions moved under council review. The bar worth reading is the empty one: not one session came out cosmetic. Each was read individually rather than sampled, by the person whose opinions they record.

Bar chart of how the author's decision moved in council sessions, from an audit of the twenty-nine most recent sessions. Of the twenty-two sessions in which the decision moved, seven were verdict-flips, where the author came in pointing one way and walked out pointing another; six were scope changes, where the direction held but the boundary moved; five were test-upgrades, where the course held but the council inserted a falsifying probe; and zero were cosmetic, shown as an empty bar labelled in amber.

// the audit that lied, gently

One of the earlier council sessions had been an operational review of the workshop itself - nineteen findings, several Critical, run five weeks ago. I went back to it as part of the audit pass to walk each finding against current state.

It turned out the review had been lying to me. Politely, but lying. Of the nineteen findings, ten had been addressed - just not in the form the review had asked for. The disaster recovery playbook wasn't at the literal filename the review had asked about; it had been written under a different one, the day after the review had requested it. The dream-consolidation soft-delete wasn't a directory; it was a deleted_at column with a seventy-two-hour recovery window. There was no local CI runner; there were per-project GitHub Actions across every active repository, which is the same thing only better.

The fixes had landed. The review had simply not known to look for them, because some later check had searched for the literal artefact names the original review had recommended, found none, and reported six unaddressed Critical findings with the slightly hurt tone that automated checks tend to adopt when they cannot find what they are looking for.2

Looked for by nameWhat had landed
Disaster recovery playbooknot at the literal filenameunder a different name, the day after it was requested
Dream-consolidation soft-deletenot a directorya deleted_at column, seventy-two-hour recovery window
CIno local CI runnerper-project GitHub Actions across every active repository
Check by function, not by artefact name. Ten of the nineteen findings had been addressed in forms the review had not asked for, so a later check that searched for the literal names reported six Critical findings as open.

Comparison table of three findings from an earlier operational review, checked two ways. Disaster recovery playbook: looked for by name, it was not at the literal filename; by function, it had been written under a different name the day after it was requested. Dream-consolidation soft-delete: not a directory, but a deleted_at column with a seventy-two-hour recovery window. CI: no local CI runner, but per-project GitHub Actions across every active repository. Every left-hand cell is marked missing and every right-hand cell present.

This is a finding worth nailing to a wall. Check by function, not by artefact name. The correct audit question is "does a disaster recovery procedure exist" rather than "does a file named DR.md exist." This applies, on reflection, to a great deal of life.

The corrected triage left a smaller pile of genuinely pending items. Most were closed before nightfall.

// the two ghosts

One of the smaller items on the remediation list was command-execve logging. snoopy preloaded via /etc/ld.so.preload, every execve on the box to /var/log/auth.log with user, ssh source, cwd, and full command line. The kind of thing one fits and forgets about, until the day one needs it.

Within a minute of it being live, the process listing turned up something I'd not been looking for.

matt  1418455  Mar 31 14:19:47  /bin/bash -c sudo apt-get install -y socat 2>/dev/null && socat TCP-LISTEN:3459,fork,reuseaddr,bind=0.0.0.0 TCP:127.0.0.1:3458 &
matt  1418655  Mar 31 14:19:48      socat TCP-LISTEN:3459,fork,reuseaddr,bind=0.0.0.0 TCP:127.0.0.1:3458

A socat listener on 0.0.0.0:3459, forwarding to 127.0.0.1:3458. Fifty-one days old. The exact shape any attentive listener-audit ought to catch on day seven, given one had been running. The kind of artefact one is, on first encounter, supposed to investigate before killing.3

Investigation took fifteen minutes. Nothing was listening on the target port. There were no active connections. No persistence in cron, systemd, or any rc file. UFW would have dropped any inbound connection from outside the box regardless of the bind address; default-deny incoming on this host, with only SSH and the anthill dashboard on the allow list. The shape was a development experiment from some earlier session - expose this local thing on the LAN so I can hit it from the other side of the room - that had simply never been switched off when whoever set it up was done with it.

Killing it took one command. Writing the cron that would have caught it in week one took half an hour. The cron lists every non-loopback TCP listener weekly, filters against an allowlist (sshd, tailscaled, docker-proxy, anthill), and drops a marker the daily-health-check surfaces until acknowledged.

The first run found another one: a Python http.server on 0.0.0.0:8765 from two days earlier, exposing a draft directory under _scratch/, launched via nohup for some testing purpose and never closed. Same shape, same lifecycle. Killed.

Forgotten listenersbind a dev server, detach, walk awaysocat on 0.0.0.0:3459forwarding to 127.0.0.1:3458;fifty-one days oldPython http.server on0.0.0.0:8765from two days earlier; a draftdirectory under _scratch/weekly cronevery Mondayevery non-loopback TCPlistenerone ss callon the allowlist?sshd, tailscaled, docker-proxy,anthillmarkerthe daily-health-check surfacesit until acknowledgedwould have beencaught in week onefound on the firstrunanything outsideit
A check that runs by itself. The socat forwarder went unnoticed for fifty-one days; with this running, the blind window is seven days at worst, for a few lines of bash and one ss call a week.

Two forgotten 0.0.0.0 listeners on the same machine in the same audit pass is the kind of finding that justifies the audit's existence. The pattern is well-attested - bind a dev server, detach, walk away - and the cost of catching it weekly rather than at fifty-one days is a few lines of bash and one ss call per week.

// the instruction files

Two other pieces of housekeeping landed in the same window, both about the documents that tell the workshop what to do.

The CLAUDE.md files across the personal workspace got an audit pass - per-project, workspace-level, global. These files are the long-term instruction surface: rules that constrain how work gets done, conventions for where things live, standing preferences. They drift the way any living document does; rules survive their original use case by months. The pass was scope-discipline: each file pruned against what its layer was actually for. The workspace-level CLAUDE.md went from 328 lines to 143 - a 56% cut, roughly 3,700 fewer tokens loaded into every context. An eight-probe adherence check before and after held its ground: one rule improved, one softened, net unchanged. Four reference memories were extracted to standalone files in the process. Three stale plugin entries in settings.json were quarantined.

In the same window, the gbrain repository's CLAUDE.md got the same treatment as a shipped PR (closing #884 and #429): 1,863 lines down to 464, roughly 85k tokens of mixed content cut by 93%. Reference material - a key-files inventory, a testing handbook, thin-client routing notes, bulk-action progress - moved out into docs/, where it can be consulted on demand rather than loaded into every context. The behavioural rules stayed inline in CLAUDE.md, where Claude actually reads them.

lines
Workspace CLAUDE.md143 - 328
328 to 143: a 56% cut, roughly 3,700 fewer tokens
gbrain CLAUDE.md464 - 1,863
1,863 to 464: roughly 85k tokens of mixed content cut by 93%
Two instruction files, before and after. The workspace file lost more than half its lines and an eight-probe adherence check came out net unchanged: one rule improved, one softened. The gbrain file's reference material moved to docs/, to be read on demand rather than every time.

Range chart of line counts for two CLAUDE.md files before and after an audit pass, on an axis from 0 to 2,000 lines. The workspace-level CLAUDE.md went from 328 lines to 143, a 56% cut and roughly 3,700 fewer tokens. The gbrain repository's CLAUDE.md went from 1,863 lines to 464, with roughly 85k tokens of mixed content cut by 93%. Each span runs from the after value on the left to the before value on the right.

Neither was dramatic work. Both produced the same quiet relief - the relief of having said clearly what had previously been implied across three places, and of having stopped saying the things that no longer applied.

// what changed about the machine

Stated plainly:

Command logging is live. snoopy preloaded via /etc/ld.so.preload; every execve on the box now hits /var/log/auth.log with full context.

A weekly listener audit runs every Monday. Non-loopback TCP listeners checked against an allowlist; anything outside it surfaces via the daily-health-check marker. The two-month-blind window that let the socat artefact survive is now seven days at worst.

The disaster recovery runbook reflects the current shape of the workshop. Future-me reading it during an actual disaster will be reading something accurate.

The council practice has a backlog. A small ordered list of tuning candidates - dashboard, schema discipline, criterion tagging - the kind of polish one queues for the moments between heavier work.

The instruction surface is sharper. Workspace CLAUDE.md trimmed; gbrain CLAUDE.md restructured per AGENTS.md scope. Less to read, more of it load-bearing.

// the thing one is left with

The thing I'm left with isn't any individual fix. It is the observation that very nearly every problem we found had been hiding behind something else.

The earlier review's "unaddressed" findings were not unaddressed; they had been addressed in places the next check hadn't known to look. The port forwarder was not malicious; it was a development experiment that had simply failed to switch itself off. The things that needed attention were not the ones one had been preparing to find.

The audits were the things that let us see the gap. Not because the audits were clever - the scripts are, frankly, rather thin - but because they ran. Once something is looking, the things hiding behind other things show up. They had been hiding the whole time. They were patient about it.

The previous weeks had been about adding ever-more-elaborate rules - the council criteria, the memory wiki ingest ripple, the verification discipline. The rules are fine; the audit said they were working. What I hadn't been adding, until this week, is the slightly different category of checks that run by themselves. The weekly cron that looks for forgotten listeners. The execve logger that puts every command into the auth log. The instruction-file pass that strips what isn't load-bearing.

Rules need someone to apply them. Checks just run.

That's the take.


  1. This is, generally, what most questions on Wednesday mornings are supposed to produce. It almost never works out that way. The asking of small questions appears to be something humans have evolved to do anyway, perhaps so that the larger questions feel less alarming when they arrive. ↩

  2. A surprising number of audits, on examination, turn out to be looking for the wrong thing in much the same spirit as the man in the well-worn joke who is searching for his keys under the streetlight because the light is better. ↩

  3. There is a particular discipline to finding an unfamiliar listener on one's own machine: the discipline of investigating before reacting. The first guess is rarely the right one. The second guess, depending on temperament, is also rarely the right one. Fifteen minutes of forensics costs less than the false-positive killing of something one's own past self had needed for a reason one had merely forgotten. ↩

Back to blog