nullhex

Verified Public Facts

7 Aug 2026·11 min read·ai

In July we produced a decision brief on a matter with real money attached, and about a third of the way down, under a heading that read Verified public facts, sat this sentence:

The site identifies [an operating brand] and, on its terms and billing pages, [a named limited company].

Names redacted, structure untouched. It is a profoundly boring sentence. It is the sort of thing that appears in every brief anyone has ever paid for, written in the flat confident register of somebody who has done the reading and is now telling you what they found.

That sentence went on to direct a solicitor conflicts check, a set of monitoring instructions, and a do-not-contact list. Three separate actions, aimed at a company we could not show existed anywhere in our own evidence.

At this point somebody at the back puts their hand up to explain that language models make things up, that this is the oldest news in the industry, and that we have essentially written a blog post about water being wet. Which would be fair, if that were the finding. It is not. The finding is that this brief went through a review round with multiple models on it, all of them asked to challenge it, and the sentence came out the other side with its confidence entirely undisturbed. Everybody looked. Nobody checked.

There is a second thing worth saying immediately, because it is the part that makes this useful rather than merely embarrassing. The company may well be real. The brief was written against the live web a day before we froze the evidence, and the pages it read that off were never admitted. It may be sitting on that page right now, exactly as described. That does not rescue the sentence. The defect was never whether the claim was true; it was that the brief could not tell you which of its sentences it could prove, and had filed this one under a heading asserting that it could.

// the sentence

The rest of the brief was not the work of an idiot. It contained, three bullets further down, an unusually good piece of reasoning about what the earliest records did and did not establish, carefully separating chronology from recognition, standing, and anything the other party could be shown to have known. That paragraph would survive a hostile reading. It is genuinely careful work.

Then, in the same list, under the same heading: six separate fee figures that appear nowhere in the evidence, quietly driving a spending recommendation. A recommended outreach email opening with a first name that appears in no admitted file. And a small chronological miracle, in that the brief is dated the day before every piece of evidence it claims to have verified.

None of these is a dramatic failure. That is exactly the problem. A brief that announces that the counterparty is a cabal of lizards gets caught in the first thirty seconds by anyone. A brief that says a company is identified on a billing page gets a slow nod, because it is the most ordinary sentence in the world, and the reader's attention has already moved on to the recommendations, which is where the reader always thinks the risk lives.

The review round did look at it. Our two scorers, working blind, described that review afterwards as a vote tally with one plan change and no per-point disposition, none of it catching the unsupported identification. Which is the sentence I would tattoo on the inside of my eyelids if I were the tattooing sort. A room full of reviewers who vote, but who are never required to resolve a specific claim against specific evidence, will ratify a confident sentence. They are not being lazy. They have simply been given a job that does not include looking anything up.

Here is the claim. Type it yourself, before I tell you the rule.

Type the claim

One sentence from the brief. The evidence available was a frozen packet: site snapshots, third-party records, internal history, one unsigned agreement template. Names are redacted here; the structure is exactly as written.

The site identifies [an operating brand] and, on its terms and billing pages, [a named limited company].

// what a claim has to declare

The mechanism that prevents all of this is unglamorous and takes about a page to describe.

We run it as a piece of in-house software called Octaval. It is not a product and it is not for sale; it is the thing we built for ourselves after noticing that we knew perfectly well how to do careful research and did not reliably do it. What Octaval holds is a case: the admitted evidence with its provenance, every material claim with its status, the contradictions and how each one was resolved, the review exchange in full, and a decision at the end with a named human owner. It decides nothing itself. It makes skipping a step expensive and advancing with an untyped claim impossible, and that turns out to be the whole trick.

The stages are Access, Provenance, Retrieval, Contradiction, Evaluation, Synthesis, Review, and Decision. The one doing the heavy lifting is Retrieval, because that is where every material claim has to be typed as exactly one of four things. SUPPORTED, which requires the file and the exact location inside it: a line number, a section heading, a JSON path. PREMISE, meaning the operator asserted it and the evidence does not contain it. Premises are never evidence. ASSUMPTION, meaning it is yours, and you are saying so. UNRESOLVED, meaning nobody can settle it yet.

And then the part everyone forgets: material absences. What the evidence verifiably does not contain, written down as a finding in its own right, because the shape of the hole is frequently the most decision-relevant thing on the table.

Under those four types, the sentence that started this post cannot happen. Not "is less likely to happen", not "would probably get caught by a sufficiently alert reviewer on a good day". Cannot. The claim has no locator, so it cannot be typed SUPPORTED; it types as PREMISE; and the synthesis stage forbids any recommendation resting on a specific absent from the evidence. The conflicts check never gets written. The failure is not caught downstream by vigilance, it is unavailable upstream by construction, which is a much better place to solve a problem than in the reviewer's attention span.

The Octaval brief worked the same case from the same evidence, and the company does not appear in it. Not once. What appears instead is a row in the claim ledger:

The other party's legal identity (company, jurisdiction, address, principals) unknown. SUPPORTED absence. No identity anywhere in the packet.

Read that second line again. The absence is not a gap in the brief; it is a finding, carrying a status, checked against the whole of the admitted evidence and recorded as carefully as any positive claim. The brief knows that it does not know who the other party is, it can tell you exactly what it searched to establish that, and it says so in the same voice it uses for the things it can prove.

And then, instead of a conflicts check aimed at a name it could not stand behind, it recommends pursuing the formal identity-disclosure channel that the evidence itself points to. Same question, same evidence. One brief invents certainty and acts on it; the other writes down the shape of its ignorance and tells you how to fix it.

The obvious objection is that this turns every piece of research into a compliance exercise, and it would, if you did it to everything. The ledger is deliberately lean: material claims only, the ones the decision actually rests on. Immaterial detail is not logged. If a claim could be wrong without changing what you do next, it does not need a locator, it needs deleting.

// the packet

None of this works if the evidence can move while you are looking at it.

So the evidence gets fixed first. An admitted file list, hashed, agreed before any work starts. Anything not on the list is not evidence, and anything on the list you have decided not to rely on gets said out loud, with the reason. Then one line per file recording what it is, where it came from, its date, and its custody weakness: self-generated, fetched snapshot, third-party record, unsigned template. A screenshot you took of your own website and a record from a neutral third party are not the same species of fact, and a method that treats them alike will eventually let you cite yourself as corroboration.

This is the stage everybody skips, and skipping it is why "the website says" is not a locator. A live website is not a fixed object. It is a performance that happens to be running when you look, and a claim sourced to one cannot be rechecked tomorrow, by you or by anybody who has to rely on your work. Our failing brief is dated the day before its own evidence: it verified its facts against a world, and then the world was replaced with a different one, and nothing in the process noticed.

// the contradiction pass runs somewhere else

One stage in the method exists purely to argue with the rest of it, and it is run by a separate context that has seen the evidence and the claim ledger but none of your reasoning. It returns a list of contradictions. Every item has to be dispositioned in writing: accepted, rebutted with a locator, or typed unresolved. No item may be quietly dropped for being inconvenient.

The rationale is that a reviewer who has read your argument reviews your argument. They follow your path, and they check whether the steps connect, which is a genuinely useful thing to do and is not at all the same as checking whether the path goes anywhere near the evidence. Hiding the reasoning was supposed to fix that.

We preregistered it as a measurable bet, ran it as a controlled sub-experiment, and it came back unsupported.

The deliberately contaminated pass, the one that got to see the synthesis first and was therefore supposed to be hopelessly anchored to it, scored nine. The clean isolated pass scored eight. Three high-materiality discoveries were unique to the contaminated pass, and one of them was a mischaracterisation of the evidence that the isolated pass missed, that the independent reviewer also missed, and that survived into the winning brief and cost it points with both scorers. Our best method shipped an error that our own contaminated control caught.

In fairness to the bet, the isolated pass surfaced five medium-materiality items the other one missed, and the conditions were not symmetric: the contaminated pass ran later, with the isolated pass's dispositions already visible, so it was building on ground already cleared. The honest reading is complementarity rather than superiority. Two passes at different points in the pipeline beat one pass run twice, and the thing we were confident about turned out to be the thing we should have measured earlier. We are keeping the bet in the record because a method that only publishes its wins is not a method, it is advertising.

// against a deep research run

It is worth being precise about what the losing brief actually was, because it was not a lazy prompt typed into a chat box at four in the afternoon. It was a deep research run: the well-known workflow where an agent plans, searches, reads across many sources, and synthesises a long report with citations. On top of that it got a review round with several models challenging it. Two real stages, both of them work, both of them the thing most people mean when they say they used AI to research something properly.

Here is what the two approaches do differently, and none of the differences are about model quality.

The evidence is fixed before the work starts. A deep research run gathers as it goes, from a live web, and the report is a snapshot of whatever was there during the search. Octaval takes an admitted file list, hashes it, and works only from that. Every claim is checkable tomorrow, by someone else, because the thing it was checked against still exists in the same form.

Every material claim carries a type and a locator. Deep research produces prose with citations, which sounds like the same thing and is not. A citation says roughly where an idea came from. A locator says exactly where, precisely enough that a reader finds the passage in under a minute, and the type says what kind of claim it is before you are allowed to build on it.

Absences are findings. This is the one with the biggest gap between the two. A deep research report tells you what it found. It has no mechanism at all for telling you what is not there, so the hole where the counterparty's identity should be does not appear anywhere in the output, and the reader has no way to notice it is missing. Octaval writes the hole down, types it, and cites it.

The challenge sees the evidence, not the argument. Both workflows have a review step. In the ordinary one, reviewers read the report and push back on it, which means they are checking whether the argument hangs together. Octaval's contradiction stage runs in a separate context that gets the evidence and the claim ledger and none of the reasoning, and every item it raises has to be dispositioned in writing.

Numbers get recomputed, not quoted. Every date interval, price, and quantity is recalculated and the arithmetic shown where it matters. The losing brief carried six fee figures it had picked up somewhere and could not resolve against anything.

The end of the document is a decision, not a recommendation. A named human owner, confidence per item, the blockers that must clear first, the alternatives that stay open and why, and the specific future facts that would reverse each call. A recommendation tells you what to do. A decision record tells you who owns it and what would change it.

None of this makes the deep research run useless. It is very good at the thing it is for, which is covering ground fast and telling you what is out there. It is simply not an instrument for deciding anything, and the failure mode when you use it as one is not that it looks unfinished. It is that it looks finished.

// what it scored

Sealed blind labels, two fresh scorers who never saw which brief was which, or the workspace, or each other's marking. The protocol was frozen before any brief existed, and the ceiling amendment, the clause governing what to conclude if the scores bunched at the top, was recorded before any arm produced a word. That last detail is the difference between a trial and an anecdote with numbers on it.

Blind scores for both briefs, primary and sensitivity scorers, with admissibility gate outcomes.
BriefPrimarySensitivityAdmissibility gates
Octaval, eight stages9997Pass, both scorers
Deep research plus multi-model review4646Fail all three, both scorers

The scorers verified eighteen to twenty locators per brief directly against the evidence and recomputed every date before scoring, which is why this took a day rather than an hour.

Fifty-three points, on both scorers' cards, in the same direction. And it has now been produced twice, on two cases, in two fields, with fresh scorers each time, which matters more than the size of it: a fifty-point gap on one case is a story, and a fifty-point gap on two is a property of the method.

Octaval spent about 400,000 tokens across five sequential stages, and about twenty-one minutes of agent time, which is roughly what a mid-sized pull request review costs. The failing brief was cheaper. Bad research always is, right up until somebody acts on it.

// what this is not

One case, one field, one operator per arm. Two scorers.

The operators, the reviewers and the scorers are all fresh contexts of the same model family in the same harness, isolated by prompt-level rules and not by anything as robust as an operating system. Correlated failure is a live problem and we have direct evidence of it: the same mischaracterisation of the evidence arose independently in three different agents' first drafts, and was caught only because something later in the pipeline was obliged to go and look.

The rubric is deliberately aligned with evidence-discipline mechanisms, which means it saturates near the top and cannot price anything a careful operator would have done anyway. Our own preregistered isolation bet came back unsupported. And the failing brief, in fairness, was written against a live web that our frozen evidence could not represent, so some of what it could not support may be perfectly true.

Its actual failure is unchanged by all of that. It stated unverifiable specifics as verified fact, and then let them drive actions with a solicitor at the end of them.

The gap between forty-six and ninety-nine was not intelligence. The same class of model wrote both briefs, on the same day, from the same evidence, and one of them was excellent. The difference was whether a claim had to declare what kind of claim it was before anything was allowed to use it.

A brief that cannot tell you which of its sentences it can prove is not a brief. It is a draft with confidence applied evenly, and confidence is the one thing in the whole process that costs nothing to produce.