The impossible revision
An invoice whose “revision” is dated a year after the original — and whose review stamp predates the document’s own authorship. Flagged as a template artifact, cited to both file versions.
We asked an AI agent — with Find and Seek as its memory — to audit six departments’ worth of company files: spend, risk, vendors, signatures, the lot. What follows is what it actually produced, with every claim cited to its source file. The files are synthetic by design, which is why we can publish all of it.
Buried in two different departments’ budget files sat a vendor called “Northstar meeldia Pty Ltd” — the company’s own name, misspelled the same way twice — billed as an external staffing vendor, at two different amounts: $110,000 in one file, $240,000 in another.
““Northstar meeldia Pty Ltd” — Data-integrity error — $110,000 ‘staffing’ line in budget_1300_924 is a garbled rendering of the company’s own name (Northstar Media), i.e., the company appears to be billing itself as an external vendor — not a legitimate third party”
— the agent’s memo, verbatim (vendor-integrity table, full report)
Why the difference? This risk only exists across team boundaries — one amount in one team’s files, another amount in another’s. That’s exactly the kind of thing team boundaries make invisible to people, and exactly what a memory that reconciles across them makes visible — the governance page explains why this changes the job itself. The ordinary agent wrote a competent memo too. It just never saw this — three times out of three.
The files contained six different “total cost” documents for the same program. The agent found them all, refused to average them, and reconciled them the way an auditor would — naming which source is authoritative and why:
| Candidate | What it claims | What the agent noted |
|---|---|---|
| Legacy stub | $54,200 | No author, no date, five bare lines — “not a genuine program-total claim” |
| Official master budget | $435,000 stated | Its own 13 line items sum to $465,000 — the agent caught the $30,000 arithmetic error in the official document, and identified the exact dropped line |
| Variant C | $744,666 summed | Entirely different vendor roster; carries the self-billing vendor |
| Variant D | $683,833 summed | Third vendor roster; $90,500 of it flagged “pending approval” |
| IT shared-drive rows | $225,000 | A law firm billed for ad buys; a packaging vendor billed for social media — flagged implausible |
Final answer, with the working shown: a reconciled program total of ≈ $547,020, and a stated variance between defensible candidates of $309,666 (41.6%) — plus the wider figure reported “for completeness but not recommended,” because the honest answer includes the answer it decided against.
An invoice whose “revision” is dated a year after the original — and whose review stamp predates the document’s own authorship. Flagged as a template artifact, cited to both file versions.
One signer’s email domain silently differs between her invoice and every other message in the record — the kind of spoofing tell that hides in plain sight. Flagged for manual verification, not asserted as fraud.
The same logistics vendor at $30,000 in the approved budget and $158,333 in a variant — and a $4,250 plumber nobody approved. Twenty-four vendors found outside the approved list, each with its source path.
Every row of the full memo carries its file path. When the record couldn’t support an answer, the agent said so instead of inventing one — “low-confidence signal, not a confirmed finding” is a phrase it uses unprompted.
Producing this memo took the agent roughly 44 tool calls and about $0.45 of inference (three-run average, live claude-sonnet-5). The same agent doing it the ordinary way read through more than a million tokens of files to write its version — and still missed Exhibit A. The point isn’t only that it’s cheaper. It’s that this class of work — cross-team, cited, repeatable — becomes something you can run routinely instead of annually.
The corpus is synthetic by design — real company document types, no customer data — which is why the complete output is publishable. There was no planted answer key: the catch shows cross-team reach, and we score it as exactly that. Both agents ran the same model over the same files; the only difference was the memory in between. Walls are per team today, per person with the enterprise build.