Live API
claude-sonnet-5. Exact billed tokens — not estimates.
The hardest generation task we measured: reconcile Spring Glow across Campaign, Finance, Legal, HR, Operations, and IT — conflicting totals, a risk register, vendor mismatches, and identity flags. Both sides finished. Find and Seek finished with far less of the company’s text in the prompt.
| Side | Tool calls | Input tokens | Cost / run | Wall-clock |
|---|---|---|---|---|
| a normal AI assistant | 93 | 1,169,902 | $0.758 | 394.3s |
| find & seek | 44 −53% | 284,855 −76% | $0.445 −41% | 287.0s −27% |
Representative single run: 99 calls · $0.84 · 442s vs 47 calls · $0.49 · 305s. Cost and speed advantages narrow as tasks span more teams — token reduction stayed strong (~76%).
Produce a comprehensive spend-and-risk audit memo for the Spring Glow (SGX) program, covering every department that touched it: Campaign, Finance, Legal, HR, Operations, and IT. (1) State total program cost across ALL departments — not just creative spend — including staffing/contractor time, IT infrastructure, logistics, and legal/compliance costs. Reconcile every conflicting figure you find and state which source is authoritative and why. (2) Produce a risk register: every invoice, contract, or commitment that is unsigned, unpaid, pending approval, or flagged non-compliant, with department and dollar amount. (3) Cross-reference vendors across departments — flag any vendor whose billed amount differs between departments’ records, or who appears in one department’s records but not in Campaign’s own approved vendor list. (4) Produce a single reconciled program-total variance: the dollar gap between the highest and lowest total figures found across all six departments, and that gap as a percentage of the highest figure. Show your work. (5) Cross-reference every person named as a sender, signer, or approver across all six departments’ documents. Flag any name that appears with inconsistent details as a possible data-integrity or fraud flag. (6) Cite every figure and claim to its source document.
claude-sonnet-5. Exact billed tokens — not estimates.
Campaign, Finance, Legal, HR, Operations, IT — like separate folders on a shared drive.
Empty or truncated attempts are excluded from the headline.
Without Find and Seek: list folders, search filenames, read whole files.
With Find and Seek: the most relevant passages with their sources — not whole-file dumps.
Token reduction held when the task got ugly. Six teams, long memos, heavy citations — Find and Seek still sent ~76% fewer input tokens. Less of the company’s material re-entering the model every turn.
Both sides independently caught planted mess. Misspelled or self-referencing vendor names. Document revisions with chronologically impossible dates. Conflicting “totals” that no honest auditor should collapse into one quiet number.
Cost advantage narrowed more than token advantage. Billed cost and raw token volume do not always move in lockstep once prompt caching is in play. We report both — and prefer the range across workloads over a single hero percentage.
Identity flags were the jaw-dropper. Same names under different titles across departments; contact domains that don’t match; templated signature blocks. That showed up because the prompt forced a cross-department people check.
One representative run, word for word. Both sides had to confront conflicting totals, open risk, vendor mismatches, and people / integrity flags. Scroll each panel — these are long on purpose.
Not “41% cheaper on every enterprise audit forever.” Not that Find and Seek replaces finance or legal judgment — it feeds the model less noise so that judgment costs less to run. Same synthetic corpus and model caveats as our other reports. Prefer the range across workloads (≈41–69% cheaper) over any one headline.