Find and Seek.
Evidence report · 1 team · Measured 13 Jul 2026

One campaign. One page. Same model — half the bill.

We asked a normal AI assistant and Find and Seek the same question over the same company files. Both wrote a grounded executive summary. One burned through whole-file re-reads. The other asked for the passages that mattered.

The numbers

3-run average. Both sides finished.

Side Tool calls Input tokens Cost / run Wall-clock
a normal AI assistant 30 150,873 $0.368 81.6s
find & seek 11 −63% 39,504 −74% $0.114 −69% 44.5s −45%

Representative single run (inside the average): 31 calls · $0.33 · 74s vs 9 calls · $0.08 · 29s. Model: claude-sonnet-5 via the live Anthropic API.

The prompt

What we asked — verbatim.

Draft a one-page executive summary of the Spring Glow creative campaign. Include: the total spend across all Spring Glow creative work, the creative concepts that were approved, and which agency or team delivered them. Ground every figure and claim in the company’s documents and note the source.

Methods

How we kept it fair.

Same model

Live API

claude-sonnet-5. Exact billed tokens — not estimates.

Same files

Campaign only

Synthetic company corpus. One permission group for this task.

Only change

How it finds

Whole-file re-reads vs relevant passages with sources.

We only cite runs where both sides completed a real deliverable. We publish the bill and the answer — not a recipe for how Find and Seek ranks or stores passages.

What surprised us

Cheaper — and still careful.

The cheaper answer was still careful. Both sides hedged where the files disagreed on spend. Both refused to treat a forged-looking sign-off email as fact. The saving was not “cheap because it retrieved nothing.”

Conflicting totals showed up on both sides. The files themselves disagree. A useful assistant surfaces the conflict instead of picking a number and sounding confident. Both did.

The bill dropped hardest on input. Input tokens dominate agent costs because context is re-sent every step. Cutting what gets sent is why the dollar line moves.

One honest miss. On a spot-check, the normal assistant’s full-file read caught a placeholder sender address on that suspicious email that Find and Seek’s bounded passages did not surface. Compact by default is cheaper and less oversharing — and sometimes a detail only a full-file read would catch. We publish that, not hide it.

Sources

Documents the answers draw on.

campaign/briefs/brief.txt campaign/budgets/spring_glow_budget.csv campaign/emails/creative_approved.eml legal/contracts/northstar_retainer.pdf operations/invoices/doc_978_6347.md finance/briefs/doc_46_3591.eml
Outputs

Read the actual answers.

One representative run, word for word. Numbers on a single run sit inside — not exactly on — the 3-run average above. Synthetic corpus: same document types a real company holds, no customer data.

A normal AI assistant

31 tool calls · $0.33 · 74s

        

Find and Seek

9 tool calls · $0.08 · 29s

        

What this does not prove

Not a claim that every Claude or ChatGPT session gets cheaper — only work that uses Find and Seek against connected files. Not a penetration test. Not proof on every model forever — this report is live claude-sonnet-5. Synthetic corpus: same types of files, not a customer’s private drive.