Where LLM spend actually goes
Enterprise AI spend grows with three things: long contexts (contracts, statements, data rooms), thinking tokens, and agents that re-read the same files at every step. Prompt tweaks shave a few percent. Moving the reading to a cheaper model saves 90%, but only if you can prove the cheaper model got it right.
That proof is the whole product. Each job comes with checks: totals must reconcile, every quoted clause must exist in the document, two different models must agree. A reading that fails is read again by a stronger model. Nothing is accepted on trust.
How it works
- A top model plans the job once. Claude Opus looks at a few samples and writes the instructions, the checks and the arithmetic. You pay for that once per kind of job, and the plan is kept for next time.
- Cheap models do the reading. A short tryout picks the cheapest open models that pass the job's checks; they read every document. Sums, dates and comparisons are done in code, not by a model.
- Every answer is checked. Totals have to reconcile and two different models have to agree. When they don't, a stronger model reads it again; only real disagreements go back to the top model.
- You get a receipt. Before the job you see what Claude, ChatGPT and Gemini would each charge alone; after it, what each document cost and which checks it passed.
Measured against Claude Opus on its own
| Job | Accuracy | Opus alone | Yellowjacket | Cheaper by |
|---|---|---|---|---|
| 80 real contracts (CUAD), checked against lawyers' answers | 286 vs 284 of 320 | $4.41 | $0.50 | 9× |
| 164 coding problems (HumanEval), checked by their tests | 164 of 164 vs 55 of 55 | $0.347 (55) | $0.033 (164) | 25× |
| 95 real receipt photos (CORD) | 93 vs 92 of 95 | $0.84 | $0.28 | 3× |
| A year of supplier statements (165k tokens) vs the ledger | both exact | $1.00 | $0.15 | 7× |
| 30-invoice exception desk, each repeat batch | both exact | $0.79 | $0.006 | 130× |
| 95 receipts as text, a small one-off job | 91 vs 95 of 95 | $0.09 | $0.10 | Opus wins |
Billed costs from our own runs. Small one-off jobs are cheaper on a top model alone, so Yellowjacket sends those straight to one; the savings come from long documents, volume and repeat work.
Ways to cut LLM costs, compared
| Approach | Typical saving | Risk |
|---|---|---|
| Shorter prompts, prompt caching | 10-50% on repeated prefixes | Low |
| Lower reasoning effort / a smaller top model | 2-3× | Accuracy drops on hard steps |
| An LLM router (pick a model per request) | 2-5× | Nothing checks the cheap model's answer |
| Plan once, cheap models read, every answer checked (Yellowjacket) | 7-25× measured, 130× on repeat batches | Checks catch the cheap model's mistakes and escalate them |
We measured the second row too: on 30 exact business questions Claude Sonnet 5.5 matched Opus (30/30 vs 29/30) for about a third of the price, so small direct questions go to Sonnet.
Questions people ask
How much can I save on LLM costs?
On our measured runs: 9x on contract review, 7x on long financial documents, 25x on code with tests, 130x on repeat batches of agent work. Small one-off questions save nothing, so they go straight to a top model at cost plus a small fee.
Does a cheaper model mean worse answers?
Not when every answer is checked. On 80 real contracts Yellowjacket scored 286/320 against lawyers' labels; Claude Opus alone scored 284/320.
Is this an LLM router?
It goes further. A router picks one model per request and trusts its answer. Yellowjacket splits a job into reading, checking and arithmetic, has the reading cross-checked by two models, and escalates only the readings that fail.
Which models does it use?
Claude Opus plans and rules on disputes. Workers are open models (for example DeepSeek V4 Flash, Qwen3 and Gemma) chosen per job by a tryout; Gemini 3.1 Pro handles hard images.