LLM cost optimization that keeps the top model's answers

Most of an AI bill is a frontier model reading long documents, one page at a time. Yellowjacket keeps the frontier model for planning and checking and moves the reading to models that cost a small fraction as much. Measured: the same accuracy as Claude Opus for 7 to 25 times less.

Where LLM spend actually goes

Enterprise AI spend grows with three things: long contexts (contracts, statements, data rooms), thinking tokens, and agents that re-read the same files at every step. Prompt tweaks shave a few percent. Moving the reading to a cheaper model saves 90%, but only if you can prove the cheaper model got it right.

That proof is the whole product. Each job comes with checks: totals must reconcile, every quoted clause must exist in the document, two different models must agree. A reading that fails is read again by a stronger model. Nothing is accepted on trust.

How it works

  1. A top model plans the job once. Claude Opus looks at a few samples and writes the instructions, the checks and the arithmetic. You pay for that once per kind of job, and the plan is kept for next time.
  2. Cheap models do the reading. A short tryout picks the cheapest open models that pass the job's checks; they read every document. Sums, dates and comparisons are done in code, not by a model.
  3. Every answer is checked. Totals have to reconcile and two different models have to agree. When they don't, a stronger model reads it again; only real disagreements go back to the top model.
  4. You get a receipt. Before the job you see what Claude, ChatGPT and Gemini would each charge alone; after it, what each document cost and which checks it passed.

Measured against Claude Opus on its own

JobAccuracyOpus aloneYellowjacketCheaper by
80 real contracts (CUAD), checked against lawyers' answers286 vs 284 of 320$4.41$0.509×
164 coding problems (HumanEval), checked by their tests164 of 164 vs 55 of 55$0.347 (55)$0.033 (164)25×
95 real receipt photos (CORD)93 vs 92 of 95$0.84$0.283×
A year of supplier statements (165k tokens) vs the ledgerboth exact$1.00$0.157×
30-invoice exception desk, each repeat batchboth exact$0.79$0.006130×
95 receipts as text, a small one-off job91 vs 95 of 95$0.09$0.10Opus wins

Billed costs from our own runs. Small one-off jobs are cheaper on a top model alone, so Yellowjacket sends those straight to one; the savings come from long documents, volume and repeat work.

Ways to cut LLM costs, compared

ApproachTypical savingRisk
Shorter prompts, prompt caching10-50% on repeated prefixesLow
Lower reasoning effort / a smaller top model2-3×Accuracy drops on hard steps
An LLM router (pick a model per request)2-5×Nothing checks the cheap model's answer
Plan once, cheap models read, every answer checked (Yellowjacket)7-25× measured, 130× on repeat batchesChecks catch the cheap model's mistakes and escalate them

We measured the second row too: on 30 exact business questions Claude Sonnet 5.5 matched Opus (30/30 vs 29/30) for about a third of the price, so small direct questions go to Sonnet.

Questions people ask

How much can I save on LLM costs?

On our measured runs: 9x on contract review, 7x on long financial documents, 25x on code with tests, 130x on repeat batches of agent work. Small one-off questions save nothing, so they go straight to a top model at cost plus a small fee.

Does a cheaper model mean worse answers?

Not when every answer is checked. On 80 real contracts Yellowjacket scored 286/320 against lawyers' labels; Claude Opus alone scored 284/320.

Is this an LLM router?

It goes further. A router picks one model per request and trusts its answer. Yellowjacket splits a job into reading, checking and arithmetic, has the reading cross-checked by two models, and escalates only the readings that fail.

Which models does it use?

Claude Opus plans and rules on disputes. Workers are open models (for example DeepSeek V4 Flash, Qwen3 and Gemma) chosen per job by a tryout; Gemini 3.1 Pro handles hard images.

Take the sting out of your AI bill

Run a job