Yellowjacket vs Claude Opus alone
| Job | Accuracy | Opus alone | Yellowjacket | Cheaper by |
|---|---|---|---|---|
| 80 real contracts (CUAD), checked against lawyers' answers | 286 vs 284 of 320 | $4.41 | $0.50 | 9× |
| 164 coding problems (HumanEval), checked by their tests | 164 of 164 vs 55 of 55 | $0.347 (55) | $0.033 (164) | 25× |
| 95 real receipt photos (CORD) | 93 vs 92 of 95 | $0.84 | $0.28 | 3× |
| A year of supplier statements (165k tokens) vs the ledger | both exact | $1.00 | $0.15 | 7× |
| 30-invoice exception desk, each repeat batch | both exact | $0.79 | $0.006 | 130× |
| 95 receipts as text, a small one-off job | 91 vs 95 of 95 | $0.09 | $0.10 | Opus wins |
Billed costs from our own runs. Small one-off jobs are cheaper on a top model alone, so Yellowjacket sends those straight to one; the savings come from long documents, volume and repeat work.
Which top model for which job
| Test | Result |
|---|---|
| 95 receipt photos, one call each | Gemini 3.1 Pro 95/95 $0.23 · Grok 4.7 94/95 $0.43 · Claude Sonnet 5.5 94/95 $0.42 · Claude Opus 5.5 92/95 $0.84 |
| 30 exact multi-step business questions | Claude Sonnet 5.5 30/30 $0.11 · Claude Opus 5.5 29/30 $0.31 |
| 20 contracts x 4 questions, cheap workers vs lawyers | DeepSeek V4 Flash 75/80 (best worker) |
That is why small text questions go to Sonnet, hard images to Gemini, and planning and rulings stay with Opus.
Prompt style vs thinking tokens
Asking Opus to think in short drafts cut output tokens 12% and got 30/30. Writing in ASD-STE100 Simplified Technical English kept accuracy but used 21% more tokens, so we don't use it for reasoning.
Data and method
Contracts: CUAD (lawyer-labelled commercial contracts). Receipts: CORD (real receipt photos). Code: HumanEval. Statements, invoices and purchase orders: generated documents with exact ground truth. Costs are what OpenRouter billed. Run records are kept with the code.
Questions people ask
Is Yellowjacket always cheaper than a top model?
No. On a small one-off job (95 short receipts as text) Opus alone was cheaper and more accurate, so such jobs go straight to a top model. Savings come from long documents, volume and repeat work.
Who ran these benchmarks?
We did, on public datasets and on generated data with exact answers, at billed prices.