Benchmarks: cost and accuracy, measured

Every figure on this site comes from a run we paid for, graded against exact answers or lawyers' labels. Including the one Opus wins.

Yellowjacket vs Claude Opus alone

JobAccuracyOpus aloneYellowjacketCheaper by
80 real contracts (CUAD), checked against lawyers' answers286 vs 284 of 320$4.41$0.509×
164 coding problems (HumanEval), checked by their tests164 of 164 vs 55 of 55$0.347 (55)$0.033 (164)25×
95 real receipt photos (CORD)93 vs 92 of 95$0.84$0.283×
A year of supplier statements (165k tokens) vs the ledgerboth exact$1.00$0.157×
30-invoice exception desk, each repeat batchboth exact$0.79$0.006130×
95 receipts as text, a small one-off job91 vs 95 of 95$0.09$0.10Opus wins

Billed costs from our own runs. Small one-off jobs are cheaper on a top model alone, so Yellowjacket sends those straight to one; the savings come from long documents, volume and repeat work.

Which top model for which job

TestResult
95 receipt photos, one call eachGemini 3.1 Pro 95/95 $0.23 · Grok 4.7 94/95 $0.43 · Claude Sonnet 5.5 94/95 $0.42 · Claude Opus 5.5 92/95 $0.84
30 exact multi-step business questionsClaude Sonnet 5.5 30/30 $0.11 · Claude Opus 5.5 29/30 $0.31
20 contracts x 4 questions, cheap workers vs lawyersDeepSeek V4 Flash 75/80 (best worker)

That is why small text questions go to Sonnet, hard images to Gemini, and planning and rulings stay with Opus.

Prompt style vs thinking tokens

Asking Opus to think in short drafts cut output tokens 12% and got 30/30. Writing in ASD-STE100 Simplified Technical English kept accuracy but used 21% more tokens, so we don't use it for reasoning.

Data and method

Contracts: CUAD (lawyer-labelled commercial contracts). Receipts: CORD (real receipt photos). Code: HumanEval. Statements, invoices and purchase orders: generated documents with exact ground truth. Costs are what OpenRouter billed. Run records are kept with the code.

Questions people ask

Is Yellowjacket always cheaper than a top model?

No. On a small one-off job (95 short receipts as text) Opus alone was cheaper and more accurate, so such jobs go straight to a top model. Savings come from long documents, volume and repeat work.

Who ran these benchmarks?

We did, on public datasets and on generated data with exact answers, at billed prices.

Take the sting out of your AI bill

Run a job