First, find where the tokens go
Output costs far more than input: Claude Opus 5.5 lists $4 per million input tokens and $20 per million output tokens, and thinking is billed as output. Agents multiply both, because every step re-sends the conversation and thinks again. On a 30-invoice accounts payable task an Opus agent cost about $0.026 per invoice; one Opus call over the whole batch cost $0.063 for all 30, roughly 12× less per invoice. Before changing anything, log input, output and thinking tokens per kind of request.
1. Send the reading to cheaper models, and check their work
Most tokens are spent reading documents, and reading is what small models can do, if something checks them. Give each answer a check software can run: totals that must reconcile, quotes that must exist in the document, two different models that must agree. Escalate only what fails. On 80 real contracts this scored 286 of 320 against lawyers' labels, against 284 for Claude Opus alone, for $0.50 of AI cost instead of $4.41. A router that picks a cheap model and trusts it skips the part that makes this safe.
2. Test models on your own work, with your checks
Public leaderboards don't tell you which cheap model reads your documents. Run candidates on a dozen real items and score them with the checks from step 1; no answer key is needed. On supplier statements one tryout scored Mistral Small 12/12, Qwen3-30B 10/12, DeepSeek V4 Flash 5/12 (invalid JSON) and Llama 3.1 8B 1/12. The same goes for top models: on 30 exact business questions Claude Sonnet 5.5 scored 30/30 for $0.11 against Claude Opus at 29/30 for $0.31, and DeepSeek V4 Flash with thinking off 30/30 for $0.0077.
3. Keep the arithmetic in code
Asking a model to add up, compare dates or match two lists costs output tokens and invites mistakes. Have the model read the numbers out and let code do the sums. Reconciling a year of supplier statements, Claude Opus needed 17,227 output tokens to work it through itself; with the matching in code, cheap models only read one statement at a time.
4. Plan once, reuse the plan
Work out the instructions, checks and code for a kind of job once with a strong model, store them, and reuse them for every later batch. Planning cost $0.103 for the statement job and $0.26 for the contract job; on the 30-invoice task the first batch cost $0.099 (mostly planning) and each later batch about $0.006.
5. Cache verified results
Cache each checked reading, keyed on the input and a version of the instructions, so a changed prompt never serves an old answer, and never cache a reading that failed its checks. Asking the statement question a second time cost $0.00 and took 0.1 seconds.
6. Cap output, but not below what the answer needs
Output limits stop runaway spend, but a limit that is too low wastes the whole call. Claude Opus with a 16,000-token limit on the statement job spent it all thinking, returned nothing and still cost $0.98; at 64,000 it answered for $1.00. Our own first cap of 600 tokens on list answers cut every reading short and sent the work up to Opus: that run cost $3.14 before we fixed it. Size the cap to the answer, report truncation as an error, and put a budget on escalations too.
7. Ask for structured output
A JSON schema makes answers short and checkable by code, which is what steps 1, 3 and 5 depend on. Prefer models that support strict structured output; fall back to JSON mode and validate when they don't. Prompt wording matters less than people hope: asking Claude Opus to think in short drafts cut output tokens by 12% on our questions, while a simplified-English style used 21% more.
8. Keep data private when you move to cheaper providers
Cheaper models often run on hosts you don't know. Route only to providers that keep no copy of the request and don't train on it (OpenRouter's zero data retention setting does this), and swap names and account numbers for placeholders before sending. On 20 contracts the same job scored 70/80 with placeholders and 71/80 without.
How Yellowjacket does this for you
- A top model plans the job once. Claude Opus looks at a few samples and writes the instructions, the checks and the arithmetic. You pay for that once per kind of job, and the plan is kept for next time.
- Cheap models do the reading. A short tryout picks the cheapest open models that pass the job's checks; they read every document. Sums, dates and comparisons are done in code, not by a model.
- Every answer is checked. Totals have to reconcile and two different models have to agree. When they don't, a stronger model reads it again; only real disagreements go back to the top model.
- You get a receipt. Before the job you see what Claude, ChatGPT and Gemini would each charge alone; after it, what each document cost and which checks it passed.
You pay 35% of what the cheapest top AI model would charge for the same job, or our AI cost plus 25% when that is higher. Small one-off jobs are cheaper on a top model alone, so Yellowjacket sends them straight to one; the savings come from long documents, volume and repeat work.
Questions people ask
What is the fastest way to reduce LLM costs?
Find which requests use the most output and thinking tokens, then try a cheaper model on exactly those, with a check that software can run on every answer. In our tests the model choice moved cost far more than prompt wording.
Does prompt caching reduce LLM costs?
It reduces the price of repeated input prefixes on providers that offer it. It doesn't touch output and thinking tokens, which are usually the larger part of a document job's bill. We haven't measured it here.
Are cheaper models accurate enough?
On their own they make more mistakes. With checks and escalation, our cheap workers matched or beat Claude Opus on real contracts (286 vs 284 of 320) and receipt photos (93 vs 92 of 95).
When is a single top-model call cheapest?
Small one-off jobs. On 95 short receipts as text, Claude Opus alone cost $0.09 and was more accurate than our engine at $0.10, so Yellowjacket sends jobs like that straight to a top model.