When a Cheap Model Is Enough
Most production LLM traffic does not need a frontier model. A practical framework for deciding which tasks can drop a tier without anyone noticing.
The default in most codebases is one model for everything, chosen once, usually the best one available at the time. It is a reasonable starting point and an expensive steady state, because the majority of requests in a typical production system are not doing anything a smaller model would fail at.
The question is which ones. Answering it by intuition produces the two standard failure modes: paying frontier prices to reformat JSON, or downgrading the one task where reasoning quality was the whole product.
The property that decides it
Forget task categories for a moment. The variable that predicts whether a cheap model suffices is how much of the answer is already present in the input.
Where the answer is contained in the prompt and the model's job is to locate, transform or reformat it, capability differences between tiers largely vanish. Where the model must supply knowledge, hold a plan across many steps, or choose between non-obvious trade-offs, they dominate.
That single distinction sorts most workloads correctly:
- Answer is in the input. Extraction, classification, reformatting, translation between known schemas, summarising a supplied document, routing a support ticket, generating boilerplate from a spec, converting natural language to a constrained query. Cheap models handle these well.
- Answer must be constructed. Debugging unfamiliar failures, multi-file refactors, architectural trade-offs, long-horizon agent work, anything where a wrong answer is plausible-looking and expensive.
The grey middle is code generation, which spans both ends depending on how well specified the task is. A function with a clear signature and tests is a transformation. "Make the checkout flow work" is not.
Second: what does a failure cost
Capability is only half the decision. The other half is the blast radius of being wrong.
use_cheap_model if:
answer_is_in_input
AND (failure_is_detectable OR failure_is_cheap)
Detectability is doing a lot of work in that expression. A classification that returns an invalid label is caught by a schema check. A summary that quietly drops the one clause that mattered is not caught by anything, and will surface as a customer problem three weeks later.
This is why structured outputs and cheap models pair so well. If the response must conform to a schema, a large class of small-model failures becomes a validation error rather than a silent wrong answer, and you can retry or escalate deterministically.
Measure it on your own data, in an afternoon
Nobody can tell you whether a cheaper model works for your task, because published benchmarks measure someone else's task. The experiment is small enough that there is no excuse for guessing.
- Take 100 real production inputs. Sampled, not curated, and including the awkward ones.
- Record current outputs from your existing model. This is your reference, not ground truth.
- Run the same inputs through the cheaper model with the same prompt. Do not tune the prompt yet.
- Define pass or fail before you look. Schema valid, required fields present, no contradiction with the source, human spot check on a random 20.
- Compare rates, then compute the cost of the gap. If the cheap model fails 4% more often and a failure costs a retry, the retry cost is what you compare against the saving.
The result is frequently surprising in both directions. Tasks people assumed needed the frontier often do not. Tasks assumed trivial sometimes have a long tail of edge cases where the tier gap is stark.
Routing patterns that work
Once you know which tasks can drop a tier, three patterns cover almost every deployment.
Static routing. Different code paths use different models, decided at build time. Boring, transparent, and correct for the large majority of systems. Start here.
Escalation. Try the cheap model, validate the output, retry on the expensive one if validation fails. Works well when failures are detectable. The economics depend on the failure rate: at a 10% escalation rate you pay roughly the cheap price plus a tenth of the expensive one, which is usually a large win. At 60% you have added latency and complexity for nothing.
Difficulty triage. A small classifier picks the model per request. Attractive in principle, and it adds a component that can be wrong in a way that is hard to debug. Justify it with measurements before building it.
Whatever you choose, log which model served each request. Without that field you cannot attribute a quality regression or a cost change to anything.
Where cheap models genuinely fail
Be honest about this or you will relearn it in production:
- Long-horizon agent work. Small errors compound across turns. A model that is 95% reliable per step is under 60% reliable over ten steps, and recovery from a bad intermediate state is exactly the kind of reasoning cheap models are worst at.
- Instruction density. Prompts with many simultaneous constraints get partially followed. The tenth requirement is the one dropped.
- Long context. Nominal window size is not the same as effective use of it. Accuracy across a large context degrades faster on smaller models.
- Tool calling under ambiguity. Choosing the right tool from twenty when the request is vague is a reasoning task wearing a formatting costume.
- Knowing when to stop. Cheap models are more likely to produce a confident wrong answer than to say the input is insufficient.
A decision rule you can apply today
- List your model calls by volume. Optimise the top three; ignore the tail.
- For each, ask whether the answer is already in the input.
- If yes, check whether failure is detectable — usually meaning a schema you can validate.
- If both hold, run the 100-input comparison before switching anything.
- Keep the frontier model on the path where the reasoning is the product, and stop apologising for the cost there. That is the one place it is clearly worth it.
The saving from getting this right is usually larger than any pricing negotiation, and unlike a rate card it does not expire. Flat-rate access changes the calculus somewhat — when the marginal token is free to you, the reason to route down becomes latency rather than cost — but the capability boundaries above hold regardless of how you are billed.
Common questions
How do I know if a cheaper model is good enough for my task?
Run 100 real production inputs through both models with the same prompt, define pass or fail criteria before looking, and compare failure rates. Published benchmarks measure other workloads and rarely predict your outcome.
Which tasks are safe to move to a cheaper model?
Tasks where the answer is already present in the input and failure is detectable: extraction, classification, reformatting, schema translation and summarising supplied text. Structured outputs make failures catchable, which is what makes the downgrade safe.
Why do cheap models struggle with agents specifically?
Errors compound across turns. A model that is 95 percent reliable per step falls below 60 percent over ten steps, and recovering from a bad intermediate state requires exactly the reasoning that smaller models are weakest at.