Model Routing: Send Cheap Work to Cheap Models
Models

Model Routing: Send Cheap Work to Cheap Models

Most requests do not need your most capable model. Routing by task cuts cost sharply and adds resilience — here is how to build it without a mess.

Sending every request to your best model is the default because it is the simplest thing that works. It is also, for most applications, the single largest avoidable cost — and it makes you fragile to one provider having a bad day.

Two separate problems

Routing picks the right model for a request. Fallback picks a different one when the first fails. They are often built together and are worth thinking about separately, because the failure modes differ.

Routing strategies, in order of complexity

Static routing by task type. You know at the call site what kind of work it is. Commit message generation goes to the cheap model; architectural review goes to the strong one. Unglamorous, predictable, and captures most of the available savings.

Routing by input characteristics. Length, language, presence of code, estimated complexity. Cheap heuristics, no extra model call.

Classifier routing. A small fast model decides which model should handle the request. Adds latency and a failure mode of its own. Worth it only at volume where the savings dominate.

Escalation. Try the cheap model, evaluate the result, escalate on failure. Powerful when you have a cheap check — tests passing, schema validating — and wasteful when you do not, because you pay twice on every escalation.

Start static. Most teams that build a classifier discover the static version captured the majority of the benefit.

Fallbacks that do not lie to you

A fallback chain is easy to get subtly wrong. Rules worth following:

  • Fall back on infrastructure failures, not quality. Timeouts, 5xx, rate limits — retry elsewhere. A response you dislike is not a failure condition; retrying until you like it is how you burn money.
  • Do not retry non-retryable errors. A 400 for a malformed request will fail identically on every provider.
  • Cap total attempts. Two providers, not five. Beyond that you are adding latency to a request that is going to fail anyway.
  • Record which model answered. Without this, debugging quality regressions is guesswork.
  • Preserve semantics. If the fallback lacks tool calling or a large enough window, it is not a fallback for that request.

The compatibility that makes this cheap

Routing is straightforward when every candidate speaks the same API shape. With an OpenAI-compatible endpoint, switching model is changing a string. Without it, every provider needs its own client, request mapping and response normalisation — and that adapter layer is where routing projects usually die.

A gateway that exposes many models behind one endpoint collapses this to configuration, which is the difference between routing being a config change and a refactor.

Watch for silent quality drift

The real risk of routing is that it works — until it does not, quietly. A cheap model handles 90% of cases well, and the 10% degrade in ways nobody notices because no error is raised.

Guard against it: sample routed requests and score them periodically, alert on fallback rate rather than only on errors, and keep the option to pin a route to one model while investigating.

Do not route what you cannot evaluate

Routing assumes you know which tasks the cheap model handles adequately. If you have no way to check, you are not routing — you are gambling and calling it optimisation.

This is where a private eval set pays for itself twice. Run the same tasks through each candidate and you learn precisely where the cheaper option stops being sufficient. That boundary is your routing rule, derived from evidence rather than intuition.

Latency is part of the decision

Cost is the usual motivation, but responsiveness often matters more to users. A cheaper model that answers in half the time can be the better choice even at equal price.

The inverse also holds: a cheap model that needs four agent turns where the expensive one needed two may be slower and more expensive overall. Measure turns to completion, not just price per token — per-token comparisons routinely mislead in agentic workloads.

A reasonable starting point

  1. Route by task type at the call site — no classifier.
  2. One fallback per route, on infrastructure failures only.
  3. Log model, latency, tokens and outcome on every call.
  4. After a fortnight, look at where the cheap model was actually sufficient and adjust.

That gets most of the savings for very little machinery, and it leaves you with the data to justify anything more sophisticated.

Common questions

Should I use a classifier model to route requests?

Usually not at first. Static routing by task type captures most of the savings without adding latency or another failure mode. Consider a classifier only at high volume.

Should I fall back to another model if I dislike the answer?

No. Fall back on infrastructure failures such as timeouts and 5xx errors. Retrying until output looks good doubles cost and hides quality problems rather than fixing them.

How do I stop routing from quietly degrading quality?

Sample and score routed requests regularly, alert on fallback rate as well as errors, and keep the ability to pin a route to a single model while you investigate.

Similar articles

Fallback Model Selection: Choosing Your Second Model Well
Models
Models·9 min read

Fallback Model Selection: Choosing Your Second Model Well

A fallback that behaves nothing like your primary turns an outage into a quality incident. How to pick a second model and prove it works.

Read
A/B Testing Two Models Without Fooling Yourself
Models
Models·9 min read

A/B Testing Two Models Without Fooling Yourself

Comparing two models on live traffic sounds simple and usually is not. Sample sizes, paired designs, and the metrics that actually settle the question.

Read
Active vs Total Parameters: The Number Spec Sheets Hide
Models
Models·9 min read

Active vs Total Parameters: The Number Spec Sheets Hide

A 2.8T model and a 27B model can be two-to-one apart on the figure that governs thinking. How to read parameter counts across the 2026 field.

Read