What is model routing?
A model router sits between your application and a pool of models. For each request, it decides which model should handle it, instead of sending everything to the same one by default. The approach works because of a simple mismatch. Most requests you send are easy, things like a summary or a quick lookup. But frontier models are priced for the hard ones, the requests that need real reasoning. By 2026, the input-token price gap between the cheapest and most capable models has stretched to roughly two orders of magnitude, so the costliest model can run about a hundred times more per token than the cheapest.
The default behaviour in most organizations is to send everything to the strongest model available, because that is the safe choice and nobody gets criticised for it. It is also how you end up paying frontier prices for classification, extraction and formatting.
Published results are consistent enough to take seriously. Cascade approaches, beginning with Stanford's FrugalGPT work in 2023, reported cost reductions up to around 98% against always calling the strongest available model, with comparable output quality. Later router work trained on preference data reported substantial savings on standard benchmarks. Real-world cascade deployments more commonly land in the region of 37 to 46% reduction while routing roughly two thirds of traffic to cheaper models.
Read those figures as proof the technique works, not as the number you will hit. The headline results are tied to specific benchmarks and specific model pairings, and savings vary sharply by workload: general question answering saves a lot, mathematically heavy work saves considerably less. Run the evaluation on your own prompts before promising a percentage to anyone in finance. The economics of this sit at agentic AI ROI; this page is about the mechanics.
What are the routing strategies?
Four are in common use, and a 2026 survey sorts them by three things. These are when the decision happens, what information feeds the router, and how it gets computed. The practical split is simpler than that. You decide either before the request, during inference, or after seeing a first answer.
Ordered by how much they cost to build, which is roughly the inverse of how accurate they are.
| Strategy | How it decides | Where it fits |
|---|---|---|
| Static | Rules you wrote. This task type goes to that model | Start here. Cheapest to build, easiest to reason about, and captures most of the available saving |
| Classifier | A small trained model predicts difficulty before the request | High volume with an identifiable easy and hard split. Adds tens of milliseconds once trained |
| Semantic | Embedding similarity against known query classes | Where task types are distinguishable by meaning rather than by an obvious feature |
| Cascade | Try the cheap model, check the answer, escalate if it falls short | The most accurate approach and the most demanding, because it needs a reliable failure signal |
The latency objection is usually a misframing
Teams often reject routing because the router adds a step. The arithmetic does not support that. At a typical median inference time around 800 milliseconds, a 100 millisecond classifier is roughly an eighth of the call, and it can pay that back several times over by sending the request to a model that answers in 300 milliseconds rather than 1,500. The router is not the bottleneck; the model choice it makes is.
The genuine exception is a router that itself calls a language model to judge difficulty. That adds a full inference round trip and should be reserved for cases where the decision cannot be made any other way.
One implementation note worth acting on: routing logic is most durable enforced at a gateway rather than written into each application. Hardcoding the choice in six services means six places to change when a model is deprecated, and no single place to see what is being routed where.
Why does cascade routing need a failure signal?
Because a cascade only works if you can tell that the cheap answer was inadequate. Try the cheap model, assess the output, escalate when it falls short. That middle step is the whole design, and if you had a reliable general way to detect a wrong answer you would have solved a considerably harder problem.
Which means cascades work brilliantly in one class of task and poorly in another, and the dividing line is worth knowing before you build one.
- Cascades work where failure is mechanically detectable. Code that does not compile, JSON that does not parse against a schema, a tool call that returns an error, an extraction missing a required field, a computation you can check. The signal is deterministic and cheap.
- Cascades struggle where the wrong answer is plausible. A summary that misses the point, an analysis with a subtle error, a confidently incorrect factual claim. Nothing about the output declares itself inadequate, so the cascade accepts it. This is the same boundary drawn at AI hallucination: what can be checked by rule and what needs judgement.
- Confidence scores are a partial substitute and a fragile one. Research on calibrated cascade routing exists and works, and it requires deliberate calibration rather than taking raw model confidence at face value. Uncalibrated confidence is exactly the signal a plausible wrong answer scores highly on.
Why are the router's two errors not equivalent?
Routing down and routing up fail in totally different ways. If you send a hard request to a cheap model, you get a wrong answer, but at least it's cheap. If you send an easy request to an expensive model, you get the right answer, but you overpaid for it. Here's the thing though, almost every router gets tuned to minimize average cost, which basically treats those two mistakes as if they are the same. But they are not.
They are not. One costs money and the other costs correctness, and correctness usually costs more.
| The error | What it costs, and how you find out |
|---|---|
| Routed up: easy request, expensive model | Money. The answer is correct, the waste appears in the bill, and it is self-limiting because somebody eventually looks at the bill. Recoverable |
| Routed down: hard request, cheap model | Correctness. The answer is wrong and fluent, the cost looks excellent, and nothing surfaces it. In an agentic system it becomes an action. Often unrecoverable |
Two consequences follow, and both run against how routing is usually configured.
Bias the router toward over-provisioning. Where the decision is close, route up. You are trading a known, visible, bounded cost against an unknown one. That is the right trade almost everywhere, and it is not what a system optimised purely for average cost per request will do.
Measure the two directions separately. A single accuracy figure for your router hides which way it errs. Report the rate of requests routed down that should have gone up, and treat that as a quality metric owned alongside the cost metric rather than inside it. The asymmetry is sharper still in an agentic system, where a cheap wrong answer becomes a step in a trajectory rather than a sentence somebody reads.
What does routing break?
There are three things the rest of this glossary insists on, and they are reproducibility, evaluation, and the audit record. A router can improve performance while staying completely opaque about how it works. The same paper that documents this improvement also names what gets worse because of it, which is that similar prompts start behaving inconsistently and fallback behaviour stays hidden from view.
This is the part missing from most routing material, and it is the part that matters once a routed system is doing something consequential.
- Reproducibility goes first. A failure you cannot attribute to a model is a failure you cannot investigate. If three models could have served that step, debugging starts with a guess about which one did.
- Evaluation fragments. If a step can run on three models, your evaluation has to cover three paths or you have tested a configuration that only sometimes runs. Most teams evaluate the system and have in fact evaluated one route. See agent evaluation.
- The audit record loses a field it needs. Which model produced a decision is part of establishing how a decision was reached, and it is not recoverable after the fact. See AI audit trail.
- Behaviour becomes inconsistent in a way users notice and cannot explain. Two near-identical prompts routed differently produce differently shaped answers, and support has no vocabulary for why.
The fix has a name in the literature: a route receipt. Record, per request, which model and version served it, what the routing decision was and on what basis, whether a cascade escalated, and what it cost. That is four fields, it is cheap, and it converts routing from something opaque into something you can reason about. It also depends on per-step tracing existing at all, which is covered at AI agent observability.
How should you adopt it?
Start with static rules, instrument before you optimise, and be sceptical of sophistication. A 2026 benchmark found that many routing methods perform similarly under unified evaluation, and that several recent approaches including commercial routers did not reliably beat a simple baseline.
That finding should shape the sequence, because it means the cheap version of this captures most of the value.
-
Write static rules first
Classification, extraction, formatting and summarisation to a cheap model. Anything requiring multi-step reasoning to a strong one. This takes an afternoon, is completely explainable, and captures a large share of the available saving before you evaluate a single vendor.
-
Add the route receipt before anything clever
Model and version, decision and basis, escalation, cost, per request. Do this before you introduce a classifier, because otherwise you cannot tell whether the classifier helped.
-
Measure both error directions on your own traffic
Not just average cost. How often was something routed down that should have gone up, and what did that cost in rework or in a wrong action. That number decides whether to route more aggressively or less.
-
Enforce at a gateway, not in each application
One place to change when a model is deprecated, one place to see what is going where, and one place to apply a spend ceiling. Hardcoding the choice in each service guarantees drift.
-
Only then consider a learned router, and hold it to a baseline
Given the benchmark position, require any classifier or commercial router to demonstrate improvement over your static rules on your prompts. If it cannot, the static rules are the better system because they are explainable.
-
Watch the escalation rate as a standing metric
For any cascade. A rising escalation rate is either a mis-set threshold or a change in your traffic, and either way it is the thing that turns a saving into a loss quietly.
One point specific to agentic systems. Routing per step rather than per request is where the real gains sit, because a single agent run may include one genuinely hard reasoning step and a dozen mechanical ones. But it multiplies the observability requirement: a run touching five models needs five receipts, or the trace tells you less than it appears to.