All glossary terms
M Architecture Economics

Model routing

Sending every request to your best model is like calling a neurosurgeon to take a temperature. Routing fixes that, and introduces a quieter problem: similar prompts start behaving differently, and nothing in your trace records which model actually ran.

Definition

Model routing means sending each request to the model that fits it, instead of running everything through one model. It exists because what a task needs and what a model can do rarely match. Most requests are simple, but frontier models are priced for the hard ones, and the price gap between the cheapest and most capable models is now wide enough to change what you pay.

What is model routing?

A model router sits between your application and a pool of models. For each request, it decides which model should handle it, instead of sending everything to the same one by default. The approach works because of a simple mismatch. Most requests you send are easy, things like a summary or a quick lookup. But frontier models are priced for the hard ones, the requests that need real reasoning. By 2026, the input-token price gap between the cheapest and most capable models has stretched to roughly two orders of magnitude, so the costliest model can run about a hundred times more per token than the cheapest.

The default behaviour in most organizations is to send everything to the strongest model available, because that is the safe choice and nobody gets criticised for it. It is also how you end up paying frontier prices for classification, extraction and formatting.

Published results are consistent enough to take seriously. Cascade approaches, beginning with Stanford's FrugalGPT work in 2023, reported cost reductions up to around 98% against always calling the strongest available model, with comparable output quality. Later router work trained on preference data reported substantial savings on standard benchmarks. Real-world cascade deployments more commonly land in the region of 37 to 46% reduction while routing roughly two thirds of traffic to cheaper models.

Read those figures as proof the technique works, not as the number you will hit. The headline results are tied to specific benchmarks and specific model pairings, and savings vary sharply by workload: general question answering saves a lot, mathematically heavy work saves considerably less. Run the evaluation on your own prompts before promising a percentage to anyone in finance. The economics of this sit at agentic AI ROI; this page is about the mechanics.

What are the routing strategies?

Four are in common use, and a 2026 survey sorts them by three things. These are when the decision happens, what information feeds the router, and how it gets computed. The practical split is simpler than that. You decide either before the request, during inference, or after seeing a first answer.

Ordered by how much they cost to build, which is roughly the inverse of how accurate they are.

Four model routing strategies, how each decides, and what each suits
Strategy How it decides Where it fits
Static Rules you wrote. This task type goes to that model Start here. Cheapest to build, easiest to reason about, and captures most of the available saving
Classifier A small trained model predicts difficulty before the request High volume with an identifiable easy and hard split. Adds tens of milliseconds once trained
Semantic Embedding similarity against known query classes Where task types are distinguishable by meaning rather than by an obvious feature
Cascade Try the cheap model, check the answer, escalate if it falls short The most accurate approach and the most demanding, because it needs a reliable failure signal

The latency objection is usually a misframing

Teams often reject routing because the router adds a step. The arithmetic does not support that. At a typical median inference time around 800 milliseconds, a 100 millisecond classifier is roughly an eighth of the call, and it can pay that back several times over by sending the request to a model that answers in 300 milliseconds rather than 1,500. The router is not the bottleneck; the model choice it makes is.

The genuine exception is a router that itself calls a language model to judge difficulty. That adds a full inference round trip and should be reserved for cases where the decision cannot be made any other way.

One implementation note worth acting on: routing logic is most durable enforced at a gateway rather than written into each application. Hardcoding the choice in six services means six places to change when a model is deprecated, and no single place to see what is being routed where.

Why does cascade routing need a failure signal?

Because a cascade only works if you can tell that the cheap answer was inadequate. Try the cheap model, assess the output, escalate when it falls short. That middle step is the whole design, and if you had a reliable general way to detect a wrong answer you would have solved a considerably harder problem.

Which means cascades work brilliantly in one class of task and poorly in another, and the dividing line is worth knowing before you build one.

  • Cascades work where failure is mechanically detectable. Code that does not compile, JSON that does not parse against a schema, a tool call that returns an error, an extraction missing a required field, a computation you can check. The signal is deterministic and cheap.
  • Cascades struggle where the wrong answer is plausible. A summary that misses the point, an analysis with a subtle error, a confidently incorrect factual claim. Nothing about the output declares itself inadequate, so the cascade accepts it. This is the same boundary drawn at AI hallucination: what can be checked by rule and what needs judgement.
  • Confidence scores are a partial substitute and a fragile one. Research on calibrated cascade routing exists and works, and it requires deliberate calibration rather than taking raw model confidence at face value. Uncalibrated confidence is exactly the signal a plausible wrong answer scores highly on.

Why are the router's two errors not equivalent?

Routing down and routing up fail in totally different ways. If you send a hard request to a cheap model, you get a wrong answer, but at least it's cheap. If you send an easy request to an expensive model, you get the right answer, but you overpaid for it. Here's the thing though, almost every router gets tuned to minimize average cost, which basically treats those two mistakes as if they are the same. But they are not.

They are not. One costs money and the other costs correctness, and correctness usually costs more.

Two error directions compared. Routing up, meaning an easy request sent to an expensive model, produces a correct answer at unnecessary cost, described as a waste that is visible in the bill and self-limiting. Routing down, meaning a hard request sent to a cheap model, produces a confident wrong answer at low cost, described as invisible in the bill and reaching the user or an action. A caption notes that tuning a router on average cost treats these as the same mistake.
The two directions a router can be wrong, and how the consequences differ
The error What it costs, and how you find out
Routed up: easy request, expensive model Money. The answer is correct, the waste appears in the bill, and it is self-limiting because somebody eventually looks at the bill. Recoverable
Routed down: hard request, cheap model Correctness. The answer is wrong and fluent, the cost looks excellent, and nothing surfaces it. In an agentic system it becomes an action. Often unrecoverable

Two consequences follow, and both run against how routing is usually configured.

Bias the router toward over-provisioning. Where the decision is close, route up. You are trading a known, visible, bounded cost against an unknown one. That is the right trade almost everywhere, and it is not what a system optimised purely for average cost per request will do.

Measure the two directions separately. A single accuracy figure for your router hides which way it errs. Report the rate of requests routed down that should have gone up, and treat that as a quality metric owned alongside the cost metric rather than inside it. The asymmetry is sharper still in an agentic system, where a cheap wrong answer becomes a step in a trajectory rather than a sentence somebody reads.

What does routing break?

There are three things the rest of this glossary insists on, and they are reproducibility, evaluation, and the audit record. A router can improve performance while staying completely opaque about how it works. The same paper that documents this improvement also names what gets worse because of it, which is that similar prompts start behaving inconsistently and fallback behaviour stays hidden from view.

This is the part missing from most routing material, and it is the part that matters once a routed system is doing something consequential.

Two traces of the same request. The first records only the request and the answer, and three questions below it are marked unanswerable: which model ran, was there a fallback, and can this be reproduced. The second adds a route receipt recording the model and version, the routing decision and why, whether a cascade escalated, and the cost, and the same three questions are marked answerable. A caption states that without the receipt a routed system cannot be reproduced, evaluated per route, or audited.
  • Reproducibility goes first. A failure you cannot attribute to a model is a failure you cannot investigate. If three models could have served that step, debugging starts with a guess about which one did.
  • Evaluation fragments. If a step can run on three models, your evaluation has to cover three paths or you have tested a configuration that only sometimes runs. Most teams evaluate the system and have in fact evaluated one route. See agent evaluation.
  • The audit record loses a field it needs. Which model produced a decision is part of establishing how a decision was reached, and it is not recoverable after the fact. See AI audit trail.
  • Behaviour becomes inconsistent in a way users notice and cannot explain. Two near-identical prompts routed differently produce differently shaped answers, and support has no vocabulary for why.

The fix has a name in the literature: a route receipt. Record, per request, which model and version served it, what the routing decision was and on what basis, whether a cascade escalated, and what it cost. That is four fields, it is cheap, and it converts routing from something opaque into something you can reason about. It also depends on per-step tracing existing at all, which is covered at AI agent observability.

How should you adopt it?

Start with static rules, instrument before you optimise, and be sceptical of sophistication. A 2026 benchmark found that many routing methods perform similarly under unified evaluation, and that several recent approaches including commercial routers did not reliably beat a simple baseline.

That finding should shape the sequence, because it means the cheap version of this captures most of the value.

  1. Write static rules first

    Classification, extraction, formatting and summarisation to a cheap model. Anything requiring multi-step reasoning to a strong one. This takes an afternoon, is completely explainable, and captures a large share of the available saving before you evaluate a single vendor.

  2. Add the route receipt before anything clever

    Model and version, decision and basis, escalation, cost, per request. Do this before you introduce a classifier, because otherwise you cannot tell whether the classifier helped.

  3. Measure both error directions on your own traffic

    Not just average cost. How often was something routed down that should have gone up, and what did that cost in rework or in a wrong action. That number decides whether to route more aggressively or less.

  4. Enforce at a gateway, not in each application

    One place to change when a model is deprecated, one place to see what is going where, and one place to apply a spend ceiling. Hardcoding the choice in each service guarantees drift.

  5. Only then consider a learned router, and hold it to a baseline

    Given the benchmark position, require any classifier or commercial router to demonstrate improvement over your static rules on your prompts. If it cannot, the static rules are the better system because they are explainable.

  6. Watch the escalation rate as a standing metric

    For any cascade. A rising escalation rate is either a mis-set threshold or a change in your traffic, and either way it is the thing that turns a saving into a loss quietly.

One point specific to agentic systems. Routing per step rather than per request is where the real gains sit, because a single agent run may include one genuinely hard reasoning step and a dozen mechanical ones. But it multiplies the observability requirement: a run touching five models needs five receipts, or the trace tells you less than it appears to.

Frequently asked questions about model routing

If your question is not here, our team will answer it directly.


Talk to a Specialist →
What is model routing?

A layer that intercepts each call and picks which model handles it. The economics behind it are simple: the bulk of what an application asks is not difficult, pricing on the strongest models reflects the difficult cases, and by 2026 the spread between budget and frontier input pricing had widened to something like a hundredfold. Sending everything to the best available model is the choice nobody gets blamed for, which is why it is the default.

How much does model routing actually save?

Cascade research has reported reductions as high as roughly 98% measured against sending everything to the strongest option, while deployments in the field more often report somewhere between 37 and 46%, with about two thirds of requests handled by cheaper models. Treat the headline figures as proof the technique works rather than as your number: they are tied to specific benchmarks and model pairings, and savings vary by workload, with general question answering saving considerably more than mathematically heavy work.

What are the four routing strategies?

Static rules you write, a trained classifier predicting difficulty before the request, semantic routing using embedding similarity against known query classes, and cascade routing which tries the cheap model first and escalates. A 2026 survey organises the design space along three axes: when the decision is made, what information feeds it, and how it is computed. Cost to build runs roughly inverse to accuracy, which is why static is the right starting point.

Does a router add too much latency?

Usually not, and the objection is generally a misframing. At a typical median inference time near 800 milliseconds, a 100 millisecond classifier is about an eighth of the call, and it can repay that several times over by choosing a model that answers in 300 milliseconds rather than 1,500. The router is not the bottleneck; the model choice it makes is. The real exception is a router that itself calls a language model to judge difficulty, which adds a full round trip.

What is cascade routing?

Sending every request to a cheap model first, checking whether the answer meets a threshold, and escalating to a stronger model only when it does not. It is the most accurate strategy and the most demanding, because the middle step is the whole design. Stanford's FrugalGPT work in 2023 established the pattern and reported cost reductions up to around 98% with comparable quality, using a small trained scoring function rather than raw model confidence.

When does cascade routing not work?

Where the wrong answer is plausible. Cascades depend on detecting that the cheap answer was inadequate, so they work well where failure is mechanical: code that will not compile, JSON that fails a schema, a tool call returning an error, an extraction missing a required field. They struggle on a summary that misses the point or a confidently incorrect claim, because nothing about the output declares itself wrong and the cascade simply accepts it.

Why does the escalation rate matter?

Because a cascade escalating most of its traffic costs more than no routing at all: you pay for the cheap attempt and the expensive one. That failure is silent, since costs rise while quality still looks fine and nobody is watching the ratio. A climbing escalation rate means either the threshold is wrong or your traffic mix has changed, and both are findings worth acting on. Treat it as the standing health metric for any cascade.

Are the two ways a router can be wrong equally bad?

No, and most routers are tuned as though they were. Sending an easy request to an expensive model produces a correct answer at unnecessary cost, which appears in the bill and is self-limiting because somebody eventually reads the bill. Sending a hard request to a cheap model produces a confident wrong answer while the cost looks excellent, and nothing surfaces it. One costs money, the other costs correctness, and correctness usually costs more.

Should a router be biased in one direction?

Yes, toward over-provisioning. Where the decision is close, route up, because you are trading a known bounded visible cost against an unknown one. That is the right trade in most settings and it is not what a system optimised purely for average cost per request will do. It also means reporting the two error directions separately, since a single router accuracy figure hides which way it errs.

What is a route receipt?

A per-request record of which model and version served it, what the routing decision was and on what basis, whether a cascade escalated, and what it cost. Four fields, cheap to add, and they convert routing from something opaque into something you can reason about. The term comes from research framing routing as a trust problem, on the argument that operators need records of what routers actually did rather than assurances about average behaviour.

What does routing break if you do not instrument it?

Reproducibility, evaluation and the audit record. A failure you cannot attribute to a model is one you cannot investigate, since debugging starts with a guess about which of three models ran. Evaluation fragments, because a step that can run on three models needs three evaluated paths or you have tested a configuration that only sometimes runs. And which model produced a decision is part of establishing how it was reached, which is not recoverable afterwards.

Should you buy a commercial router?

Make it prove itself against something trivial first. When a 2026 benchmark put routing systems on a common footing, the differences between them largely washed out, and a number of recent entrants, vendor products among them, failed to beat a naive comparator consistently. So write static rules, instrument them, and require any classifier or vendor router to demonstrate improvement on your own prompts. If it cannot beat your rules, the rules are the better system because they are explainable.

Which model actually ran?
A routed system without receipts is a system you cannot debug

SERAA Cortex records every model call per agent with tokens, cost and actor attribution, inside your own perimeter, so which model served a step is a field in the trace rather than a guess, alongside the spend ceilings that make a routing decision enforceable.