Trust is earned, not given

A different perspective

2026-01-08 · Projects

AI Frontiers, part 59: Model routing β€” choosing a model per request, and proving it works

Part 59from the AI Frontiers series · 65 parts in all

The moment a system has more than one model available, every request becomes a policy decision, even if nobody wrote it down. Most teams make it implicitly: one model in the config, changed by a pull request when someone feels like it. The teams that got good at this over the last two years made the decision explicit, measured it, and discovered that the largest available efficiency gain in an applied AI system is usually not a better prompt or a smaller model β€” it is sending each request to the cheapest resource that can handle it.

The idea is old. FrugalGPT framed it as a cascade: try a cheap model, score the answer, and escalate only when the cheap answer is insufficient, choosing the sequence that minimizes cost subject to a quality constraint (Chen et al.). What changed by 2026 is that routing matured from a trick into an infrastructure component with a literature, a set of published evaluations, and a recognizable set of failure modes. This entry is about building one and, more importantly, proving it works β€” a harder problem than it looks, because a router is evaluated on counterfactual questions it can never directly observe.

The decision, properly framed

A router is a function from request to model, and it should optimize a stated objective under constraints. Four dimensions compete, and being explicit about their relative weights is the first useful step.

Quality, which is task-specific and should be measured with your own evaluation set rather than a leaderboard. Cost, dominated by output tokens and retries, per part 50. Latency, where the router itself adds a step and where interactive and batch traffic deserve completely different policies. And jurisdiction, which part 58 argued is a design variable: a request may not be legally routable to the best model for it.

Two structural consequences follow. First, routing is a constraint-satisfaction problem before it is a machine learning problem, and the biggest wins usually come from hard rules β€” "anything containing a patient identifier goes to the in-region deployment" β€” rather than from a classifier. Second, the objective should be cost per successful task rather than cost per call, because the cheapest model that fails and escalates costs more than the expensive model that succeeds.

Five routing strategies, from trivial to fragile

Static rules. Route by request type, user tier, tenant, or feature. No model, no training, fully auditable, and it captures most of the available win whenever the workload has obvious strata β€” which it usually does. If your product has a free tier and a paid tier, you already have a router; deciding it deliberately is the entire improvement.

Cascades with a confidence signal. Run the small model, measure whether the answer is acceptable β€” with a self-consistency check, a verifier model grading it against a rubric, or a cheap heuristic β€” and escalate on failure. The cost is the verifier, which must be cheaper than the difference between the models or the cascade loses money. This is the structure FrugalGPT evaluated, and it remains the best available option when the acceptability of an answer can be assessed without ground truth.

Learned routers. Train a classifier to predict which model will succeed. RouteLLM trained routers on preference data from model comparisons and reported large cost reductions at comparable quality, with the important detail that the router was trained on a preference dataset and evaluated on its ability to preserve responses' quality while cutting spend (Ong et al.). Hybrid LLM set the problem up explicitly as a quality-aware binary decision (Ding et al.), and earlier work routed among experts using reward-guided signals (Lu et al., "Routing to the Expert").

Bandit and contextual-bandit policies. Treat model selection as a sequential decision problem with an exploration cost. Classical results give you regret bounds and a principled way to balance exploration against exploitation (Auer et al.), which is exactly the right frame when you are learning a policy in production and cannot run a clean A/B test for every request.

Ensembles and judges. Query several models and pick the best response with a judge. Highest quality, highest cost, and useful mainly as a reference point to measure how much the cheaper policies give up.

Why published router numbers overstate reality

The published evaluations are honest and still systematically optimistic about deployment, for four reasons worth internalizing before you promise anyone a percentage.

Benchmark prompts are not your prompts. Routers are trained and evaluated on general instruction-following or chat data whose difficulty distribution has little to do with a specific product's traffic. Where the distribution differs, the router's decision boundary is in the wrong place, and the error is not random β€” it is concentrated in whatever makes your traffic unusual.

Feedback loops make the training distribution endogenous. Once the router is live, the data you collect only contains outcomes for the arm the router chose. High-quality answers from the cheap model are visible; answers the cheap model would have gotten wrong are not, because they never happened. This is the bandit problem's central difficulty, and it means naive retraining on logged data progressively overestimates the cheap model.

Quality is measured with a proxy. LLM-as-judge scores correlate usefully with human preference (Zheng et al., "Judging LLM-as-a-Judge") and are still proxies, with known biases toward length and self-similarity. A router optimized against a biased judge learns the bias.

The router has costs the benchmark ignores. It adds latency, it may require its own model call, and β€” a subtle one β€” splitting traffic across models destroys prompt-cache locality, which can consume a meaningful share of the savings on systems that rely on caching.

Proving a router works

Because you can only observe the outcome of the arm you chose, evaluating a routing policy requires deliberate counterfactual machinery. Three approaches, in increasing cost.

Uniform exploration on a slice. Send a small random fraction of traffic to each candidate model regardless of the router's decision, and use that slice as an unbiased measurement of every arm. Simple, statistically clean, and the only approach that gives you ground truth for all arms on the same distribution. The cost is the quality and latency sacrifice on that slice, which is why you keep it small and exclude anything regulated.

Off-policy evaluation. Use logged data with the router's propensities to estimate what a different policy would have done, using inverse-propensity or doubly robust estimators (DudΓ­k et al.). These techniques are standard in recommendation and advertising, they are implementable in an afternoon with the right logs, and they let you screen candidate policies before risking traffic on them. They require that you log the probability of each action, which is the piece teams forget to add.

Online experiments with a proper design. The gold standard, and the one that requires the discipline of sequential testing, pre-registered metrics, and enough traffic to detect the effect you care about β€” the material of any serious continuous evaluation practice (Kohavi et al.).

Every one of these produces a different number, and the honest way to report a router's benefit is as a range with the measurement method attached: "reduced cost per resolved task by thirty to forty percent at equal judged quality, measured on a ten-percent uniform-exploration slice over four weeks." A single unqualified percentage is a claim nobody should believe.

Failure modes that show up after launch

Rare-class regression. A router optimizes the average and is worst on the uncommon request types, which are frequently the most valuable ones. Slice the evaluation by request type and set a floor for each; do not accept an aggregate improvement that hides a category collapse.

Verifier gaming. If the cascade's verifier can be satisfied superficially β€” a fluent answer, a well-formed citation, a plausible number β€” the cheap model learns to produce exactly that. This is the same dynamic as synthetic preference data in part 44, and the mitigation is the same: keep an independent, human-validated measurement that the pipeline cannot see.

Router drift. Traffic distributions shift, models are updated underneath you, and a policy tuned in January is miscalibrated by June. Recalibrate on a schedule and monitor the route distribution itself β€” a sudden change in the fraction of requests sent to the expensive model is a signal, not a curiosity.

Two sources of failure. Every routed system has the model's failure modes plus the router's. When a quality complaint arrives, the first question is which model answered, and without routing information in the trace that question is unanswerable. Log the decision, the reason for it, and the policy version alongside everything else.

Where this leaves a production team

My summary recommendation is deliberately unambitious: start with static rules by request type, add a cascade with an executable verifier where one exists, and invest in the measurement infrastructure β€” propensities in the logs, a small exploration slice, per-slice metrics β€” before investing in a learned router. The learned router is a genuine improvement, and it is also the component most likely to be tuned against a proxy that does not match your business. The teams with the best routing outcomes in 2026 were not the ones with the cleverest classifiers; they were the ones that knew, to within a few percentage points, what their accuracy and unit cost actually were, and could tell when either moved.

Routing is also, quietly, the maturity marker of an applied AI organization. When a team can name the model on every path, say why, and show the measurement that justified it, they have stopped treating model choice as taste and started treating it as engineering. That is the same transition every other part of this stack has gone through, and it is where the interesting work begins.

What the logs must contain

Every measurement problem in this entry reduces to a logging problem, so it is worth being specific about the fields, because they are cheap to add at build time and impossible to reconstruct later.

The decision and the reason. Which model handled the request, which policy version made the choice, and why β€” the rule that matched, the classifier's score, the cascade's verification outcome. When a quality complaint arrives, this is the first question, and without it the investigation starts by re-running the request and hoping.

Propensities. For a learned or exploratory router, the probability with which each action was selected. This one field is what makes off-policy evaluation possible; without it you cannot estimate any policy other than the one you ran, and you are limited to online experiments forever.

Outcome labels, as they become available. The retry, the edit, the reopen, the thumbs-down, the human override. Delayed outcomes need to be joined back to the routing decision, which means the routing record has to carry a stable identifier the rest of the system can reference.

Cost and latency by arm. Tokens in, tokens out, dollars, and end-to-end latency including the router's own overhead and any verifier call. A router's economics are decided by the ratio of verifier cost to model-cost savings, and that ratio should be visible per configuration rather than estimated.

Cache and locality effects. Whether the request hit a prompt cache, and whether the routing decision broke cache locality. This is the effect most likely to be discovered late and the hardest to reconstruct after the fact, because it requires knowing the prefix hash and which model served the prefix previously.

None of this is exotic, and all of it is easier to add before the router exists. The teams that built the fields in from the start could answer questions about their routing policy in an afternoon. The teams that did not spent their time reconstructing what happened from partial logs, which is the same lesson part 49 made about traces: instrument the decision, not just the outcome.

Works Cited

Auer, Peter, NicolΓ² Cesa-Bianchi, and Paul Fischer. "Finite-Time Analysis of the Multiarmed Bandit Problem." Machine Learning, vol. 47, 2002, pp. 235–256. Accessed 8 Jan. 2026.

Chen, Lingjiao, Matei Zaharia, and James Zou. "FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance." arXiv, 2023, arxiv.org/abs/2305.05176. Accessed 8 Jan. 2026.

Ding, Dujian, et al. "Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing." arXiv, 2024, arxiv.org/abs/2404.14618. Accessed 8 Jan. 2026.

DudΓ­k, Miroslav, John Langford, and Lihong Li. "Doubly Robust Policy Evaluation and Learning." arXiv, 2011, arxiv.org/abs/1103.4601. Accessed 8 Jan. 2026.

Kohavi, Ron, Diane Tang, and Ya Xu. Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing. Cambridge University Press, 2020. Accessed 8 Jan. 2026.

Lu, Keming, et al. "Routing to the Expert: Efficient Reward-Guided Ensemble of Large Language Models." arXiv, 2023, arxiv.org/abs/2311.08692. Accessed 8 Jan. 2026.

Ong, Isaac, et al. "RouteLLM: Learning to Route LLMs with Preference Data." arXiv, 2024, arxiv.org/abs/2406.18665. Accessed 8 Jan. 2026.

Zheng, Lianmin, et al. "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena." arXiv, 2023, arxiv.org/abs/2306.05685. Accessed 8 Jan. 2026.