Model routing

Most requests in a real product are easy. Sending all of them to your best model is the most expensive habit in the industry.

Part of the Cheap and fast track on lAItest.

Most requests in a real product are easy. Sending all of them to your best model is the most expensive habit in the industry.

Routing is a decision per request, not per project

Every vendor ships a ladder: a small model, a middle model, a flagship. Adjacent rungs differ by several times in price. Routing means classifying the request first — a rewrite, a lookup, a hard multi-step task — and sending it to the cheapest rung that reliably passes your evals. The simplest version is escalation: run the cheap model, check the result, re-run only the failures on the expensive one.

One vendor's ladder, input side

Read from a live model index at page-render time, each figure linked to the vendor page it came from.
WhatValueProvenance
Claude Haiku 4.5$1 per 1Msource, verified .
Claude Sonnet 5$2 per 1Msource, verified .
Claude Opus 5$5 per 1Msource, verified .
Claude Fable 5$10 per 1Msource, verified . Note which of these is the newest model. It is not the most expensive one.

A common misconception

Commonly believed: Newer is cheaper and better, so route everything to the latest version and move on.

Actually: Not reliably. Google's newer Flash model is priced lower on output than the one it appears to replace, while Google's own copy still calls the older one its most intelligent model for sustained agentic and coding work — so an automatic upgrade can quietly regress. The newest xAI flagship has half the context window of its predecessor. Newer is a different model, not a better one on your task.

Try it

Move a million requests one rung down the ladder and watch what happens to the monthly bill. This step is an interactive widget; open the lesson to use it.

In one sentence

Routing is the largest single lever on cost, and the one nobody can automate for you, because it needs your evals to say what good enough means.