Tiers and surcharges

A quoted rate assumes a short request on the standard tier. Tier and length are separate multipliers, and they compound.

Part of the Cheap and fast track on lAItest.

The number on the pricing page is the price for a short request, on the middle tier, on a good day.

Most vendors sell more than one speed of the same model

Google sells Gemini in four tiers, not two: standard, batch, flex which is priced like batch, and priority which costs more than standard in exchange for being served ahead of everyone else. Same weights, four prices. A cost model that only knows about standard and batch is missing half the surface.

Long requests cost more per token, not just more tokens

Several vendors raise the rate once your input crosses a threshold. OpenAI does it on GPT-5.5, Google does it on Gemini, xAI does it on its current flagship, MiniMax does it on M3, and every one of them picked a different threshold. Kimi K3 is the notable exception: one flat rate across its entire window.

A common misconception

Commonly believed: The worst case is paying the priority premium on a request.

Actually: The worst case is paying it on a long request. On Gemini the priority multiplier and the long-context multiplier apply at the same time, so a large call served at priority bills at a rate several steps above the headline. Two multipliers compound. Most spreadsheets only count one of them.

The windows those thresholds are measured against

Read from a live model index at page-render time, each figure linked to the vendor page it came from.
WhatValueProvenance
Kimi K3 context1Msource, verified . One flat rate across all of it — unusual.
Claude Opus 5 context1Msource, verified .
Largest frontier context1.1Msource, verified .

In one sentence

A quoted price is a rate on one axis. Tier and request length are two more axes, and they multiply rather than add.