Small and on-device
Small models win on latency, privacy, offline operation and licence — not on difficulty. Self-hosting to save money often loses to the hosted floor.
Part of the Cheap and fast track on lAItest.
A speech model with 0.6 billion parameters transcribes audio 887 times faster than real time.
Small is not a consolation prize.
Small models are a different product, not a worse one
Alibaba shipped one open-weight generation in eight sizes, from 397 billion parameters down to 0.8 billion. A text-to-speech model called Kokoro has 82 million parameters, was trained for roughly a thousand dollars of compute, and runs where nothing larger will. These models win where the constraint is latency, privacy, offline operation or licence. They do not win where the constraint is difficulty.
On a device, the ceiling is memory and it is unforgiving
A three-billion-parameter speech model like Orpheus wants roughly eight to twelve gigabytes of video memory and falls apart if it has to spill out of it. There is no autoscaling on a phone. Retrieval followed the same path: Qdrant ships an embedded edition for robots and kiosks that runs with no background service and no network at all.
A common misconception
Commonly believed: Self-hosting a small open-weight model is how you cut the bill.
Actually: Sometimes. Check the floor first. Hosted small models are already priced far under the flagships, and a rented GPU bills by the hour whether or not anyone is using it. Self-hosting pays off on steady high volume, on data that cannot leave your network, and where there is no network. It rarely pays off on spiky traffic.
The hosted floor you would be trying to beat
| What | Value | Provenance |
|---|---|---|
| DeepSeek Flash input | $0.30 per 1M | source, verified . |
| MiniMax M3 input | $0.30 per 1M | source, verified . |
| Claude Haiku 4.5 input | $1 per 1M | source, verified . |
| Cheapest frontier input | $0.75 per 1M | source, verified . A rented GPU has to beat this while sitting idle too. |
In one sentence
Choose a small model for where it has to run, not for what it costs. The cost argument is usually the weaker one.