Fine-tune or prompt?
Most problems people bring to fine-tuning turn out to be prompt, retrieval or model-choice problems in disguise.
Part of the How models are made track on lAItest.
The first question is not how to fine-tune. It is whether you want different behaviour, different knowledge, or a different model.
A common misconception
Commonly believed: The model does not know about our product, so we need to fine-tune it on our documentation.
Actually: Fine-tuning is a poor way to install facts. Facts change, and every change means another training run. The model gives no citation and no way to separate what it absorbed from what it invented. Putting the documents in front of the model at request time, in the prompt or retrieved, updates instantly, cites its source and is far easier to debug. Tune for behaviour. Retrieve for knowledge.
What fine-tuning is genuinely good at.
A consistent output format you cannot get reliably from instructions. A tone that takes three hundred words of prompt to describe. A narrow, high-volume classification where a small tuned model matches a large prompted one at a fraction of the cost and latency. The common thread is behaviour that is easier to demonstrate than to describe.
The long prompt you pay for on every single call
| What | Value | Provenance |
|---|---|---|
| Claude Sonnet 5, input | $2 per 1M | source, verified . |
| Claude Haiku 4.5, input | $1 per 1M | source, verified . |
| DeepSeek Flash, input | $0.30 per 1M | source, verified . A tuned small model can beat a prompted large one on cost. |
| Median frontier input price | $2 per 1M | source, verified . |
Try it
Price the same job twice: a long prompt on a large model on every call, against a short prompt on a small one. This step is an interactive widget; open the lesson to use it.
In one sentence
Prompt first, retrieve for facts, change model for capability, and fine-tune last, for behaviour you can demonstrate but cannot describe.