Fine-tune or prompt?

Most problems people bring to fine-tuning turn out to be prompt, retrieval or model-choice problems in disguise.

Part of the How models are made track on lAItest.

The first question is not how to fine-tune. It is whether you want different behaviour, different knowledge, or a different model.

A common misconception

Commonly believed: The model does not know about our product, so we need to fine-tune it on our documentation.

Actually: Fine-tuning is a poor way to install facts. Facts change, and every change means another training run. The model gives no citation and no way to separate what it absorbed from what it invented. Putting the documents in front of the model at request time, in the prompt or retrieved, updates instantly, cites its source and is far easier to debug. Tune for behaviour. Retrieve for knowledge.

What fine-tuning is genuinely good at.

A consistent output format you cannot get reliably from instructions. A tone that takes three hundred words of prompt to describe. A narrow, high-volume classification where a small tuned model matches a large prompted one at a fraction of the cost and latency. The common thread is behaviour that is easier to demonstrate than to describe.

The long prompt you pay for on every single call

Read from a live model index at page-render time, each figure linked to the vendor page it came from.
WhatValueProvenance
Claude Sonnet 5, input$2 per 1Msource, verified .
Claude Haiku 4.5, input$1 per 1Msource, verified .
DeepSeek Flash, input$0.30 per 1Msource, verified . A tuned small model can beat a prompted large one on cost.
Median frontier input price$2 per 1Msource, verified .

Try it

Price the same job twice: a long prompt on a large model on every call, against a short prompt on a small one. This step is an interactive widget; open the lesson to use it.

In one sentence

Prompt first, retrieve for facts, change model for capability, and fine-tune last, for behaviour you can demonstrate but cannot describe.