Stop sequences and max_tokens

A model has no sense of how long your answer should be. Two settings decide when it ends.

Part of the Turning the dials track on lAItest.

Left alone, a model will keep producing tokens. Something outside it has to say when.

max_tokens is a hard ceiling. A stop sequence is a trapdoor.

max_tokens caps how many tokens one response may contain. Hit the cap and generation is cut off wherever it happened to be, often mid-word; the text is not wrong, it is truncated, and the response tells you that is why it ended. A stop sequence is a string you nominate in advance. The moment the model produces it, generation ends and the sequence itself is not returned.

Output ceilings vary far more than context windows do

Read from a live model index at page-render time, each figure linked to the vendor page it came from.
WhatValueProvenance
Claude Opus 5, max output128Ksource, verified .
Claude Haiku 4.5, max output64Ksource, verified .
DeepSeek V4-Pro, max output384Ksource, verified . The widest published output ceiling in the index.
Cohere Command A+, max output64Ksource, verified .

A common misconception

Commonly believed: If I leave max_tokens alone I get whatever length the model is capable of.

Actually: Not always. At least one major provider defaults this ceiling far below what its model can actually produce, so long answers are silently truncated until you set the value yourself. Check the default for the API you call rather than assuming it matches the maximum in the docs.

A response comes back cut off in the middle of a word. What most likely happened?

Answer: You hit the max_tokens ceiling. A stop sequence ends at a boundary you chose, so it never lands mid-word. Running out of context is an error raised before generation starts, not a cut halfway through. A response that dies mid-word is the ceiling, and the API reports that as the reason it stopped.

In one sentence

max_tokens is a budget, not a target — and a truncated answer is a settings bug, not a model failure.