Alignment
A model that will do anything you ask is a product decision, and so is one that refuses. Alignment is the name for that decision and the work behind it.
Part of the Trust and evals track on lAItest.
A model that will do anything you ask is not a neutral tool. Somebody decided that, the same way somebody decides a refusal.
Alignment is the name for the decision and the work behind it.
What the word covers
A pretrained model predicts text and holds no view about what it is used for. Alignment is the training and the machinery that make a deployed model behave the way its makers intend: helpful in the ways they chose, unhelpful in the ways they chose. Most of it happens during training, through feedback from humans and from other models. A thinner layer sits around the model in production as classifiers that watch what goes in and what comes out.
A refusal is a state your code has to handle
On the newest frontier Claude models a declined request is not an error. It arrives as a normal successful response with a stop reason of refusal and a category attached: cyber, bio, reasoning extraction, frontier LLM, or none. Code that reaches straight into the first block of the response breaks on it. There is even a billing consequence: a refusal that happens before any output is not billed at all, while one that arrives mid-stream is billed for the part already sent.
A common misconception
Commonly believed: Alignment is a filter bolted on top. Underneath sits a raw model that would answer anything, and a jailbreak peels the filter off.
Actually: Most of it is in the weights. The model was trained to behave this way and the behaviour is not a component you can detach. The classifier layer is genuinely separable, though, and that is visible in the market: in 2026 one lab shipped two models with the same underlying capabilities, one with safety classifiers and one without, the second available only by invitation for defensive security work. The layer moved. The training did not.
Your code reads the first content block of every API response. A frontier model declines a request. What happens?
Answer: You get a normal successful response with a refusal stop reason, and your code breaks on it. A refusal is a successful call that produced a different shape. The status is a normal 200, the stop reason says refusal, and a category names the policy that fired. Treating safety behaviour as an error path is a reliable way to ship a crash; it is a state to handle, like any other stop reason.
In one sentence
Alignment is not a switch on the side of the model. It is training, plus a classifier layer, plus product decisions you can see in the API.