Choosing between the three

Most tasks announce which kind of model they want, as soon as you ask whether you need a sentence or a decision.

Part of the Fast, slow, and neither track on lAItest.

Ask one question first: do I need prose, or do I need a choice?

The three, in one breath.

Write, summarise, explain, converse, generate code — a plain language model, because the alternatives literally cannot. A multi-step problem where checking beats guessing — a reasoning model at high effort, where the extra compute is demonstrably worth its cost. Classify, route, gate, triage or extract — a decision model, or something even simpler.

A common misconception

Commonly believed: For classification, the new decision model is obviously the right tool.

Actually: If you have labelled data, benchmark a small supervised model first: a fine-tuned embedding classifier, or even TF-IDF. In the independent runs published so far, which are small, single-author and not peer reviewed, those beat Jev on their own corpora. It is the cheapest, fastest, fully-owned option and it is the baseline almost everyone skips. The decision model earns its place when you do NOT have labelled data and your categories keep changing, because then you specify the task at request time and never train anything.

And the case where none of them is the answer.

If the decision has to be auditable, none of these gives you an audit trail. A reasoning model produces a rationale that may not be the real reason — that is the previous lesson. A System One model produces no rationale at all, by construction. A probability is not an explanation, and neither is a plausible paragraph.

You need to route support tickets into categories that change every few weeks, and you have no labelled data. What is the strongest argument for a decision model here?

Answer: You can change the categories at request time without retraining. Specifying the task in the request, with no training run, is the part of the claim nobody disputes. The speed and price multiples are contested and inconsistent even in the vendor’s own material; the "no training run" property is structural.

In one sentence

Generation and decision are different jobs. Pick by the shape of the output you need, measure on your own data, and distrust any comparison offered as a single multiple.