Who decides where to cut
The tokenizer is a frozen list of chunks, learned by merging the most common pairs of characters over and over.
Part of the What is this thing? track on lAItest.
The list of chunks was built by a program that started with single characters and kept gluing together whichever pair it saw most.
That is very nearly the entire algorithm.
Byte pair encoding
Start with an alphabet of single characters. Scan a mountain of text for the most common neighbouring pair — say t followed by h. Merge it into one new symbol, th. Scan again. Merge again. Do that tens of thousands of times. What is left is a vocabulary: a fixed, ordered list of chunks. This is byte pair encoding, published in 2015, and it is still the backbone of how models read.
Try it
Compare plain English with digits, punctuation and a non-Latin script. Watch which ones cost far more chunks. This step is an interactive widget; open the lesson to use it.
A common misconception
Commonly believed: Tokenization is a shared standard, so a sentence is worth the same number of tokens everywhere.
Actually: Every model family has its own vocabulary, built from its own text. The same sentence splits differently on different models, and makers change the vocabulary between versions — Anthropic did exactly that for its newer Claude models, which now count the same English text higher than the older ones did. Any token budget you memorise belongs to one model, not to text in general.
The unfairness that gets baked in
The vocabulary is learned from whatever text the makers fed it. Text that resembles that text is cheap; everything else is expensive. English prose lands in few chunks. Other scripts, unusual names and long numbers get spelled out piece by piece. Since usage is billed per token, some languages simply cost more money to say the same thing in.
In one sentence
A tokenizer is a frozen list of chunks learned by merging common pairs — and whatever it saw most is what is cheapest to say.