The pieces are not letters
The model never sees letters or words. It sees tokens: chunks of text drawn from a fixed list.
Part of the What is this thing? track on lAItest.
A model can write a decent poem about a strawberry and still miscount the letter r inside the word.
The reason is what it is allowed to see.
It reads in chunks, not letters.
Before a model touches your text, the text is cut into pieces called tokens. A token is usually a common word, a piece of a longer word, or a punctuation mark. "strawberry" may arrive as two or three chunks. The individual letters inside a chunk are about as visible to the model as the individual pen strokes in this sentence are to you.
Try it
Type anything and watch where the cuts land. Try your name, a long number, and a word in another language. This step is an interactive widget; open the lesson to use it.
Why anyone would do it this way
Letters would make every input enormously long and carry almost no meaning per step. Whole words would need a list containing every word in every language, plus every name and typo, and would still break on the next new word. Chunks are the compromise: common things are one piece, rare things get spelled out of several.
Roughly what is a token?
Answer: A chunk of text — often a whole word, often part of one. Tokens sit between letters and words. Short common words are usually one token; long or unusual words break into several. As a rough average, one token is about four characters of English text — which is why the count never matches your word count.
In one sentence
The model never sees your text. It sees a sequence of chunks, and everything it can and cannot do with spelling follows from that.