How Language Models Work

Lesson 03

Tokens

It never sees letters. It sees chunks.

Before anything else happens, your text gets chopped into pieces called tokens. A token is a chunk of a word, and the model stores each chunk as a plain id number. Your spelling never reaches it. Some of the most famous AI failures start right here.

A koi folded from teal paper, its body a run of separate faceted scales.
It never sees your letters. It sees chunks — scale after scale after scale — and those pieces, never your spelling, are what it reads.

1Pick some text (or type your own underneath)

Do thisIn step 2, flip between What you see and What the model sees. Try to count the r’s in strawberry in each view. Then type your own name into the box above and see how many pieces it costs.

2The same text, two ways

These splits are real: the page carries its own small tokenizer, with a fixed list of about 7,700 pieces, and it chops whatever you type by the same rule a production one uses: longest piece on the list wins. A real model’s list holds 50,000 to 200,000 pieces, so it keeps more words whole than this one does. Everything else about the mechanism is identical, including what happens to a word that isn’t on the list.