How Language Models Work

Lesson 10

Embeddings

Every word becomes a dot on a map. Similar words sit close together.

Each token id gets turned into a long list of numbers. Think of that list as an address: it parks the word at one exact spot on a huge map. Nobody ever tells the model what “dog” means. It just puts “dog” next to “puppy.” That map is called an embedding.

A butterfly folded from cobalt blue paper, wings open and matched.
Every word it knows, parked somewhere on one enormous map. Nobody placed them: words used alike end up neighbours, the way two wings match without being told to.

Do thisTap any word on the map in step 1. The “closest to” panel then lists the five words parked nearest it, which is this model’s entire idea of “means something similar.”

1The map (40 words, placed by meaning)

Swipe the map sideways to see the rest of the words.

2Word math (because directions on the map mean something too)

The arrow from one word to another is a direction, and a direction can be reused. Press either button below: it picks up the arrow that runs from man to woman, puts its tail on king, and shows you where the tip lands.