How Language Models Work

Lesson 16

Beyond text

Pictures and sound get chopped into chunks too.

When an AI “looks at” your photo or “listens to” your voice, nothing new happens inside it. The picture gets cut into squares and the sound gets cut into slices, each one becomes a chunk with an id, and those chunks go into the same window as your words, which is all multimodal means. Making a picture, though, is a completely different machine, called diffusion.

An owl folded from violet paper, wide facial discs turned forward.
It sees and it hears, and past the first step neither is special. A picture and a voice arrive as the same kind of chunk your words do.

1Pick a sense

Do thisPress Next step and watch where the chunks come from. Notice that after the second step, this is lesson 01 again.

2Step through it

3The picture

4What the model has to work with

Real systems cut a picture into hundreds or thousands of squares, not sixteen, and the ids here are stand-ins. The shape of what happens is exactly this.