Lesson 16
Beyond text
Pictures and sound get chopped into chunks too.
When an AI “looks at” your photo or “listens to” your voice, nothing new happens inside it. The picture gets cut into squares and the sound gets cut into slices, each one becomes a chunk with an id, and those chunks go into the same window as your words, which is all multimodal means. Making a picture, though, is a completely different machine, called diffusion.
1Pick a sense
Do thisPress Next step and watch where the chunks come from. Notice that after the second step, this is lesson 01 again.
2Step through it
3The picture
4What the model has to work with
Real systems cut a picture into hundreds or thousands of squares, not sixteen, and the ids here are stand-ins. The shape of what happens is exactly this.