How Language Models Work

Lesson 02

Rolling the dice

Ask the same thing twice and you get two different answers.

Lesson 01 ended with a ranked list of guesses. Something still has to pick one, and that picker rolls dice. Two dials control how loaded those dice are. They are the reason a chatbot can answer the same question differently every time you hit refresh.

A fox folded from burnt orange paper, standing alert with its tail out.
Same question in, a different answer out. The dial decides how far from the safe path it is willing to wander, and a fox is what wandering looks like.

1Pick a question to ask, over and over

Do thisPress Ask it 25 times in step 4 and see how many different answers come back. Then drag the risk dial in step 3 down to 0 and ask 25 more.

2The ranked list it produces (the same one, every single time)

    3The two dials that load the dice

    Called temperature. It flattens or sharpens the bars in step 2. At 0 the top answer wins every single time.

    Called top-p. Go down the list adding up the chances, and stop there. Everything below the line is struck out and can’t be picked at all.

    4Now roll the dice against that list

      A real model ranks every chunk it knows, roughly 100,000 of them. These seven are the top of that list, with the rest left out so the bars stay readable.