Mental model
Decoding turns scores into a choice
Direct answer
The model produces raw next-token scores; decoding decides how those scores become the next visible token. Greedy decoding always chooses the largest probability, while sampling allows controlled variation among plausible candidates.
Follow the mechanism
- The model produces one logit for every vocabulary token.
- Temperature rescales the logits before softmax. Lower values magnify score differences; higher values compress them.
- Top-p keeps the smallest high-probability candidate set whose cumulative mass reaches the chosen threshold.
- The decoder selects or samples a token from the allowed set, appends it, and repeats with a new distribution.
Running example
For “The capital of France is”, use logits Paris: 3.8, Lyon: 1.5, and France: 0.8. At temperature 1, softmax gives approximately 87%, 9%, and 4%. At temperature 0.5, Paris rises to about 99%; at temperature 2, it falls to about 65% while Lyon and France become more likely. The model has not learned or forgotten anything between these settings. Only the selection distribution changed. With top-p at 0.90, the permitted set also changes as the temperature changes because cumulative probabilities are recalculated at every step.
What this does not mean
Higher temperature does not make a model more knowledgeable, and lower temperature does not make an answer correct. For a machine-consumed label or tool call, a closed schema and validator provide a stronger guarantee than temperature. For brainstorming, diversity can be useful, but it should still be measured against task quality and safety.
Key points
- Logits → temperature → softmax → top-p set → selection is the decoding chain.
- Temperature changes probability concentration; it does not add knowledge.
- Top-p creates a different candidate set for each next-token distribution.
- Choose decoding for the task, then enforce validity with explicit application controls.
Try it: Predict the distribution before moving a control
Use the Paris, Lyon, and France logits shown in the animation.
- 1.Predict whether Paris becomes more or less likely when temperature moves from 1 to 0.5.
- 2.Predict which candidates remain when top-p is reduced after the distribution becomes concentrated.
- 3.Run both changes one at a time and explain why neither result verifies the factual answer.