LLM learning

LLM Decoding & Sampling

Compare greedy decoding, temperature, and top-p through visible probability changes.

Not started3 min explanation

Visualize, practice, and deep-dive material are optional—use only what helps you learn.

Explanation

A focused 3-minute explanation using the topic's authored material.

Learning goals and prerequisites

After this lesson

  • Compare greedy, temperature, and top-p decoding
  • Choose settings from task requirements
  • Build repeatable decoding tests

Helpful before starting

  • How LLMs generate text
  • Basic probability intuition

Mental model

Decoding turns scores into a choice

Direct answer

The model produces raw next-token scores; decoding decides how those scores become the next visible token. Greedy decoding always chooses the largest probability, while sampling allows controlled variation among plausible candidates.

Follow the mechanism

  1. The model produces one logit for every vocabulary token.
  2. Temperature rescales the logits before softmax. Lower values magnify score differences; higher values compress them.
  3. Top-p keeps the smallest high-probability candidate set whose cumulative mass reaches the chosen threshold.
  4. The decoder selects or samples a token from the allowed set, appends it, and repeats with a new distribution.

Running example

For “The capital of France is”, use logits Paris: 3.8, Lyon: 1.5, and France: 0.8. At temperature 1, softmax gives approximately 87%, 9%, and 4%. At temperature 0.5, Paris rises to about 99%; at temperature 2, it falls to about 65% while Lyon and France become more likely. The model has not learned or forgotten anything between these settings. Only the selection distribution changed. With top-p at 0.90, the permitted set also changes as the temperature changes because cumulative probabilities are recalculated at every step.

What this does not mean

Higher temperature does not make a model more knowledgeable, and lower temperature does not make an answer correct. For a machine-consumed label or tool call, a closed schema and validator provide a stronger guarantee than temperature. For brainstorming, diversity can be useful, but it should still be measured against task quality and safety.

Key points

  • Logits → temperature → softmax → top-p set → selection is the decoding chain.
  • Temperature changes probability concentration; it does not add knowledge.
  • Top-p creates a different candidate set for each next-token distribution.
  • Choose decoding for the task, then enforce validity with explicit application controls.
Try it: Predict the distribution before moving a control

Use the Paris, Lyon, and France logits shown in the animation.

  1. 1.Predict whether Paris becomes more or less likely when temperature moves from 1 to 0.5.
  2. 2.Predict which candidates remain when top-p is reduced after the distribution becomes concentrated.
  3. 3.Run both changes one at a time and explain why neither result verifies the factual answer.

Was this lesson helpful?

Submit to the team when server feedback is available; otherwise this browser keeps a local copy and tells you so.