Mental model
How one token becomes a response
Direct answer
An LLM generates text by predicting one token at a time. At each step it scores the possible next tokens, a decoding rule selects one, and that token is appended to the context before the process repeats. The model is not retrieving a finished paragraph from storage.
Follow the mechanism
- The current text is converted into token IDs.
- The model turns those IDs into a logit, or raw score, for every token in its vocabulary.
- Softmax converts the logits into relative probabilities, and the decoding policy selects one candidate.
- The selected token becomes part of the next input. Generation ends only when a stop token, length limit, completed schema, or application rule says to stop.
Running example
Start with “The opposite of hot is”. Suppose the visible candidates are cold: 72%, warm: 18%, and <stop>: 10%. Greedy decoding selects “cold”, so the next context is “The opposite of hot is cold”. The model now computes a new distribution; it does not reuse the old one. If <stop> receives 81% at the next step, the application can stop with a complete answer. Changing the original context to “A mild day feels” would change the scores because the model is now answering a different continuation problem.
What this does not mean
A 72% next-token probability does not mean there is a 72% chance that the completed statement is factually true. It means “cold” is relatively likely under this model, this context, and this decoding setup. Facts still require evidence, calculation, a trusted tool, or review.
Key points
- Context → logits → probabilities → selected token → updated context is the full generation loop.
- Every selected token changes the input used for the next prediction.
- Token likelihood describes continuation behavior; it does not verify a claim.
- Stopping and validation are application responsibilities, not hidden model guarantees.
Try it: Predict the next state before running the animation
Use the “opposite of hot” example. Make a prediction before advancing the next-token animation.
- 1.Rank cold, warm, and <stop> for the initial context and explain the signal you used.
- 2.Predict which candidate should rise after “cold” is appended.
- 3.Run the animation, compare the observed distribution with your prediction, and name why the context changed.