LLM learning

LLM Tokenization

Understand how text becomes versioned token IDs, context usage, and model cost.

Not started3 min explanation

Visualize, practice, and deep-dive material are optional—use only what helps you learn.

Explanation

A focused 3-minute explanation using the topic's authored material.

Learning goals and prerequisites

After this lesson

  • Explain why tokens are not words
  • Inspect domain-specific token splits
  • Version and test a tokenizer contract

Helpful before starting

  • How LLMs generate text
  • Basic familiarity with words and characters

Mental model

Text becomes model-specific pieces

Direct answer

A tokenizer converts visible text into a sequence of integer IDs that a model can process. A token can be a whole word, part of a word, punctuation, whitespace, or a byte-like piece, so tokens are not the same thing as words or characters.

Follow the mechanism

  1. The tokenizer applies its configured normalization rules to the input text.
  2. It finds a sequence of known vocabulary pieces that represents the text.
  3. Each piece is replaced by its vocabulary ID and sent to the model as one position.
  4. Decoding uses the same vocabulary to turn IDs back into text.

Running example

Use the sentence “GraphRAG answers with cited evidence.” A transparent teaching tokenizer may expose Graph, RAG, spaces, ordinary words, and the final period as separate pieces. Another production tokenizer may split “GraphRAG” differently because it learned a different vocabulary. The sentence still looks like five words to a reader, yet the model may receive many more token positions. That difference affects context capacity, latency, and cost. It also explains why punctuation-heavy identifiers, source code, names, and multilingual text deserve their own tests.

What this does not mean

A token ID has no portable meaning by itself. ID 1842 can represent different pieces in different vocabularies, and cached IDs from one tokenizer must not be sent to a model that expects another. Word count and character count can help with rough planning, but only the deployed tokenizer can enforce the real request budget.

Key points

  • Text → normalization → vocabulary pieces → token IDs is the tokenizer interface.
  • Token boundaries depend on the exact tokenizer and vocabulary version.
  • More pieces consume more context positions even when the visible meaning is unchanged.
  • Budget and compatibility tests must use the tokenizer paired with the deployed model.
Try it: Predict the boundaries before revealing them

Before using the tokenizer explorer, mark where you think “GraphRAG answers with cited evidence.” will split.

  1. 1.Mark predicted boundaries for GraphRAG, the spaces, and the final period.
  2. 2.Run the explorer and record where its transparent teaching rules disagree with your word boundaries.
  3. 3.Change GraphRAG to Graph-RAG and explain why a production budget test must include both strings.

Was this lesson helpful?

Submit to the team when server feedback is available; otherwise this browser keeps a local copy and tells you so.