Mental model
Text becomes model-specific pieces
Direct answer
A tokenizer converts visible text into a sequence of integer IDs that a model can process. A token can be a whole word, part of a word, punctuation, whitespace, or a byte-like piece, so tokens are not the same thing as words or characters.
Follow the mechanism
- The tokenizer applies its configured normalization rules to the input text.
- It finds a sequence of known vocabulary pieces that represents the text.
- Each piece is replaced by its vocabulary ID and sent to the model as one position.
- Decoding uses the same vocabulary to turn IDs back into text.
Running example
Use the sentence “GraphRAG answers with cited evidence.” A transparent teaching tokenizer may expose Graph, RAG, spaces, ordinary words, and the final period as separate pieces. Another production tokenizer may split “GraphRAG” differently because it learned a different vocabulary. The sentence still looks like five words to a reader, yet the model may receive many more token positions. That difference affects context capacity, latency, and cost. It also explains why punctuation-heavy identifiers, source code, names, and multilingual text deserve their own tests.
What this does not mean
A token ID has no portable meaning by itself. ID 1842 can represent different pieces in different vocabularies, and cached IDs from one tokenizer must not be sent to a model that expects another. Word count and character count can help with rough planning, but only the deployed tokenizer can enforce the real request budget.
Key points
- Text → normalization → vocabulary pieces → token IDs is the tokenizer interface.
- Token boundaries depend on the exact tokenizer and vocabulary version.
- More pieces consume more context positions even when the visible meaning is unchanged.
- Budget and compatibility tests must use the tokenizer paired with the deployed model.
Try it: Predict the boundaries before revealing them
Before using the tokenizer explorer, mark where you think “GraphRAG answers with cited evidence.” will split.
- 1.Mark predicted boundaries for GraphRAG, the spaces, and the final period.
- 2.Run the explorer and record where its transparent teaching rules disagree with your word boundaries.
- 3.Change GraphRAG to Graph-RAG and explain why a production budget test must include both strings.