LLM learning

Embeddings & Context Windows

Learn how models represent text and how finite context changes application design.

Not started3 min explanation

Visualize, practice, and deep-dive material are optional—use only what helps you learn.

Explanation

A focused 3-minute explanation using the topic's authored material.

Learning goals and prerequisites

After this lesson

  • Distinguish token and retrieval embeddings
  • Construct a complete token budget
  • Choose long context or retrieval from evidence

Helpful before starting

  • LLM tokenization
  • Basic coordinate or similarity intuition

Mental model

Similarity chooses evidence; context carries it

Direct answer

An embedding represents an item as a learned vector so related items can be ranked by geometric similarity. A context window is the finite token budget available to one model request. Embeddings can help choose what enters that budget, but similarity and context capacity answer different questions.

Follow the mechanism

  1. Candidate passages are embedded and compared with the question.
  2. Authorization rules remove passages the user is not allowed to read.
  3. The application ranks the remaining passages and adds only those that fit after instructions, the question, and output capacity are reserved.
  4. The generation model receives the selected text as context and produces an answer; it does not see omitted passages.

Running example

Assume a teaching context limit of 18 tokens. Instructions, the question, and reserved output capacity use 7. A refund-policy passage costs 6 tokens with relevance 0.94, a returns FAQ costs 5 with relevance 0.82, and a shipping guide costs 5 with relevance 0.31. The first two exactly fill the budget: 7 + 6 + 5 = 18. The shipping guide is omitted even though it exists in the corpus. If the refund passage were unauthorized, it would need to be removed before ranking and budgeting rather than hidden after generation.

What this does not mean

High embedding similarity does not prove that a passage is true, authorized, current, or sufficient to support a claim. Likewise, a large advertised context window does not guarantee that the model will use every position equally well. Retrieval relevance, permissions, omissions, position effects, and answer support need separate tests.

Key points

  • Question → similarity ranking → authorization → token budget → model context is the application chain.
  • Embedding similarity is ranking evidence, not truth or permission.
  • Every context plan must reserve capacity for fixed instructions, the user input, and output.
  • Maximum context size and effective use of context must be measured separately.
Try it: Predict which evidence will fit

Use the 18-token refund example before moving the context-budget control.

  1. 1.Calculate which passages fit after the fixed 7-token allocation.
  2. 2.Predict what will be omitted when the limit falls below 18 and what risk that creates.
  3. 3.Run the animation, then explain why an omitted relevant passage is a context-construction failure rather than a generation failure.

Was this lesson helpful?

Submit to the team when server feedback is available; otherwise this browser keeps a local copy and tells you so.