Mental model
Similarity chooses evidence; context carries it
Direct answer
An embedding represents an item as a learned vector so related items can be ranked by geometric similarity. A context window is the finite token budget available to one model request. Embeddings can help choose what enters that budget, but similarity and context capacity answer different questions.
Follow the mechanism
- Candidate passages are embedded and compared with the question.
- Authorization rules remove passages the user is not allowed to read.
- The application ranks the remaining passages and adds only those that fit after instructions, the question, and output capacity are reserved.
- The generation model receives the selected text as context and produces an answer; it does not see omitted passages.
Running example
Assume a teaching context limit of 18 tokens. Instructions, the question, and reserved output capacity use 7. A refund-policy passage costs 6 tokens with relevance 0.94, a returns FAQ costs 5 with relevance 0.82, and a shipping guide costs 5 with relevance 0.31. The first two exactly fill the budget: 7 + 6 + 5 = 18. The shipping guide is omitted even though it exists in the corpus. If the refund passage were unauthorized, it would need to be removed before ranking and budgeting rather than hidden after generation.
What this does not mean
High embedding similarity does not prove that a passage is true, authorized, current, or sufficient to support a claim. Likewise, a large advertised context window does not guarantee that the model will use every position equally well. Retrieval relevance, permissions, omissions, position effects, and answer support need separate tests.
Key points
- Question → similarity ranking → authorization → token budget → model context is the application chain.
- Embedding similarity is ranking evidence, not truth or permission.
- Every context plan must reserve capacity for fixed instructions, the user input, and output.
- Maximum context size and effective use of context must be measured separately.
Try it: Predict which evidence will fit
Use the 18-token refund example before moving the context-budget control.
- 1.Calculate which passages fit after the fixed 7-token allocation.
- 2.Predict what will be omitted when the limit falls below 18 and what risk that creates.
- 3.Run the animation, then explain why an omitted relevant passage is a context-construction failure rather than a generation failure.