## RAG

*Context and knowledge*

*Last edited: 28 September 2026*

RAG (retrieval-augmented generation) means finding passages in your documents and pasting them into the prompt before the question. The model answers from the text it was given instead of from memory, so its knowledge can be current, private and cited, but answer quality depends mostly on what retrieval finds.

**In plain words:** An open-book exam. Instead of relying on the student’s memory, you put the right page of the textbook in front of them just before the question. Give them the wrong page and they will answer wrongly with the same confidence, because they read what they were given.

*Interactive widget on the page: Step through the pipeline. See which chunks drop out at retrieval and reranking, and what changes in the answer.*

### Indexing and chunking

- The index is built ahead of time, off the question path: parsing documents, splitting them into chunks, embedding each chunk and storing it in a vector index and a text index. How embeddings and vector search work is covered in “Embeddings and vector search”. Parsing PDFs, tables and scans is a common source of errors: no retrieval will find text that was mangled during parsing. Layout-aware parsers and vision models help, or retrieval over page images skips text extraction altogether (ColPali, 2024).
- Chunk size is a trade-off. Too small loses context (“during this period you are entitled to…”, but which period?); too large blurs the match and eats the window. A starting point is a few hundred tokens, split along the document’s structure (headings, sections, paragraphs), not by character count. You settle the final size with evals.
- Contextual Retrieval (Anthropic, 2024): a model prepends 50–100 tokens of context from the whole document to each chunk before the chunk goes into both indexes. In their tests, combined with BM25 and reranking, this cut the share of relevant chunks missing from the top 20 (1 − recall@20) from 5.7% to 1.9%.
- Every chunk gets metadata: source, date, version, permissions. A changed document has to be re-indexed, a deleted one removed from the index, and a change of embedding model means re-embedding the whole collection.

### Retrieval and reranking

- Hybrid search (vectors plus BM25, merged with RRF) and cross-encoder reranking are covered in “Embeddings and vector search”. In the pipeline you decide the numbers: how many candidates to fetch (Anthropic took the top 150), how many chunks to put into the prompt after reranking (usually a few to a dozen or so) and how much latency you can spend on it.
- A chat message is often a poor query (“so what was the deal with that leave?”). It helps to have a model rewrite it into a standalone query that takes the conversation history into account.
- Permissions are filtered in the index query, based on the user’s identity and ACLs synced with the source, before a chunk reaches the prompt. An instruction like “don’t show confidential documents” in the prompt guarantees nothing, and documents can contain injected instructions (see “Prompt injection”).

### Query-side and index-side tricks

- HyDE (Gao et al., 2022): a model writes a hypothetical answer, and retrieval uses its embedding instead of the question’s. It helps when questions and documents are phrased completely differently, e.g. a casual question versus the language of a policy. It costs an extra model call before every search. It hurts when the model doesn’t know the domain: made-up names and numbers pull retrieval towards the wrong neighbours. Embedding models with a separate prefix or instruction for queries partly close this gap without an extra call.
- The reverse direction, on the index side: a model generates, for each chunk, the questions that chunk answers, and you index them alongside it (doc2query, 2019). You pay once, at indexing time, not on every question. The index grows, and the gain is only as large as the coverage of users’ real questions.
- Multi-query extends the query rewriting from the previous section: several versions of the query, searched in parallel and merged with RRF. It raises recall when you don’t know which words the answer is written in, at the cost of a model call and several searches. A multi-step question (“who has more leave, Anna’s team or Ben’s?”) is broken into sub-questions. When each one depends on the result of the previous one, it is already agentic retrieval.
- Small-to-big (parent-document retrieval): you index small chunks, because they give precise matches, and the model gets the larger whole, a section or a page, because it gives context. The price is prompt tokens, and several hits from the same section have to be merged into one passage.
- Late interaction (ColBERT, 2020): a vector for every token instead of one per chunk, and relevance is the sum of the best matches of the question’s tokens against the document’s tokens (MaxSim). It sits between a fast bi-encoder and an accurate cross-encoder, at the cost of an index many times larger.

### The prompt and citations

- Chunks go into the prompt with an identifier and a source, and the instruction says to answer only from them, cite the identifiers and admit when the answer isn’t there. Without that way out, the model fills the gaps from memory (see “Hallucinations”).
- With long material, the documents go before the question. Anthropic reports that putting the question at the end can improve quality by up to 30% on complex inputs with many documents. The variable chunks go after a stable system prompt so they don’t break the prompt cache (see “Prompt caching”).
- Check citations in code: was the cited chunk in the context, and is the quote really in it? APIs with built-in citations, such as Citations in Claude, return the cited text with a pointer to its span in the document, and the API guarantees the pointer is valid.

### Evaluation, agents and alternatives

- Retrieval and generation are evaluated separately. Retrieval: a set of questions with the right chunks labelled, and recall@k, the share of the labelled relevant chunks that make the top k (with one relevant chunk: whether it is there at all), with k equal to the number of chunks in the prompt. Generation: faithfulness (does every claim follow from the chunks provided) and completeness, usually scored by a model as judge that has been checked against human ratings (see “Evals”).
- Agentic retrieval: search is a tool, and the model composes the queries, reads the results and searches further on its own (see “The agent loop”). It handles multi-step questions and comparisons that a single search can’t, at the cost of several turns, time and tokens.
- Long context instead of RAG makes sense for a small, stable collection, especially with prompt caching. In 2024 Anthropic put the limit at about 200k tokens (roughly 500 pages), which was then Claude’s whole window. With 1M-token windows (September 2026), cost per question, latency and quality loss with length decide well before size does (see “The context window and agents”). RAG wins for a large or changing collection, and when citations and permissions matter. Fine-tuning teaches style and format, but it teaches facts unreliably, and they are hard to update afterwards (see “Fine-tuning and LoRA”).

### Check yourself

**Question:** How does RAG work, and where does it most often fail?

**Short answer:** RAG gives the model knowledge through its context instead of relying on its memory. Offline, documents are parsed, chunked along their structure, tagged with metadata and permissions, and indexed both as vectors and in BM25. At query time a hybrid search runs with a permission filter, a reranker picks the few best chunks, and the prompt says to answer only from them, with citations, and to admit when the answer is missing. What fails most often is parsing and retrieval, not the model: the right chunk simply is not in the prompt. So measure retrieval recall and answer faithfulness separately.

### Follow-up questions

- **Users report wrong answers. How do you find the cause?** Go through the traces stage by stage: was the right chunk in the search results, did it survive reranking, did it reach the prompt, and did the model use it? Missing from the results points to parsing, chunking or the query. Present but misused points to the prompt or the model. Stale content points to the indexing process.
- **RAG or long context?** Long context is simpler for a small, stable collection, especially with prompt caching. RAG wins for a large or changing collection on cost per question, latency, permissions and citations, and a shorter context also means less quality loss from length.
- **How do you handle permissions?** ACLs stored in each chunk’s metadata and synced with the source, a filter in the index query based on the user’s identity, and a test that someone without access doesn’t get the chunk. Never through an instruction in the prompt.
- **When do you use agentic retrieval instead of a single search?** When the question needs several steps or a comparison of sources and can’t be answered with one query. Start with a single search, measure which questions it fails on, and pay for the agent’s turns and latency only there.
- **When does HyDE hurt?** When the model doesn’t know the domain: the hypothetical answer contains made-up names, numbers and terms, so its embedding lands next to documents about something else. It also doesn’t help when searching for identifiers and exact phrases, where BM25 wins, or under a tight latency budget. Decide by comparing recall@k with and without HyDE on your own question set.

### Sources

- [Anthropic: Contextual Retrieval](https://www.anthropic.com/engineering/contextual-retrieval)
- [Faysse et al.: ColPali, retrieval over page images (arXiv, 2024)](https://arxiv.org/abs/2407.01449)
- [Lewis et al.: Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (arXiv)](https://arxiv.org/abs/2005.11401)
- [Gao et al.: Precise Zero-Shot Dense Retrieval without Relevance Labels, i.e. HyDE (arXiv)](https://arxiv.org/abs/2212.10496)
- [Claude docs: Citations](https://platform.claude.com/docs/en/build-with-claude/citations)

Interactive page: https://howaiworks.dev/rag/
