Skip to content

31 - RAG (Retrieval-Augmented Generation)

Quick reference for building RAG systems: loading documents, chunking, embedding, retrieval, reranking, prompting with citations, evaluating and improving answer quality.

Last verified: 2026-09-27. For newer changes, check the Official docs links in the Introduction.

Introduction

Before you start

You should know: how to call an LLM with a prompt (27, 28) and how embeddings find similar text (30).

The problem it solves: an LLM does not know your company's documents, policies or anything newer than its training data, and when it does not know, it may invent an answer. You want answers grounded in your documents, with sources, that update when the documents change.

Before RAG: the options were retraining or fine-tuning a model on your documents (expensive, slow to update, and the model still cannot cite where a fact came from) or pasting whole documents into every prompt (too long and too costly). RAG (named in a 2020 research paper) retrieves only the relevant passages at question time and hands them to the model.

Think of it like: an open-book exam (the mental model below). The student does not memorise the textbook; a librarian finds the right pages for each question, and the student answers from them and cites them.

What is RAG?

Retrieval-Augmented Generation means: before the LLM answers, your app retrieves the most relevant pieces of your own documents and puts them into the prompt, so the model generates its answer from that material. It is how you build "chat with your documents", support bots that know your policies, and assistants over company knowledge, without retraining the model.

Mental model: an open-book exam

An LLM on its own answers from memory (training data): it may not know your documents and may make things up. RAG turns it into an open-book exam: a librarian (the retriever) finds the right pages, and the student (the LLM) answers using those pages, citing them.

=========================== INDEXING (offline, when documents change) ===========================

 PDFs, Word, web pages,   LOAD        CHUNK              EMBED               STORE
 Confluence, SQL rows  -> text  ->  ~300-800 token  ->  vector per    ->  vector DB
                                    pieces + metadata   chunk              (text + vector + metadata)

============================ QUERYING (online, for every question) =============================

 "What is our refund      EMBED       RETRIEVE           (RERANK)          GENERATE
  window for jackets?" -> query   ->  top 20 similar  ->  best 5 by a   ->  LLM prompt:
                          vector       chunks (+ filter)  reranker          instructions + chunks + question
                                                                              |
                                                                              v
                                                   "Jackets can be returned within 30 days [source: policy.pdf p.3]"

Two separate quality problems to keep in mind:

  1. Retrieval quality: did we find the right chunks? (If not, the best LLM cannot answer.)
  2. Generation quality: did the LLM use them correctly, completely and honestly?

Why use RAG?

  • Your private / up-to-date knowledge without fine-tuning.
  • Fewer hallucinations: answers grounded in real sources, with citations users can check.
  • Easy updates: change a document, re-index it; no retraining.
  • Access control: retrieve only documents the user may see.
  • Cheaper than long context for large knowledge bases (send 5 chunks, not 5,000 pages).

Key terms

Term Meaning
Corpus / knowledge base All documents you can retrieve from
Chunk A piece of a document stored and retrieved as a unit
Chunk overlap Repeated text between neighbouring chunks so ideas are not cut in half
Retriever Component that finds relevant chunks for a query
Top-k Number of chunks retrieved
Reranker Model that re-scores retrieved chunks more precisely
Hybrid search Vector + keyword search combined
Grounding / faithfulness Answer supported by the retrieved text
Citation Reference to the source of each claim
Agentic RAG The LLM decides when and what to search, using a search tool

Where it fits: built on 30 - Embeddings and Vector DBs, 27 - LLM APIs and 28 - Prompt Engineering; agentic version with 29 - Tool Use / 32 - AI Agents; frameworks in 33; measured with 35 - Evals; served via 40 - FastAPI with a UI from 39.

Official docs

Where to read the latest, authoritative documentation:

Resource Link
Anthropic: contextual retrieval https://www.anthropic.com/news/contextual-retrieval
Claude citations https://platform.claude.com/docs/en/build-with-claude/citations
LlamaIndex https://docs.llamaindex.ai/
Ragas (RAG evaluation) https://docs.ragas.io/
Docling (document parsing) https://docling-project.github.io/docling/
pypdf https://pypdf.readthedocs.io/

Contents

  1. When to Use RAG (and When Not)
  2. Loading Documents
  3. Cleaning Text
  4. Chunking Strategies
  5. Chunking in Code
  6. Metadata
  7. Embedding and Indexing
  8. Retrieval
  9. Reranking
  10. The Generation Prompt (with Citations)
  11. Complete Minimal RAG App
  12. Improving Retrieval
  13. Contextual Retrieval
  14. Agentic RAG
  15. Conversational RAG (Follow-up Questions)
  16. Structured Data: Text-to-SQL
  17. Evaluating RAG
  18. Keeping the Index Fresh
  19. Security and Permissions
  20. RAG Frameworks
  21. Common Failures and Fixes
  22. Try It

1. When to Use RAG (and When Not)

Choosing between RAG, long context, fine-tuning and tools. Match the size and nature of the knowledge to the technique.

Use it before building anything.

Situation Best approach
Few documents that fit easily in the context window Put them all in the prompt (+ prompt caching)
Large or growing knowledge base (hundreds+ docs) RAG
Data in databases / APIs (orders, metrics) Tools / text-to-SQL, not embeddings
Need a specific style / format / narrow skill Prompting, then fine-tuning (37)
Frequently changing facts RAG (re-index) or live tools
Open-ended research across sources Agentic RAG (32)

2. Loading Documents

Getting plain text (and structure) out of files. Use a loader per file type; keep page numbers, headings and source paths as metadata.

Use it at the start of the indexing pipeline.

from pathlib import Path

from pypdf import PdfReader                    # pip install pypdf


def load_pdf(path: Path) -> list[dict]:
    """Return one record per page with text and metadata."""
    reader = PdfReader(path)
    return [
        {"text": page.extract_text() or "", "source": path.name, "page": i + 1}
        for i, page in enumerate(reader.pages)
    ]


def load_markdown(path: Path) -> list[dict]:
    return [{"text": path.read_text(encoding="utf-8"), "source": path.name, "page": None}]


records = []
for p in Path("docs").rglob("*"):
    if p.suffix == ".pdf":
        records += load_pdf(p)
    elif p.suffix in {".md", ".txt"}:
        records += load_markdown(p)
File type Libraries
PDF (text) pypdf, pymupdf
PDF with tables / layout, scans docling, unstructured, OCR (pytesseract), or send pages to a vision LLM
Word / PowerPoint python-docx, python-pptx, docling
HTML / web trafilatura, beautifulsoup4
Many formats unstructured, markitdown, LlamaIndex / LangChain loaders

3. Cleaning Text

Removing noise that hurts retrieval. Normalise whitespace, drop headers / footers / page numbers, fix broken hyphenation, remove boilerplate.

Use it after loading, before chunking.

import re


def clean(text: str) -> str:
    text = re.sub(r"-\n(\w)", r"\1", text)           # join words hyphenated across lines
    text = re.sub(r"[ \t]+", " ", text)              # collapse spaces
    text = re.sub(r"\n{3,}", "\n\n", text)           # max one blank line
    text = re.sub(r"Page \d+ of \d+", "", text)      # page footers
    return text.strip()

Look at your extracted text before indexing; garbage in = garbage retrieved.

4. Chunking Strategies

Splitting documents into retrievable pieces. Chunks should be big enough to contain a complete idea and small enough to be specific; overlap avoids cutting ideas in half.

Use this when every RAG index. Chunking is one of the biggest quality levers.

Strategy How Good for
Fixed size + overlap Every N tokens / characters, overlap 10 to 20% Simple baseline
Recursive / by separators Split by paragraphs, then sentences, until under size General text (default choice)
By structure Split on Markdown / HTML headings, sections, pages Manuals, docs, wikis
Semantic Split where the topic changes (embedding similarity drops) Long flowing text
Per record One chunk per FAQ entry / ticket / product Naturally separate items
Parent-child Retrieve small chunks, send the larger parent section to the LLM Precision + context

Starting point: 300 to 800 tokens per chunk, 10 to 20% overlap, split on paragraph / heading boundaries, and add the document title / section heading to each chunk. Then measure and tune.

5. Chunking in Code

A simple, dependency-free recursive chunker. Split by paragraphs; merge paragraphs until the size limit; carry overlap into the next chunk.

Use it for learning and small projects (frameworks have ready-made splitters).

CHUNK_CHARS = 2000          # about 500 tokens of English
OVERLAP_CHARS = 300


def chunk_text(text: str, size: int = CHUNK_CHARS, overlap: int = OVERLAP_CHARS) -> list[str]:
    """Split text on paragraph boundaries into chunks of about `size` characters."""
    paragraphs = [p.strip() for p in text.split("\n\n") if p.strip()]
    chunks, current = [], ""
    for para in paragraphs:
        if len(current) + len(para) + 2 <= size:
            current = f"{current}\n\n{para}" if current else para
            continue
        if current:
            chunks.append(current)
            current = current[-overlap:] + "\n\n" + para   # overlap keeps context across the boundary
        else:
            current = para
        while len(current) > size:                       # a single huge paragraph: hard split
            chunks.append(current[:size])
            current = current[size - overlap:]
    if current:
        chunks.append(current)
    return chunks

6. Metadata

Extra fields stored with each chunk. Keep source, page, section, date, language, access level, a stable chunk ID.

Use it for citations, filtering, permissions, updating / deleting documents.

chunks = []
for rec in records:
    for i, piece in enumerate(chunk_text(clean(rec["text"]))):
        chunks.append({
            "id": f"{rec['source']}::p{rec['page']}::c{i}",   # stable ID -> easy updates
            "text": piece,
            "source": rec["source"],
            "page": rec["page"],
            "tenant_id": "acme",
        })

7. Embedding and Indexing

Turning chunks into vectors and storing them. Batch-embed chunk texts; store id, text, vector and metadata in a vector store.

Use it for initial build and whenever documents change.

import chromadb
from sentence_transformers import SentenceTransformer

embedder = SentenceTransformer("BAAI/bge-m3")                        # multilingual, strong
db = chromadb.PersistentClient(path="./rag_db")
col = db.get_or_create_collection("kb", metadata={"hnsw:space": "cosine"})

BATCH = 64
for start in range(0, len(chunks), BATCH):
    batch = chunks[start:start + BATCH]
    col.upsert(
        ids=[c["id"] for c in batch],
        documents=[c["text"] for c in batch],
        embeddings=embedder.encode([c["text"] for c in batch], normalize_embeddings=True).tolist(),
        metadatas=[{k: v for k, v in c.items() if k not in ("id", "text") and v is not None} for c in batch],
    )

8. Retrieval

Finding the chunks most relevant to the question. Embed the query with the same model, search top-k with filters; optionally combine with keyword search.

Use it in every question.

TOP_K = 20


def retrieve(question: str, tenant_id: str, k: int = TOP_K) -> list[dict]:
    q_vec = embedder.encode(question, normalize_embeddings=True).tolist()
    res = col.query(query_embeddings=[q_vec], n_results=k, where={"tenant_id": tenant_id})
    return [
        {"text": doc, "source": meta["source"], "page": meta.get("page"), "distance": dist}
        for doc, meta, dist in zip(res["documents"][0], res["metadatas"][0], res["distances"][0])
    ]

Typical settings: retrieve 10 to 30 candidates, then rerank down to 3 to 8 for the prompt.

9. Reranking

A second, more accurate scoring of retrieved chunks. A cross-encoder reads the question and each chunk together and scores relevance (slower than embeddings, so only on the top candidates).

It is almost always worth it once basic RAG works; big quality gain for little code.

from sentence_transformers import CrossEncoder

reranker = CrossEncoder("BAAI/bge-reranker-v2-m3")
FINAL_K = 5


def rerank(question: str, candidates: list[dict], k: int = FINAL_K) -> list[dict]:
    scores = reranker.predict([(question, c["text"]) for c in candidates])
    ranked = sorted(zip(candidates, scores), key=lambda x: x[1], reverse=True)
    return [c for c, _ in ranked[:k]]

Hosted rerankers exist too (Cohere, Voyage, Azure AI Search semantic ranker).

10. The Generation Prompt (with Citations)

Asking the LLM to answer from the retrieved chunks only, and cite them. Put numbered, tagged sources first, the question last, rules for citations and for missing information.

Use it in every RAG answer.

RAG_SYSTEM = (
    "You answer questions for Acme employees using only the provided sources. "
    "Cite sources after each claim like [1] or [2][3]. "
    "If the sources do not contain the answer, say you could not find it in the documentation "
    "and suggest who to ask. Do not use outside knowledge."
)


def build_prompt(question: str, chunks: list[dict]) -> str:
    sources = "\n".join(
        f'<source id="{i}" file="{c["source"]}" page="{c["page"]}">\n{c["text"]}\n</source>'
        for i, c in enumerate(chunks, start=1)
    )
    return f"<sources>\n{sources}\n</sources>\n\nQuestion: {question}"

Provider features like Claude's citations on document blocks can return exact cited spans automatically (27).

11. Complete Minimal RAG App

All steps together in one small program. Retrieve -> rerank -> prompt -> Claude answer with sources.

Use it as a template for your first RAG project.

import anthropic

client = anthropic.Anthropic()


def answer(question: str, tenant_id: str = "acme") -> dict:
    """Answer a question from the knowledge base with citations."""
    candidates = retrieve(question, tenant_id)
    if not candidates:
        return {"answer": "No documents found.", "sources": []}
    top = rerank(question, candidates)
    response = client.messages.create(
        model="claude-opus-5",
        max_tokens=16000,
        system=RAG_SYSTEM,
        messages=[{"role": "user", "content": build_prompt(question, top)}],
    )
    text = "".join(b.text for b in response.content if b.type == "text")
    return {"answer": text, "sources": [{"id": i, "file": c["source"], "page": c["page"]}
                                         for i, c in enumerate(top, start=1)]}


print(answer("How many days do customers have to return a jacket?"))

Serve it with FastAPI (40) and a chat UI (39).

12. Improving Retrieval

Techniques to find better chunks. Change the query, the index or the ranking; measure each change.

Use it when the right answer exists in the docs but is not retrieved.

Technique How Helps when
Hybrid search Vector + BM25 keyword search, merged Exact codes, names, jargon
Reranking Cross-encoder on top-k Right chunk retrieved but ranked low
Better chunking Structure-aware, right size, headings in chunks Answers split across chunks
Query rewriting LLM rewrites the question into a clear search query Vague / conversational questions
Multi-query LLM generates 3 to 5 query variants, merge results Different wording in docs
HyDE Embed a hypothetical answer instead of the question Short questions, long documents
Metadata filters Date / product / language filters Too many near-duplicate docs
Parent-child Retrieve small, send larger surrounding section Precise match but missing context
Better embedding model Stronger or domain / multilingual model Generally poor matches

13. Contextual Retrieval

Adding a short explanation of each chunk's context before embedding it. For each chunk, an LLM writes 1 to 2 sentences situating it in the whole document ("This section of the 2025 refund policy covers jackets..."); prepend that to the chunk text for embedding and keyword indexing.

Use it for chunks that are ambiguous on their own ("It must be returned within 30 days" - what is "it"?). Use prompt caching on the full document to keep the cost low.

CONTEXT_PROMPT = """<document>
{document}
</document>
Here is a chunk from the document:
<chunk>
{chunk}
</chunk>
Write 1-2 sentences that situate this chunk within the overall document to improve search retrieval.
Answer only with the context."""

14. Agentic RAG

Letting the LLM decide when and what to search, possibly several times. Expose retrieval as a tool (search_docs(query, filters)); the model searches, reads results, refines the query, and answers when it has enough (29, 32).

Use it for complex questions needing several lookups or comparisons; mixed sources (docs + DB + web).

User: "Compare our 2024 and 2025 refund policies for electronics."
Agent: search_docs("refund policy electronics 2024")   -> chunks A
       search_docs("refund policy electronics 2025")   -> chunks B
       answer comparing A and B with citations

Trade-offs: better on hard questions, but slower, costlier and less predictable than a fixed pipeline.

15. Conversational RAG (Follow-up Questions)

Handling questions that depend on earlier turns ("and what about shoes?"). Before retrieval, rewrite the latest question into a standalone query using the chat history.

Use it in every chat-style RAG app.

REWRITE_PROMPT = """Given the conversation and the follow-up question, rewrite the follow-up
as a standalone search query. Reply with the query only.

<conversation>
{history}
</conversation>
Follow-up: {question}"""

16. Structured Data: Text-to-SQL

Answering questions about tables by generating SQL instead of embedding rows. Give the LLM the schema; it writes a query; your code runs it with a read-only user; the LLM explains the result.

Use it for numbers, aggregates, filters over databases ("revenue by region last quarter"). Embeddings are bad at arithmetic over rows.

Safety: read-only DB user, allow-listed tables, row limits, statement timeouts, validate / parse the SQL before running. See 20 - SQL.

17. Evaluating RAG

Measuring retrieval and answer quality separately. Build a test set of questions with expected answers and the source chunks that contain them; score each stage.

Use it before launch, and after every change to chunking, models or prompts.

Stage Metric Question it answers
Retrieval Recall@k / hit rate Is the right chunk in the top k?
Retrieval MRR How high is the first right chunk?
Generation Faithfulness / groundedness Is every claim supported by the sources?
Generation Answer relevance / correctness Does it answer the question correctly and completely?
Generation Citation accuracy Do citations point to the right source?
End to end "Don't know" accuracy Does it refuse when the answer is not in the docs?
def recall_at_k(test_set: list[dict], k: int = 5) -> float:
    """Share of questions where an expected chunk id appears in the top k results."""
    hits = 0
    for case in test_set:
        ids = [c["id"] for c in retrieve_with_ids(case["question"], k=k)]
        hits += any(e in ids for e in case["expected_chunk_ids"])
    return hits / len(test_set)

Generation quality is often graded with an LLM-as-judge and a rubric. Tools: Ragas, DeepEval, promptfoo, Langfuse (35).

18. Keeping the Index Fresh

Updating the index when documents change. Stable chunk IDs, content hashes to detect changes, upsert changed chunks, delete removed ones.

Use it in any living knowledge base.

  • Store a hash of each document; re-process only changed files.
  • Delete all chunks of a document (by source metadata) before re-adding its new version.
  • Schedule re-indexing (cron / GitHub Actions / queue workers, 42, 44).
  • Record the embedding model version; re-embed everything when it changes.

19. Security and Permissions

Making sure users only get answers from documents they may see, and that documents cannot hijack the model. Filter by user permissions at retrieval; treat retrieved text as untrusted data.

Use it in any multi-user or internal-document RAG.

  • Store access metadata (tenant, team, role) with every chunk; filter on every query in code.
  • Retrieved documents can contain prompt-injection text ("ignore previous instructions..."); tag them as data and do not give the RAG model risky tools (38).
  • Remove secrets and personal data before indexing where possible.
  • Log questions and sources for auditing (respect privacy rules).

20. RAG Frameworks

Libraries that provide loaders, splitters, retrievers and pipelines. They wrap the steps above; you trade control for speed of setup.

Use it for many file types, complex pipelines, or quick prototypes. Plain code is often clearer for small apps.

Framework Strength
LlamaIndex Data loading, indexing and retrieval for RAG
LangChain / LangGraph Broad integrations, chains and agent graphs
Haystack Production search / RAG pipelines
Azure AI Search + Azure OpenAI "on your data" Managed RAG on Azure

More in 33 - Agent Frameworks.

21. Common Failures and Fixes

Symptom Likely cause Fix
"I couldn't find it" but the doc has it Retrieval miss Check recall; hybrid search; better chunking; rewrite queries
Answer uses the wrong document Similar docs / versions Metadata filters (date, product); reranking
Answer is partial Info spread across chunks Bigger chunks / parent-child; retrieve more; agentic multi-search
Hallucinated details Weak grounding instructions or poor chunks "Only use sources", allow "don't know", require citations, faithfulness evals
Tables / numbers garbled PDF extraction Layout-aware loaders (docling) or vision model on pages
Slow responses Too many / large chunks, slow reranker Fewer chunks, smaller reranker, cache, stream the answer
Costs high Huge prompts Fewer chunks, prompt caching of stable parts, smaller model for rewriting
Works in tests, fails for users Test questions unlike real ones Build eval set from real user questions
Follow-up questions fail No query rewriting Conversational rewrite step (section 15)

22. Try It

Short exercises to practise this guide. Try each task yourself first, then open the solution.

Use it right after reading the guide, or later as a quick self-test.

Exercise 1: Run the retrieval eval

In examples/, run the retrieval eval, then add a question phrased with different words ("When do I get my money back?") and run it again.

Solution
cd examples
uv run python -m docs_chatbot.evals

The paraphrase fails with the offline hashing embedder (it matches spelling, not meaning). Switching to SentenceTransformerEmbedder fixes it: that is exactly why real RAG uses semantic embeddings.

Exercise 2: System prompt

Write a RAG system prompt that requires citations and defines what to say when the answer is missing.

Solution

See SYSTEM_PROMPT in examples/docs_chatbot/answer.py: answer only from numbered sources, cite [n] after each claim, reply with an exact fallback sentence when the sources do not contain the answer, no outside knowledge.

Exercise 3: Debug a wrong answer

The document contains the answer but the bot says it cannot find it. What do you check, in order?

Solution
  1. Retrieval: is the right chunk in the top-k (print hits and scores)? 2. Chunking: is the answer split across chunks? 3. Wording: add hybrid search or query rewriting. 4. Thresholds: min_score too high? 5. Only then the prompt / model.