Module 8: RAG
Read the transcript
1. Why bolt retrieval onto a model at all
Host: Welcome to Module 8, where we’re tackling RAG — retrieval-augmented generation. Before we get into chunking strategies and re-ranking and all the architectural decisions, I want to start with the basic question: why do we even need to bolt retrieval onto a language model in the first place? Why isn’t the model enough on its own?
Guest: Because a model’s knowledge is frozen the moment training ends, and it never had access to your private documents to begin with — your internal wikis, your customer data, whatever’s specific to you. You could retrain or fine-tune every time that data changes, but that’s slow and expensive, and your data probably changes daily. RAG sidesteps both problems: instead of baking knowledge into the model’s weights, you retrieve the relevant pieces of your own data at the moment of the request and hand them to the model as context alongside the query.
2. Two pipelines, one hard ceiling
Host: So RAG isn’t one pipeline, it’s two. Walk me through the split, because I think people picture it as a single flow from question to answer.
Guest: Right, there’s ingestion, which happens ahead of time — you chunk your documents, embed them, and write them into a vector index. Then there’s the query pipeline, which runs per request: the incoming question gets embedded with that same embedding model, compared against the index, and whatever matches best becomes the context you hand to the model. Almost every RAG failure traces back to these two drifting out of sync — different chunking assumptions, a mismatched embedding model, a stale index that doesn’t reflect what changed.
Host: And if that context is wrong or incomplete, is there any amount of clever prompting that saves you?
Guest: None. If the retrieved context doesn’t contain the answer, the model can’t produce it — it’s not being stubborn, the information simply isn’t there. That’s why this module measures retrieval and generation as two separate systems instead of one blended score, because you can pour all your effort into generation and get nothing if retrieval was the actual bottleneck.
3. Chunking: the decision nobody thinks is a decision
Host: So before any of that retrieval machinery even runs, someone has to decide how to cut the documents into pieces in the first place. That sounds like the most boring step in the whole pipeline — why does it actually matter?
Guest: Because that cut determines what the model is even allowed to see later. The naive approach is fixed-size chunking — just slice at a fixed length — and it’s simple, but it doesn’t know or care what’s in the document. It’ll happily cut a table in half, or split a procedure’s numbered steps across two chunks, and now neither half retrieves as useful because the piece that made it complete is sitting in the neighboring chunk. Semantic chunking, splitting on paragraph or section boundaries, respects that structure but gives you variable, less predictable chunk sizes to manage.
Host: And I assume size itself is a whole separate knob — you can’t just pick ‘small’ or ‘large’ and move on.
Guest: Right, it’s a real trade-off in both directions. Too small and you lose the surrounding context that made the chunk meaningful on its own; too large and you dilute relevance — a chunk that matches on one sentence but is mostly noise, while also burning more of your context-window budget per retrieved item. Overlapping chunks help hedge against a boundary slicing through something important, but that redundancy costs you extra storage and embedding compute, so even that fix isn’t free.
4. Finding the right chunks: embeddings, ANN, hybrid search, re-ranking
Host: Okay, so once the chunks exist, how does the system actually find the right ones at query time? I assume it’s not scanning every chunk in the database for every question.
Guest: Right, both the chunks and the incoming query get turned into dense vectors by an embedding model, and you’re looking for the closest vectors by something like cosine similarity. But exact nearest-neighbor search over millions of vectors is too slow to be practical, so production systems use approximate nearest-neighbor indexes — HNSW is the one you’ll see most. You’re deliberately trading a little bit of recall for a huge speed win, and that’s a real knob you tune, not something to leave on default.
Host: But embeddings capture meaning, not exact strings — so what happens when someone types an actual product SKU or an error code and the semantically ‘closest’ chunk isn’t the one with that literal string in it?
Guest: That’s exactly the gap — dense retrieval is great at conceptual similarity but can miss exact keyword or entity matches, while sparse keyword search like BM25 is the opposite: it nails exact matches but is blind to paraphrasing. So hybrid search runs both and fuses the results, usually with something like reciprocal rank fusion, since the two approaches fail in different ways. And on top of that, a common pattern is to over-fetch a larger candidate set cheaply with ANN, then run a more expensive cross-encoder re-ranker over just that smaller set to pick the final top few that actually reach the model — trading extra latency for meaningfully better precision.
5. Freshness is a consistency problem wearing a search-index costume
Host: So once the top-k chunks are actually landing on the model’s desk, there’s still the question of whether they’re even up to date. This feels like it’s the same staleness question from the distributed systems module, just wearing a different hat.
Guest: It’s exactly that question, just pointed at a search index instead of a database. Module 2’s whole argument was that consistency gets chosen per invariant, not globally — a payment ledger needs strong consistency, a search index can usually tolerate being eventually consistent. Here the invariant is ‘how stale can retrieved content be before it misleads someone,’ and that answer is wildly different for a legal document repository than for a general knowledge base.
Host: And the support-docs case makes that concrete — reindex-on-change gets you near-real-time freshness but adds moving parts, versus a periodic full reindex that’s simple to operate but serves anything changed since the last run as confidently stale.
Guest: Right, and either choice is fine if you actually made it. What’s not fine is drifting into whichever was easiest to build, because then you get a support bot confidently quoting yesterday’s pricing page. From outside that looks like a generation failure — the model sounds so sure of itself — but it’s purely a retrieval-staleness failure, and no amount of prompt tuning fixes an index that’s quietly out of date.
6. Access control belongs inside the query, not after it
Host: Okay, so staleness is one failure mode, but there’s a scarier one lurking in the same architecture: what stops a retrieval system from just handing one tenant’s documents to another tenant’s user?
Guest: Nothing stops it, unless you build the filter into the search itself rather than bolting it on after. If your pipeline fetches the fifty best-matching chunks across the whole corpus and only then checks who’s allowed to see what, you’ve already lost — those unauthorized chunks influenced the ranking, and in a naive setup they can end up in the model’s context before anyone checks permissions.
Host: So the fix isn’t a permissions check somewhere in the application layer that a developer has to remember to call. It’s structural.
Guest: Exactly — in the implementation we walk through, tenant_id is a required parameter on the vector index’s search method itself, not an optional filter tacked on afterward. That means forgetting to scope a query by tenant isn’t a bug someone has to catch in code review, it’s a type error that fails before the code even runs — the query can’t compile without saying whose data it’s allowed to touch.
7. When retrieval quietly breaks: confident wrong answers and silent embedding drift
Host: So let’s talk about the failures that don’t announce themselves. What happens when the retrieved context just doesn’t have the answer in it?
Guest: The model doesn’t say ‘I don’t know’ — it generates something fluent and confident anyway, because that’s what language models do with whatever context they’re given. There’s also a silent one on the ingestion side — if you re-embed with a new model version but don’t fully rebuild the index, you’re comparing new-model query vectors against old-model chunk vectors, and similarity scores become meaningless without a single error being thrown.
8. Running RAG at scale: latency budgets, tuning knobs, and growing indexes
Host: So let’s say the pipeline is correct and the access control is solid — now it just has to run fast at scale. Where does retrieval latency actually sit in the request timeline?
Guest: It’s part of your time-to-first-token budget, full stop — the same layered latency thinking from the networking module, just applied to a retrieval hop instead of a network hop. Your ANN index has tuning knobs, HNSW being the common one, that trade recall for speed, and leaving them at default is a choice you’re making without realizing it. And every query also pays for embedding the query itself, which is a real per-request cost in latency and, if it’s a hosted API, in money — track it like you’d track generation cost.
Host: And once the index itself gets big, does that latency problem get worse on its own?
Guest: Eventually yes — past a certain size you have to shard the vector index across nodes, and now ‘search the index’ is a distributed query with its own consistency headaches. What trips people up is that ingestion and query don’t scale together — ingestion is bursty batch work tied to document volume, query scaling is tied to request volume, so you need separate capacity planning for each. And if you’re multi-tenant, you choose between separate indices per tenant, clean isolation but heavy operational overhead, or one shared index with metadata filtering, which scales better operationally but means every single query has to get that filter exactly right or isolation quietly fails.
9. Measuring the right thing: retrieval and generation as separate scores
Host: So once the thing is actually running at scale, how do you know if it’s any good? I feel like teams just eyeball the final answers and call it a day.
Guest: That’s exactly the trap. Retrieval is a distinct system from generation, and it needs its own scorecard. You build a labeled set of query-to-relevant-chunk pairs and measure recall and precision at k, completely independent of what the model does with those chunks once it has them. Generation gets scored separately, using the same evaluation harness pattern from the infrastructure module, against the final answer quality.
Host: And if you skip that separation, you just… guess at the fix?
Guest: Right, and you usually guess wrong. Poor end-to-end quality with strong retrieval recall means your generation or prompt construction is broken — the right chunks are there and the model’s botching them. The same poor quality with weak retrieval recall is a totally different problem, and no amount of prompt tuning fixes it. Without measuring the two separately, teams routinely burn weeks rewriting prompts when retrieval was the actual bottleneck the whole time.
10. Building it yourself: the hybrid retrieval lab and the tenant-isolation test
Host: So if someone wants to stop nodding along and actually build this, where do they start? Is there something that puts chunking, hybrid search, re-ranking and evaluation all in one place?
Guest: That’s exactly the hybrid retrieval and evaluation lab — structure-aware chunking, BM25 and vector search fused by rank rather than raw score, a re-ranking stage, groundedness checking, and a harness that spits out precision at k, recall at k, and MRR so you can watch the numbers we’ve been talking about all episode move. It’s labelled production-shaped because the pipeline is real but the embedding function is a deterministic hashing stand-in, no real semantic similarity, and it’s upfront about that. And then take it further yourself: implement the VectorIndex protocol against a real local vector store, seed it with two different tenant IDs, and write a test that queries as tenant A using content deliberately close to tenant B’s documents, then assert zero tenant B chunks ever leak through no matter how high the similarity score. If that test passes, you’ve proven the point of this module.
Not covered
The planner wanted these and found nothing in the source to support them:
- Comparative benchmarking of specific vector database products
- Fine-tuning or training a custom embedding model for a domain corpus
- Multi-modal (image/audio) retrieval in a RAG pipeline
- Cost comparison of hosted vector DB services versus self-hosted ANN indexes
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Executive Summary
Section titled “Executive Summary”Retrieval-augmented generation exists because a model’s training data is fixed at a point in time and doesn’t include your private documents, and retraining or fine-tuning it every time that data changes is slow and expensive. RAG sidesteps both problems: retrieve the relevant pieces of your own data at request time, and hand them to the model as context alongside the query. The engineering challenge isn’t the idea — it’s keeping two independent pipelines (index-building and query-time retrieval) correctly in sync, and treating retrieval quality as the hard ceiling on answer quality.
Mental Model
Section titled “Mental Model”RAG is two pipelines, not one, and almost every RAG failure traces back to them drifting out of sync with each other:
- Ingestion runs ahead of time: documents are chunked, embedded, and written into a vector index.
- Query runs per request: the incoming query is embedded with the same embedding model, compared against the index, and the top matches become context for generation.
Retrieval quality is a ceiling, not a knob
A better model cannot compensate for bad retrieval — if the context handed to the model doesn’t contain the answer, no amount of prompting or model capability recovers it. Every effort spent improving generation quality is wasted if retrieval is the actual bottleneck, which is why this module treats retrieval and generation as two separately measured systems, not one blended “RAG quality” number.
Architecture
Section titled “Architecture”flowchart TB
subgraph Ingestion["Ingestion pipeline (runs ahead of time)"]
Docs[Source documents] --> Chunk[Chunk]
Chunk --> EmbedIngest[Embed]
EmbedIngest --> Index[(Vector index)]
end
subgraph Query["Query pipeline (runs per request)"]
Q[User query] --> EmbedQuery[Embed]
EmbedQuery --> ANN["ANN search + access-control filter"]
Index -.-> ANN
ANN --> Rerank["Re-rank (optional second stage)"]
Rerank --> Prompt["Construct prompt: instructions + retrieved chunks"]
Prompt --> Model[Model]
Model --> Response[Response]
endThe access-control filter sitting inside the query pipeline’s retrieval step — not bolted on afterward — is the architectural decision this module returns to repeatedly: a retrieval system that fetches first and filters for permissions second has already leaked the existence and content of unauthorized documents into the ranking and, in a naive implementation, into the model’s context.
Deep Dive
Section titled “Deep Dive”Chunking. Documents are split into chunks small enough to embed and retrieve individually. Fixed-size chunking is simple but indifferent to document structure — it can split a table or a procedure’s steps across two chunks, destroying the context that made either half useful alone. Semantic chunking (splitting on paragraph, section, or heading boundaries) respects document structure at the cost of variable chunk sizes. Chunk size itself is a tuning trade-off: too small loses surrounding context a chunk needed to make sense; too large dilutes relevance (a chunk matching on one sentence but mostly irrelevant) and wastes context-window budget. Overlapping chunks (a chunk sharing some content with its neighbor) reduce the chance a boundary splits something important, at the cost of redundant storage and embedding compute.
Embeddings and similarity search. A chunk and a query are both mapped to dense vectors by an embedding model; retrieval finds the chunks whose vectors are closest to the query’s (typically by cosine similarity). Exact nearest-neighbor search doesn’t scale past a modest index size, so production systems use approximate nearest-neighbor (ANN) index structures — HNSW is the dominant one in practice — trading a small amount of recall for a large speed improvement at scale.
Hybrid retrieval. Dense (embedding) retrieval is good at conceptual/semantic similarity but can miss exact keyword or entity matches a user actually typed (a product SKU, an error code). Sparse retrieval (BM25-style keyword matching) is the reverse: excellent at exact matches, blind to paraphrasing. Combining both — hybrid search, typically merged via reciprocal rank fusion — usually outperforms either alone, because they fail in different, complementary ways.
Re-ranking. A common two-stage pattern: cheap ANN retrieval over-fetches a larger candidate set (say, top 50) than what’s actually needed, then a more expensive cross-encoder re-ranker scores that smaller candidate set and reorders it down to the final top-k (say, top 5) that actually goes to the model. This trades extra latency and compute for meaningfully better precision than single-stage ANN retrieval alone provides.
Freshness is a consistency problem in disguise. How stale can the index be relative to the source data before it matters? This is exactly Module 2’s per-invariant consistency question, applied to a search index instead of a database: a legal document repository probably needs near-real-time reindexing on change, while a general knowledge base might tolerate a daily batch reindex just fine.
Research Note
The paper that named and formalized RAG — worth reading for the original framing of retrieval as a way to give a model access to knowledge without baking it into model weights.
Source: Lewis et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks"
Implementation
Section titled “Implementation”The shape of a retrieval query with access control enforced inside the search, not after it — the concrete implementation of this module’s Architecture decision:
Tenant-scoped retrieval with a two-stage rerank
from __future__ import annotations
from dataclasses import dataclassfrom typing import Protocol
@dataclassclass Chunk: id: str text: str tenant_id: str source_document: str
class VectorIndex(Protocol): async def search( self, query_vector: list[float], top_k: int, tenant_id: str ) -> list[Chunk]: """Implementations MUST filter by tenant_id inside the index query itself — never fetch candidates first and filter afterward in application code.""" ...
class Reranker(Protocol): async def rerank(self, query: str, candidates: list[Chunk]) -> list[Chunk]: ...
async def retrieve( query: str, tenant_id: str, embed: "Callable[[str], list[float]]", index: VectorIndex, reranker: Reranker, candidate_k: int = 50, final_k: int = 5,) -> list[Chunk]: query_vector = embed(query) candidates = await index.search(query_vector, top_k=candidate_k, tenant_id=tenant_id) reranked = await reranker.rerank(query, candidates) return reranked[:final_k]The VectorIndex.search signature takes tenant_id as a required parameter, not an optional
filter applied afterward — this is a deliberate API design choice, not a style preference. It makes
“forgot to filter by tenant” a signature mismatch a type checker catches, instead of a runtime data
leak a code reviewer has to remember to look for on every call site.
Production Example
Section titled “Production Example”A support-docs RAG system’s source documents get updated throughout the day. Two reindexing strategies, with different freshness/cost trade-offs: reindex-on-change (a webhook or change-data- capture event triggers re-chunking and re-embedding of just the changed document, near-real-time freshness, more moving parts) versus periodic full reindex (rebuild the whole index on a schedule, simpler to operate, but any document changed since the last run is served stale until the next cycle). A system that silently picked “whatever was easiest to build” without an explicit freshness requirement is how a support bot ends up confidently quoting a pricing page that changed yesterday — not a generation failure, a retrieval-staleness failure that looks identical to one from the outside.
Failure Modes
Section titled “Failure Modes”Irrelevant retrieval producing confident wrong answers
If the retrieved context doesn’t actually contain the answer, the model often generates a fluent, confident-sounding response anyway — retrieval quality is genuinely the ceiling on answer quality, per this module’s Mental Model section, and no amount of prompt engineering recovers a retrieval miss.
Chunk boundaries splitting relevant content
A procedure’s numbered steps or a table’s rows split across two chunks can each retrieve as individually low-relevance, when the whole would have scored highly and answered the query completely. This is exactly why chunking strategy (Deep Dive) is a real design decision, not an arbitrary implementation detail.
Embedding model version mismatch
If the ingestion pipeline re-embeds with a new embedding model version but the existing index isn’t fully rebuilt, queries end up comparing new-model query vectors against old-model chunk vectors — similarity scores become meaningless. Any embedding model change requires a full reindex, not an incremental one; this is easy to miss because nothing errors, retrieval quality just quietly degrades.
Missing per-tenant access control at retrieval time
If a vector index doesn’t enforce tenant or user permissions inside the search itself — filtering candidates by authorization after fetching them, or not at all — a multi-tenant RAG system can retrieve, and hand to the model, content the requesting user was never authorized to see. This is a correctness and security failure simultaneously, and it survives even a well-guarded application layer if the index query itself doesn’t enforce it. See this module’s Security section.
Stale index after source data changes
Without a defined freshness requirement and a reindexing strategy to match it (this module’s Production Example), an index silently drifts from the source of truth — the same distributed- systems staleness problem Module 2 covers generally, showing up here as a search index instead of a replicated database.
Trade-offs
Section titled “Trade-offs”Chunk size: small vs. large
Small chunks retrieve precisely but can lose surrounding context a passage needed to make sense standalone. Large chunks preserve more context but dilute relevance scoring and cost more tokens per retrieved item — fewer of them fit in a given context budget.
Pure dense retrieval vs. hybrid (dense + sparse)
Dense-only retrieval is simpler to operate (one index, one similarity metric) but misses exact keyword and entity matches. Hybrid retrieval catches both semantic and exact matches at the cost of maintaining two retrieval mechanisms and a fusion strategy between their results.
Single-stage retrieval vs. retrieve-then-rerank
Single-stage ANN retrieval is faster and simpler. A second-stage cross-encoder re-rank meaningfully improves precision on the final result set, at the cost of the added latency and compute of scoring every candidate in the larger first-stage set.
Security
Section titled “Security”- Enforce access control inside the retrieval query itself, per this module’s Implementation section — never as a post-filter on results already fetched, and never left entirely to the application layer to remember on every call site.
- Prompt injection via retrieved content — a malicious or compromised document in the corpus can contain instructions the model follows once retrieved into context, the same failure mode Module 4 and Module 6 cover for other untrusted-content sources, arriving here through the corpus itself.
- Embeddings can leak information about source content even when the index doesn’t store raw text — a real, if less common, concern for highly sensitive corpora where even vector proximity could reveal something about restricted content.
Performance
Section titled “Performance”- Retrieval latency is part of the overall time-to-first-token budget, not a separate concern — see Module 3’s treatment of layered latency budgets, applied here to the retrieval step specifically.
- ANN index parameters trade recall for speed — HNSW’s tuning knobs directly control how much approximate search sacrifices exact-nearest-neighbor recall for query speed; this is a real, measurable trade-off to tune against your actual latency and quality requirements, not a default to leave untouched.
- Query-time embedding computation is a per-request cost, in both latency and (if using a hosted embedding API) money — track it the same way Module 4’s cost-tracking pattern tracks generation cost.
Scaling
Section titled “Scaling”- Index size growth eventually requires sharding the vector index across multiple nodes, turning “search the index” into a distributed query with its own latency and consistency characteristics.
- Ingestion and query workloads scale independently and have different resource profiles — ingestion is bursty and throughput-oriented (batch-embed a large document set), query is latency-sensitive and needs to scale with request volume, not document volume.
- Multi-tenant index isolation is itself a scaling decision: separate indices per tenant give the cleanest isolation guarantee but multiply operational overhead per tenant, while a shared index with metadata filtering (this module’s Implementation section) scales operationally better but puts all the isolation weight on that filter being correct on every single query.
Interview Questions
Section titled “Interview Questions”How would you design a RAG system for a multi-tenant SaaS product?
Anchor on access control enforced inside the retrieval query (never a post-filter), an explicit chunking strategy matched to the document types involved, hybrid retrieval if exact-match queries matter, and a defined freshness requirement driving the reindexing strategy — not just “embed the docs and search them.”
How do you keep a vector index in sync with a changing data source?
Either reindex-on-change (event-driven, near-real-time, more infrastructure) or periodic full reindex (simpler, accepts bounded staleness) — the right choice depends on an explicit freshness requirement, per this module’s Production Example, not on whichever is easiest to stand up first.
How do you evaluate retrieval quality separately from generation quality?
Retrieval evaluation uses recall/precision @k against a labeled set of query-to-relevant-chunk pairs, independent of what the model does with those chunks. Generation evaluation, per Module 4’s evaluation harness, scores the final answer. Conflating the two hides which one is actually the bottleneck when quality is poor.
Why this separation matters in practice
A system with poor end-to-end answer quality but strong retrieval recall has a generation or prompt-construction problem; the same poor end-to-end quality with weak retrieval recall has a completely different fix. Without separate measurement, teams commonly spend weeks tuning prompts when the actual problem was retrieval the whole time.
How do you prevent one tenant's documents from leaking into another tenant's retrieval results?
Enforce the tenant filter as a required, non-optional parameter to the index search itself (this module’s Implementation section) so it’s structurally impossible to call retrieval without it — not a convention documented in a wiki that a future call site can forget.
Hands-on Lab
Section titled “Hands-on Lab”Hybrid Retrieval and Evaluation implements this module’s
retrieval pipeline end to end: structure-aware chunking, BM25 and vector retrieval fused by rank
rather than score, a reranking stage, groundedness checking, and an evaluation harness reporting
precision@k, recall@k, and MRR. It is labelled production-shaped — the pipeline and metrics are
real, but the embedding function is a deterministic hashing stand-in with no semantic similarity,
which the lab is explicit about.
Prove cross-tenant isolation in a retrieval query
Implement the VectorIndex protocol from this module’s Implementation section against any
embeddable local vector store, seed it with documents from two different tenant_id values, and
write a test that queries as tenant A with content deliberately similar to tenant B’s documents —
asserting that no tenant B chunks ever appear in tenant A’s results, regardless of similarity
score.
References
Section titled “References”- Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”
- Anthropic, “Introducing Contextual Retrieval” — a concrete technique for improving chunk retrieval quality by adding chunk-level context before embedding.
- Module 2: Distributed Systems — the consistency and staleness concepts this module applies to index freshness.
- Module 4: AI Infrastructure — the evaluation harness pattern generation quality is measured against, separately from retrieval.
Revision History
Section titled “Revision History”| Version | Date | Change |
|---|---|---|
| 1.0.0 | 2026-08-05 | Initial publication. |