RAG
Read the transcript
1. Grounding as the whole point
Host: So everyone’s heard of RAG at this point, and I think most people would say the point is better retrieval — get the model the right documents and it gives a better answer. But I want to push back on that framing right out of the gate, because I don’t think that’s actually the value.
Guest: Right, the retrieval is just the mechanism. The actual payoff is traceability — when the answer’s wrong, you can point to the document that caused it instead of shrugging at a black box, and when the answer’s missing entirely, that’s a retrieval bug you can go fix, not some unexplainable hallucination. And I’ll say this up front too: RAG has no version. It’s a pattern, not a product, so when people say ‘RAG 2.0’ they’re not actually naming anything — there’s no spec that changed, just a diagram everyone learns that’s way simpler than the system teams actually end up running.
2. The diagram you learn vs. the system you run
Host: Okay so walk me through that diagram, because I’ve seen it in basically every intro tutorial. Sources go to loaders, loaders to splitters, splitters to embeddings, into a vector store, then a retriever pulls context and hands it to the LLM for an answer. Nine boxes, one line, done.
Guest: Right, and that diagram isn’t wrong, it’s just a sketch, not a blueprint. What teams actually run collapses into four stages, each with its own internal machinery: ingestion — parsing, chunking, metadata, embeddings, indexing; retrieval — query transformation, semantic or lexical or hybrid search, filtering, candidate retrieval, reranking; generation — context assembly, prompt, LLM, citations; and evaluation — retrieval quality, answer correctness, relevance, groundedness. That’s four systems, not one arrow.
Host: So the tutorial diagram is basically the trailer, and this is the whole movie. Where does that gap actually bite people in practice?
Guest: Almost everywhere — that’s the honest answer. Nobody’s chunking strategy survives contact with real documents unchanged, nobody’s retriever stays pure semantic search once they hit edge cases, and nobody skips evaluation once wrong answers start shipping. Most production RAG work is just discovering, one stage at a time, that the simple line was hiding a sub-system.
3. The vocabulary that actually does the work
Host: Let’s actually walk the vocabulary, because half of these terms get used loosely and it costs people later. Start at the beginning — loaders and splitters. Why do those two get treated as an afterthought when you say they break everything?
Guest: Because they’re boring until they’re not. A loader’s job is just getting bytes into documents, but that’s exactly where PDFs and tables quietly turn into garbage text nobody notices until retrieval quality tanks. Splitters compound it — fixed-size chunking on a document with headings will happily slice a table in half or separate a heading from its own content, and structure-aware splitting is the fix almost nobody does first. Then downstream, your embedding model has to be identical on ingestion and query sides, your vector store serves approximate nearest-neighbor over whatever survived that chunking, and a retriever sits on top as a swappable interface — none of which saves you if the chunk itself was already broken.
Host: So the fix-it-later stages get all the attention — hybrid search, reranking, citations — while the actual damage happened three steps earlier. Walk me through those, and then the three shapes, because I want to understand 2-step versus agentic versus hybrid as a real tradeoff, not just jargon.
Guest: Hybrid search fuses semantic and lexical — usually with reciprocal rank fusion — because BM25 still beats embeddings on identifiers and error codes, things semantic similarity is bad at. Reranking is a slower, stronger model re-scoring that shortlist — it’s literally where you buy precision, and metadata filtering has to happen at retrieval time, not after, or you leak across tenants. Citations are the payoff of all of it, the pointer from claim back to chunk, which is what makes groundedness checkable instead of assumed. And the three shapes are just how much of this you let happen deterministically: 2-step is query-retrieve-generate, one shot, predictable latency; agentic lets the model decide whether and what to retrieve, possibly repeatedly, trading predictability for coverage; hybrid RAG keeps deterministic retrieval where you need control and adds agentic or validation steps where you don’t.
4. Numbers that matter and the ways it quietly breaks
Host: Let’s get concrete for a second, because I think the abstractions can hide how unforgiving this stuff is. What’s the number that trips people up first?
Guest: Embedding dimension. It has to match the index exactly, and people treat swapping embedding models like a config change when it’s actually a full reindex. Worse, your ingestion pipeline and your query pipeline have to use the same model, or retrieval silently degrades to noise instead of loudly breaking — which is the worst failure shape, because nothing tells you it’s happening.
Host: And that silence seems to be a theme. What about the gotchas that actually leak data or lose answers outright?
Guest: Access control is the sharp one — if you filter after retrieval, you’ve already put another tenant’s document in the prompt, the damage is done. Then top-k is a hard recall ceiling, not a quality knob, so if the answer never made it into your k candidates, no reranker saves you. And citations that nobody checks are pure decoration — a model will point a claim at a chunk that doesn’t support it, and a footnote is not proof, it’s a guess with formatting.
5. Where the deeper engineering lives
Host: So if someone’s sitting there thinking my access control is probably fine, or wondering how you’d actually build the hybrid retrieval and groundedness checking you just described, where do they go from what we’ve talked about today?
Guest: Module 8 has the full treatment of chunking, hybrid retrieval, and re-ranking, including why access control has to sit at retrieval time, not after. If you want to see tenancy and freshness treated as first-class requirements rather than afterthoughts, there’s an enterprise RAG architecture built exactly that way. And there’s a hands-on lab called hybrid-retrieval with running, tested code for BM25, dense retrieval, reciprocal rank fusion, reranking, and a groundedness checker — plus a reference on what a vector DB actually does and doesn’t do for you. This conversation was the map; that’s the territory, and it’s worth actually walking.
Not covered
The planner wanted these and found nothing in the source to support them:
- A worked cost breakdown for running RAG at scale (embedding cost, storage cost, per-query cost)
- A side-by-side vendor comparison of vector databases or rerankers
- A concrete walkthrough of the disconnect-aware streaming implementation or networking latency budget
- A step-by-step chunking strategy tutorial with recommended sizes
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
At a Glance
Section titled “At a Glance”Ground a model’s answer in retrieved context instead of in whatever it memorized. The value is not the retrieval — it is that a wrong answer becomes traceable to a document, and a missing answer becomes a retrieval bug rather than a mystery.
RAG has no version. It is an architectural pattern, so “RAG 2.0” names nothing. This page uses
LangChain 1.4.3 as a concrete implementation baseline where an API surface is needed, and
keeps the conceptual material framework-neutral — every stage below is swappable and most teams end
up swapping several.
The pipeline everyone learns first:
Sources → Loaders → Splitters → Embeddings → Vector store → Retriever → Context → LLM → AnswerThe system you actually operate:
Ingestion parsing → chunking → metadata → embeddings → indexingRetrieval query transformation → semantic / lexical / hybrid search → filtering → candidate retrieval → rerankingGeneration context assembly → prompt → LLM → citationsEvaluation retrieval quality → answer correctness → relevance → groundednessThe gap between those two is where most production RAG work lives.
Key Concepts
Section titled “Key Concepts”| Concept | What it does |
|---|---|
| Loader | Gets bytes out of a source and into documents. Where PDFs and tables quietly break everything downstream |
| Splitter | Cuts documents into retrievable units. Structure-aware beats fixed-size on anything with headings |
| Embedding model | Maps text to a dense vector. Both pipelines — ingestion and query — must use the same one |
| Vector store | Holds vectors plus metadata and serves approximate nearest-neighbor search |
| Retriever | The interface over search: takes a query, returns candidates. Swappable independently of the store |
| Query transformation | Rewriting, expansion, or decomposition before search. The cheapest large recall win available |
| Lexical / BM25 search | Keyword matching. Still the better retriever for identifiers, error codes, and exact names |
| Hybrid search | Semantic and lexical candidates fused into one list — usually by reciprocal rank fusion |
| Metadata filtering | Restricting candidates by tenant, ACL, recency, or source. Enforced at retrieval, not after |
| Reranking | A stronger, slower model re-scoring a shortlist. Where precision is actually bought |
| Context assembly | Choosing what fits in the window, in what order, with what attribution |
| Citation | The pointer from a claim back to its chunk. What makes the answer auditable |
| 2-step RAG | Query → Retrieve → Generate. Deterministic retrieval, predictable latency, one shot |
| Agentic RAG | The agent decides whether and what to retrieve, possibly across several sources, possibly repeatedly |
| Hybrid RAG | Deterministic retrieval plus agentic or validation steps — control where you need it, flexibility where you do not |
| Groundedness | Whether the answer is supported by the retrieved context. Distinct from whether it is correct |
Numbers That Matter
Section titled “Numbers That Matter”| Quantity | Value | Why it matters |
|---|---|---|
| RAG’s version | None | It is a pattern. Any “RAG n.0” in a document is that document telling on itself |
| Implementation baseline here | LangChain 1.4.3 |
Released 2026-09-28; the framework-neutral material does not depend on it |
| Embedding dimension | Must match the index exactly | Changing embedding models is a full reindex, not a config change |
| Pipelines that must agree | 2 | Ingestion and query embed with the same model, or retrieval silently degrades to noise |
| Reranker input | A shortlist, not the corpus | Retrieve wide, rerank narrow. Reranking the top 5 cannot rescue a retriever that missed at 5 |
| Chunk size and overlap | No correct default | A tuning parameter against your own eval set. Anyone quoting a universal number has not measured |
| Evaluation dimensions | 4 | Retrieval quality, answer correctness, relevance, groundedness — measured separately, because they fail separately |
| Retrieval latency | Inside the time-to-first-token budget | Not a separate budget. See Module 3 |
| Agentic RAG round trips | Unbounded unless you bound them | Each retrieval decision is another model call and another chance to loop |
Common Gotchas
Section titled “Common Gotchas”- Retrieval quality and answer quality are different failures with different fixes. A wrong answer over correct context is a generation problem; a confident answer over irrelevant context is a retrieval problem. Measuring only end-to-end correctness tells you which is happening: nothing.
- Access control applied after retrieval has already leaked. Filtering the model’s output does not undo putting another tenant’s document in the prompt. Filter in the query.
- The two pipelines drift. Change the embedding model, the splitter, or the normalization on one side only, and retrieval degrades gradually rather than breaking — the worst failure shape there is.
- Semantic search is bad at exact identifiers. Error codes, SKUs, function names, and version strings are lexical matches. This is the case hybrid retrieval exists for, not a tuning problem.
- Top-k is a recall ceiling, not a quality knob. If the answer is not in the k candidates, no amount of reranking or prompting recovers it.
- Citations that are not checked are decoration. A model will happily attribute a claim to a chunk that does not support it. Groundedness has to be measured, not assumed from the presence of a footnote.
- Agentic RAG trades predictability for coverage. It can consult several sources and skip retrieval entirely when it should — and its latency and cost distributions get much wider. Choose 2-step when the latency budget is tight and the query shape is known.
- Chunking decisions are made once and paid for indefinitely. Re-chunking means re-embedding and re-indexing the whole corpus; it is the most expensive thing on this page to get wrong.
- An evaluation set built from questions you already answer well measures nothing. Build it from real failures, and grow it every time retrieval surprises you.
Where to Go Deeper
Section titled “Where to Go Deeper”- Module 8: RAG — chunking, hybrid retrieval, and re-ranking in full, including why access control belongs at retrieval time.
- Enterprise RAG Platform — the same system with tenancy, freshness, and access control as first-class requirements.
hybrid-retrieval— BM25, dense retrieval, reciprocal rank fusion, reranking, and a groundedness checker as running, tested code.- Vector DB — the storage layer underneath, and what it does and does not do for you.