Semantic Cache
Read the transcript
1. The false-hit curve nobody reports
Host: Every dashboard for a semantic cache shows you the same headline number: hit rate. It’s the metric everyone brags about, because higher hit rate means faster responses and lower compute costs. But today we’re going to talk about the number nobody puts on that dashboard.
Guest: Right, the false-hit rate. That’s when the cache confidently serves you a wrong answer, and it comes back just as fast and just as polished as a right one. Downstream, there’s no way to tell the difference — it looks exactly like a correct response until someone acts on it.
Host: So walk me through what happens when you actually sweep the similarity threshold on a real corpus. I’d assume there’s some sweet spot where you get a healthy hit rate and precision stays high.
Guest: That’s the finding that should worry people: there isn’t one. At 0.90 you already get false hits with zero precision. At 0.90 and 0.85, hit rate and false-hit rate move together exactly. Loosen it further, past 0.80, and hit rate starts pulling ahead as precision actually rises — but it never gets safe, it just gets less bad. And it’s worse than overlapping distributions — they’re inverted. The EU refund question scores 0.875 against the US version, higher than ‘how do I reset my password’ scores against its own genuine paraphrase at 0.833. The pair that must never be conflated is literally the more similar one.
2. What’s real about it, and what a real embedding would fix
Host: So how much of that inversion is a real property of the data, and how much is just an artifact of the toy embedding you built for this test?
Guest: The mechanism is real — a short entity token like EU versus US gets swamped by all the shared framing around it, and that’s a genuine weak spot even for real dense embeddings, not just a lexical stand-in. The inversion itself is probably an artifact of that stand-in; a real model would likely separate those two groups better and soften the flip into overlap rather than a clean reversal. But take the toy embedding away and the engineering point still holds: tightening the threshold only trades a false hit for a missed one, it can’t erase the risk, and there’s no universal number you can borrow — the threshold only means something once you’ve measured it against your own corpus and your own model.
Host: So there’s no number you can just lift from someone else’s benchmark and paste into your config — you have to build the labeled near-duplicate set yourself and find your own curve.
3. The key you forgot, and who’s supposed to notice
Host: Okay, so we’ve got a threshold we can’t fully trust. Is there anything about this cache that’s actually a solid rule, not a judgment call?
Guest: Yes — the key itself, before you even touch similarity. A cached answer is only valid for one model, one prompt version, one tenant, and that has to be filtered before the vector search runs, not after. Filter afterward and you’ve already read another tenant’s answer and are just deciding what to do about it, which is the exact same gotcha vector databases have with cross-tenant leaks. Prompt version is the one everyone forgets, because prompts get edited weekly while models change rarely, so an edit that changes the right answer just keeps serving the old one until the TTL finally expires.
Host: Which means the real open questions aren’t technical at all — nobody’s put a price on what a wrong cached answer costs, so the threshold gets set by whoever’s staring at the hit-rate dashboard that day. Nothing downstream can tell a false hit from a real one apart from sampling cached against fresh, which almost nobody bothers to do. And if your eval suite runs through the cache, you’re not grading the model anymore, you’re grading the cache — and it’ll keep reporting yesterday’s score long after the model’s changed.
Not covered
The planner wanted these and found nothing in the source to support them:
- A concrete dollar figure for what a wrong cached answer costs — the excerpts pose this as the open question a team must answer, but don’t supply one
- A worked example of detecting entity-swap near-duplicates automatically, beyond naming it as ‘the fix’
- A production case study of this cache deployed at scale alongside the Model Router lab’s cost tradeoffs
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Module 4 names caching as a real lever on both cost and latency. This lab measures the other side of that lever: what the similarity threshold costs in wrong answers.
Source: labs/semantic-cache
The lookup path
Section titled “The lookup path”flowchart TD
Q["Incoming question"] --> Scope{"Exact key hit?<br/>tenant + model + prompt_version + text"}
Scope -->|Yes, unexpired| Exact["EXACT hit<br/>always correct"]
Scope -->|No| Filter["Filter candidates by scope<br/>BEFORE similarity, never after"]
Filter --> Near["Nearest neighbour<br/>cosine similarity"]
Near --> Cut{"score >= threshold?"}
Cut -->|No| Miss["MISS<br/>call the model<br/>similarity still reported, so<br/>the threshold can be tuned"]
Cut -->|Yes| Risk{"Is it the same question?"}
Risk -->|"Genuine paraphrase<br/>0.45 - 0.83 here"| True["TRUE hit<br/>cost saved"]
Risk -->|"Entity swap<br/>EU/US · 2025/2026 · free/paid<br/>0.60 - 0.91 here"| False["FALSE hit<br/>wrong answer, served fast<br/>indistinguishable downstream"]
True --> Note["The ranges overlap.<br/>Tightening to exclude the right column<br/>also excludes the left one."]
False --> NoteThe finding
Section titled “The finding”Sweeping the threshold produces two curves. Hit rate is the one that gets reported. False-hit rate — a wrong cached answer, served fast, downstream indistinguishable from a right one — is the one that decides whether the cache is a saving or a correctness bug.
thresh hit rate false-hit precision true false miss 0.99 0.00% 0.00% 100.00% 0 0 9 0.95 0.00% 0.00% 100.00% 0 0 9 0.90 22.22% 22.22% 0.00% 0 2 7 0.85 44.44% 44.44% 0.00% 0 4 5 0.80 55.56% 44.44% 20.00% 1 4 4 0.60 88.89% 55.56% 37.50% 3 5 1At every threshold that serves anything at all, precision is below 100%. Tightening until the wrong answers disappear removes every right one too. There is no safe operating point on this corpus.
The two distributions are not merely overlapping — they are inverted:
| Similarity range | |
|---|---|
| Genuine paraphrases, which should hit | 0.45 – 0.83 |
| Near-duplicates with different answers, which must not | 0.60 – 0.91 |
“What is the refund policy for EU customers?” scores 0.875 against the US version. “How do I reset my password?” scores 0.833 against “How can I…”. The pair that must never be conflated is the more similar one.
How much of that is the simulation
Section titled “How much of that is the simulation”This is the part worth being careful about, because the result is dramatic and the embedding is fake.
The mechanism is real. The token carrying the semantic difference — EU/US, 2025/2026,
free/paid — is short, low-weight, and swamped by the long shared framing around it. Paraphrases
differ in function words, which are a larger fraction of a short question. Entity swaps are a known
hard case for dense embeddings too.
The inversion is probably not. A real dense model would separate these groups better than a
lexical proxy, and the inversion would most likely soften into an overlap. The tests document this
explicitly — test_the_distributions_are_inverted_not_merely_overlapping exists so a reader
reproducing this with a real model knows which assertion is expected to weaken.
What survives the change of model is the engineering conclusion:
A tighter threshold is not the fix. The fix is detecting that two questions differ on an entity that changes the answer. Threshold tuning trades one error for the other; it cannot eliminate both, and on some corpora there is no usable point at all.
And the process point: a threshold cannot be inherited. It is a joint property of your corpus and your embedding model, and the only way to know yours is to build a labelled set of near-duplicates and measure both curves.
The other half: the scope key
Section titled “The other half: the scope key”A cached answer is an answer from a particular model, under a particular prompt version, for a particular tenant. Tests pin that changing any of the three misses. Prompt version is the one people forget, because prompts get edited far more often than models change — and an edit that alters the answer keeps serving the old shape until the TTL runs out.
Scope is filtered before the similarity search. Filtering afterwards means the nearest neighbour was already read out of another tenant’s data, which is the Vector DB lookup’s first gotcha in cache form.
Run it
Section titled “Run it”cd labs/semantic-cacheuv venv .venv && uv pip install --python .venv/bin/python -e '.[dev]'./.venv/bin/python -m ruff check . && ./.venv/bin/python -m mypy src && ./.venv/bin/python -m pytest -q15 tests, ruff clean, mypy --strict clean.
Principal-level discussion points
Section titled “Principal-level discussion points”- What does a wrong answer cost here? Until that is a number, the threshold is being chosen by whoever last looked at the hit-rate dashboard.
- Who notices a false hit? Nothing downstream can distinguish one from a correct answer. The only detection is sampling cached responses against fresh ones, and almost nobody does it.
- Cache invalidation on prompt edits. Prompts change weekly. If the prompt version is not in the key, every edit silently keeps serving the previous behaviour.
- A cache changes your evaluation. An eval suite that runs through the cache is measuring the cache, not the model — and will keep reporting yesterday’s scores after a model change.
Related
Section titled “Related”- Architecture: Semantic Response Caching — the design review: the scope key, why the filter belongs inside the query, and which parts of the threshold finding above survive a real embedding model.
- Module 4: AI Infrastructure — caching as a cost lever, and the cost-latency-quality triangle.
- Model Router — the other lever on the same cost problem, with its own uncomfortable measurement.
- Vector DB — similarity search, metadata filtering, and why the filter belongs in the query.