Skip to content

Semantic Cache

Listen to this page4:00
Read the transcript

1. The false-hit curve nobody reports

Host: Every dashboard for a semantic cache shows you the same headline number: hit rate. It’s the metric everyone brags about, because higher hit rate means faster responses and lower compute costs. But today we’re going to talk about the number nobody puts on that dashboard.

Guest: Right, the false-hit rate. That’s when the cache confidently serves you a wrong answer, and it comes back just as fast and just as polished as a right one. Downstream, there’s no way to tell the difference — it looks exactly like a correct response until someone acts on it.

Host: So walk me through what happens when you actually sweep the similarity threshold on a real corpus. I’d assume there’s some sweet spot where you get a healthy hit rate and precision stays high.

Guest: That’s the finding that should worry people: there isn’t one. At 0.90 you already get false hits with zero precision. At 0.90 and 0.85, hit rate and false-hit rate move together exactly. Loosen it further, past 0.80, and hit rate starts pulling ahead as precision actually rises — but it never gets safe, it just gets less bad. And it’s worse than overlapping distributions — they’re inverted. The EU refund question scores 0.875 against the US version, higher than ‘how do I reset my password’ scores against its own genuine paraphrase at 0.833. The pair that must never be conflated is literally the more similar one.

2. What’s real about it, and what a real embedding would fix

Host: So how much of that inversion is a real property of the data, and how much is just an artifact of the toy embedding you built for this test?

Guest: The mechanism is real — a short entity token like EU versus US gets swamped by all the shared framing around it, and that’s a genuine weak spot even for real dense embeddings, not just a lexical stand-in. The inversion itself is probably an artifact of that stand-in; a real model would likely separate those two groups better and soften the flip into overlap rather than a clean reversal. But take the toy embedding away and the engineering point still holds: tightening the threshold only trades a false hit for a missed one, it can’t erase the risk, and there’s no universal number you can borrow — the threshold only means something once you’ve measured it against your own corpus and your own model.

Host: So there’s no number you can just lift from someone else’s benchmark and paste into your config — you have to build the labeled near-duplicate set yourself and find your own curve.

3. The key you forgot, and who’s supposed to notice

Host: Okay, so we’ve got a threshold we can’t fully trust. Is there anything about this cache that’s actually a solid rule, not a judgment call?

Guest: Yes — the key itself, before you even touch similarity. A cached answer is only valid for one model, one prompt version, one tenant, and that has to be filtered before the vector search runs, not after. Filter afterward and you’ve already read another tenant’s answer and are just deciding what to do about it, which is the exact same gotcha vector databases have with cross-tenant leaks. Prompt version is the one everyone forgets, because prompts get edited weekly while models change rarely, so an edit that changes the right answer just keeps serving the old one until the TTL finally expires.

Host: Which means the real open questions aren’t technical at all — nobody’s put a price on what a wrong cached answer costs, so the threshold gets set by whoever’s staring at the hit-rate dashboard that day. Nothing downstream can tell a false hit from a real one apart from sampling cached against fresh, which almost nobody bothers to do. And if your eval suite runs through the cache, you’re not grading the model anymore, you’re grading the cache — and it’ll keep reporting yesterday’s score long after the model’s changed.

Not covered

The planner wanted these and found nothing in the source to support them:

  • A concrete dollar figure for what a wrong cached answer costs — the excerpts pose this as the open question a team must answer, but don’t supply one
  • A worked example of detecting entity-swap near-duplicates automatically, beyond naming it as ‘the fix’
  • A production case study of this cache deployed at scale alongside the Model Router lab’s cost tradeoffs

Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.

Module 4 names caching as a real lever on both cost and latency. This lab measures the other side of that lever: what the similarity threshold costs in wrong answers.

Source: labs/semantic-cache

Sweeping the threshold produces two curves. Hit rate is the one that gets reported. False-hit rate — a wrong cached answer, served fast, downstream indistinguishable from a right one — is the one that decides whether the cache is a saving or a correctness bug.

thresh hit rate false-hit precision true false miss
0.99 0.00% 0.00% 100.00% 0 0 9
0.95 0.00% 0.00% 100.00% 0 0 9
0.90 22.22% 22.22% 0.00% 0 2 7
0.85 44.44% 44.44% 0.00% 0 4 5
0.80 55.56% 44.44% 20.00% 1 4 4
0.60 88.89% 55.56% 37.50% 3 5 1

At every threshold that serves anything at all, precision is below 100%. Tightening until the wrong answers disappear removes every right one too. There is no safe operating point on this corpus.

The two distributions are not merely overlapping — they are inverted:

Similarity range
Genuine paraphrases, which should hit 0.45 – 0.83
Near-duplicates with different answers, which must not 0.60 – 0.91

“What is the refund policy for EU customers?” scores 0.875 against the US version. “How do I reset my password?” scores 0.833 against “How can I…”. The pair that must never be conflated is the more similar one.

This is the part worth being careful about, because the result is dramatic and the embedding is fake.

The mechanism is real. The token carrying the semantic difference — EU/US, 2025/2026, free/paid — is short, low-weight, and swamped by the long shared framing around it. Paraphrases differ in function words, which are a larger fraction of a short question. Entity swaps are a known hard case for dense embeddings too.

The inversion is probably not. A real dense model would separate these groups better than a lexical proxy, and the inversion would most likely soften into an overlap. The tests document this explicitly — test_the_distributions_are_inverted_not_merely_overlapping exists so a reader reproducing this with a real model knows which assertion is expected to weaken.

What survives the change of model is the engineering conclusion:

A tighter threshold is not the fix. The fix is detecting that two questions differ on an entity that changes the answer. Threshold tuning trades one error for the other; it cannot eliminate both, and on some corpora there is no usable point at all.

And the process point: a threshold cannot be inherited. It is a joint property of your corpus and your embedding model, and the only way to know yours is to build a labelled set of near-duplicates and measure both curves.

A cached answer is an answer from a particular model, under a particular prompt version, for a particular tenant. Tests pin that changing any of the three misses. Prompt version is the one people forget, because prompts get edited far more often than models change — and an edit that alters the answer keeps serving the old shape until the TTL runs out.

Scope is filtered before the similarity search. Filtering afterwards means the nearest neighbour was already read out of another tenant’s data, which is the Vector DB lookup’s first gotcha in cache form.

Terminal window
cd labs/semantic-cache
uv venv .venv && uv pip install --python .venv/bin/python -e '.[dev]'
./.venv/bin/python -m ruff check . && ./.venv/bin/python -m mypy src && ./.venv/bin/python -m pytest -q

15 tests, ruff clean, mypy --strict clean.

  • What does a wrong answer cost here? Until that is a number, the threshold is being chosen by whoever last looked at the hit-rate dashboard.
  • Who notices a false hit? Nothing downstream can distinguish one from a correct answer. The only detection is sampling cached responses against fresh ones, and almost nobody does it.
  • Cache invalidation on prompt edits. Prompts change weekly. If the prompt version is not in the key, every edit silently keeps serving the previous behaviour.
  • A cache changes your evaluation. An eval suite that runs through the cache is measuring the cache, not the model — and will keep reporting yesterday’s scores after a model change.
  • Architecture: Semantic Response Caching — the design review: the scope key, why the filter belongs inside the query, and which parts of the threshold finding above survive a real embedding model.
  • Module 4: AI Infrastructure — caching as a cost lever, and the cost-latency-quality triangle.
  • Model Router — the other lever on the same cost problem, with its own uncomfortable measurement.
  • Vector DB — similarity search, metadata filtering, and why the filter belongs in the query.