Architecture: Semantic Response Caching
Read the transcript
1. Why this is the highest-leverage lever, and the catch
Host: So if you look at the cost and latency breakdown of pretty much any AI product, model calls dominate both. And a huge chunk of those calls are people asking something the system has already answered, just phrased a little differently. That feels like it should be an easy win.
Guest: It should be, but exact-match caching basically can’t touch it. ‘How do I reset my password’ and ‘how can I reset my password’ are the same question to a human and two completely different cache keys to a string comparison. So semantic caching comes in and matches on embedding similarity instead of exact text, and it becomes one of the highest-leverage cost levers you have.
Host: But you said it’s also the one that quietly changes what the system returns. That sounds like the catch.
Guest: It is the catch, and it’s a category difference, not just a tuning knob. An exact-match cache can only ever be slow or empty — worst case, you eat a model call. A semantic cache can be wrong. It can serve a confidently incorrect answer, fast, and nothing downstream can tell the difference: not the caller, not the model, not your logs, not your eval suite. So the real engineering question isn’t how to push the hit rate up, it’s how do you even know what that hit rate is costing you.
2. Two orderings that carry the whole design
Host: Okay, so if the whole risk is a confidently wrong answer served indistinguishably from a right one, where do you even start controlling that? You said there are two orderings that carry the design — what are they?
Guest: First one’s cheap and obvious once you say it out loud: always check exact-match before you ever touch semantic. It’s a hash lookup, zero false-hit risk, and it peels off a real chunk of traffic before you’ve asked any similarity question at all. The second ordering is the one that actually determines correctness, and it’s easy to get backwards — scope filter before similarity search, not after.
Host: Walk me through why the order there matters so much, because on the surface filtering feels like filtering either way.
Guest: If you filter after the search, the nearest neighbor was already read out of another tenant’s data before you discarded it — the read already happened, you’re just deciding what to do about it afterward. That’s the Vector DB lookup’s first gotcha showing up here as a cache: the ANN index and the tenant boundary have to be the same query, not two sequential steps, or you’ve built a leak that looks like a filter.
3. What makes a cache hit valid
Host: So given that gotcha, what actually has to be true for you to trust a cache hit at all? Let’s lay out the checklist, because it sounds like hit rate alone tells you nothing.
Guest: Start with an exact tier before you even touch semantic — a string match is free and certain, so let it absorb whatever share of traffic it can before you take on any risk. Then you need a scope key that makes a hit valid: model, prompt version, and tenant at minimum, because an answer is only an answer from a particular model, under a particular prompt, for a particular tenant. Beyond that you need false-hit rate measured continuously in production, not just hit rate; invalidation triggered by prompt edits and model changes instead of a TTL clock that has no idea what changed; and a bypass path so eval and debugging can always reach the real model instead of getting served a cached answer.
4. The threshold you can’t borrow, and the entity-swap trap
Host: So if I found a similarity threshold that works great in someone’s blog post, or even in your own reference implementation, can I just use that number?
Guest: No, and this trips people up constantly. The threshold is a joint property of your corpus and your embedding model — it’s telling you where similar-meaning text tends to cluster for that specific model on that specific kind of question, and neither of those transfers. Worse, when the threshold is wrong in the risky direction, there’s no symptom: nothing in the response distinguishes a false hit from a correct answer, so the only way to catch it is deliberately sampling cached answers against fresh ones, which almost nobody bothers to set up. And the case that breaks people’s intuition hardest is the entity swap — ‘pricing in the EU’ versus ‘pricing in the US,’ or ‘2025’ versus ‘2026’ — because that swapped token is short and carries almost no weight against a long shared sentence of framing, while genuine paraphrases differ mostly in function words, which are a bigger fraction of a short question. So the pair you must never conflate can score more similar than the pair that should legitimately hit.
Host: That’s a nasty inversion — the thing that actually matters least by length ends up mattering most by consequence. So how do you even start building a defense against something that leaves no trace once it’s wrong?
5. The lab’s inverted curves
Guest: You start by actually sweeping the threshold and printing both curves side by side, not just the one that makes the dashboard look good. On this corpus, at 0.90 you get a 22% hit rate, and every single one of those hits is false — zero precision. Push down to 0.60 and you’re hitting 89% of the time, but more than half of those hits are wrong answers served with total confidence.
Host: So there’s no dial you can turn where it’s just… fine. Every setting that serves anything is serving some garbage.
Guest: Right, precision never touches 100% at any threshold that returns results, and that’s before you get to the part that actually kept me up at night. The EU refund policy question scores 0.875 against the US version, and a genuine paraphrase — ‘how do I’ versus ‘how can I’ reset a password — scores 0.833. The pair that must never be conflated is literally more similar than the pair that’s supposed to be a safe hit.
6. What survives a real embedding model, and what doesn’t
Host: Okay, before we all panic and swear off semantic caching forever — how much of that inversion is a real property of language, and how much of it is just an artifact of whatever embedding model you happened to run the lab with?
Guest: The mechanism is real, the exact numbers aren’t. The lab used a lexical stand-in, so it’s basically counting shared words, and ‘EU’ versus ‘US’ is one short low-weight token drowning in a sea of identical framing, while ‘how do I’ versus ‘how can I’ changes function words that make up a huge fraction of a short question. A real dense model would likely separate those groups better, and the tests even name that expectation — there’s an assertion checking the distributions are inverted, not merely overlapping, specifically flagging which result should weaken if you rerun this with a real embedding.
Host: So if I swap in a proper model and the gap softens into overlap instead of a flip, does that mean the danger mostly goes away?
Guest: No, and that’s the conclusion that actually survives the swap — overlap is still fatal, it just looks like a harder tuning problem instead of an obvious inversion. Tightening the threshold trades false hits for false misses, it doesn’t eliminate either, and on some corpora there’s no usable point at all.
7. The scope key: model, prompt version, tenant
Host: So we’ve got the threshold pinned down as best it can be. But you keep mentioning a ‘scope key’ alongside it — what actually goes into that, and why isn’t the embedding similarity enough on its own?
Guest: Because a cached answer isn’t just an answer, it’s an answer from a specific model, under a specific prompt version, for a specific tenant — and the tests show that changing any one of those three invalidates the hit even if the embedding still lines up perfectly. Prompt version is the one everyone forgets, because prompts get edited constantly while models barely change, so a wording tweak that actually shifts the answer keeps quietly serving the old cached shape until the TTL finally expires.
Host: And where does that filtering actually happen relative to the similarity search itself?
Guest: Before, always before — scope has to gate the candidate set going into the search, not clean up after it. If you filter afterward, the nearest neighbor was already read out of another tenant’s vectors before you discarded it, which is exactly the vector store’s metadata-filtering and namespace story showing up in cache form: filtering is supposed to be evaluated as part of the search, and namespaces are the usual unit of tenant isolation, so the cache needs to lean on those same mechanics rather than bolt scope on as a post-hoc check.
8. How this fails in production
Host: So let’s walk through how this actually breaks in the wild, starting with the one everyone thinks they’ve solved: threshold tuning. Someone looks at a hit-rate dashboard, decides it’s too low, and turns the dial.
Guest: Right, and that dial only moves in trade, not in improvement. Loosen the threshold and hit rate goes up, but so does the false-hit rate, and only the hit rate shows up on the cost dashboard someone’s watching. Tightening doesn’t fix it either — it just swaps false hits for misses, because if the two populations overlap, no single cutoff separates them; you’re not solving a tuning problem, you’re discovering that similarity was never equivalence.
Host: And that’s separate from the tenant leakage failure we just covered — post-filtering looks fine in every test until the day it isn’t.
Guest: Exactly, it passes tests because the filter does remove the wrong-tenant result — until the day a result gets returned before the filter runs, or the filter itself gets dropped in a refactor. There was never structural isolation, just a check that happened to run every time until it didn’t. Then there’s the one that hits hardest because it’s so mundane: someone edits the prompt, the model doesn’t change, and the cache keeps serving the old answer for the full TTL — so you get a deploy log saying it shipped and a product behaving like it didn’t. Worse, the eval suite is scoring the cache’s answers, not the model’s, so it stays flat right through a real regression, and a TTL only tells you how old an answer is, never whether it’s still correct, so a fix just waits out an interval that was picked for reasons that have nothing to do with what changed.
9. Scaling the lookup path
Host: Let’s talk about what happens as this thing actually grows, because a lookup path that works at a thousand entries isn’t automatically the one that works at a million. Where’s the line, and what changes when you cross it?
Guest: Brute-force cosine scan is fine for longer than people expect — below a few thousand entries per scope, it’s cheaper to just compare against everything than to run an index. Past that you need ANN, but you’ve traded exact nearest-neighbour for approximate nearest-neighbour, so now the closest match the cache finds might not actually be the closest match that exists. Scope keys help enormously here because they partition the search space naturally — you’re scanning one tenant’s entries under one prompt version, not the whole corpus, so the same key that gives you isolation also keeps the candidate set small. But two things bite regardless of index choice: every single lookup pays for an embedding call, so you’ve traded a small model call for a large one on every request, and the embedding endpoint now inherits the traffic of your entire product. And when a popular entry expires, every concurrent caller piles onto the model at once unless you single-flight that miss — otherwise the cache amplifies exactly the load spike it was built to absorb, right when write volume is already peaking from a prompt change invalidating everything at once.
10. The security surface of a store of answers
Host: We’ve talked about this cache as a performance layer, but it’s also just a database sitting there with real answers in it — so what does the security surface look like once you frame it that way?
Guest: The headline risk is cross-tenant leakage, and the fix has to be structural, not a check you remember to add later — the scope filter belongs inside the query the vector store runs, not in application code you hope every call site respects. Beyond that, a cache is a store of answers, so it inherits the data classification of whatever’s in those responses, even though it usually gets a retention policy written for performance infrastructure instead of sensitive data; and the prompt version in your key is doing security work too, because if that version bump was the fix for a jailbreak or a data-exposure bug, stale cached entries under the old version are stale vulnerabilities. There’s also a subtler leak: timing itself is an oracle, since a fast response means a hit, which means somebody else asked something similar recently, and in a multi-tenant system that can reveal that another tenant exists or is active — usually fine to accept, but it’s a decision you want to make on purpose rather than discover in an audit.
11. Counting the real cost, and the trade-offs nobody wants to make explicit
Host: Let’s put a number on this. A hit swaps a full model call for an embedding call and a lookup — that’s two orders of magnitude cheaper, but you’re paying that embedding cost on every single request, hit or miss. So a cache with a low hit rate isn’t free, it’s just a small tax instead of a saving, right?
Guest: Exactly, and storage is basically free — it’s text, TTLs keep it bounded, nobody loses sleep over that line item. The cost nobody puts on an invoice is a false hit: a wrong answer costs whatever it costs your product, a support deflection that misinforms, a policy quoted for the wrong region.
Host: So without a dollar figure for a wrong answer, every threshold conversation is really just people arguing about a dashboard. Given that, what’s the actual defensible fallback — exact-only, and TTLs or versioned keys?
Guest: Exact-only is a legitimate final answer when correctness turns on entities the user supplies — you give up most of the savings but you can never be wrong, whereas semantic buys you the paraphrase traffic at the cost of a new correctness surface. For invalidation, run both: versioned keys for correctness since they cost nothing at read time, and a TTL as the backstop for the version bump somebody forgot. And the last honest question is whether you sample hits against a fresh call to get a real false-hit rate — it costs you a slice of the savings, but skip it and you won’t know your error rate until an incident tells you.
12. What good looks like: metrics, deployment checklist, and the interview questions that test for it
Host: So if I’m standing in front of a dashboard trying to decide whether this thing is healthy, what actually earns a place on it — not hit rate, since we’ve spent this whole episode complicating that one?
Guest: False-hit rate from sampling is the one metric here that doesn’t come for free from the cache’s own logs — you have to manufacture it by re-running a slice of hits against a fresh model call, and it’s the number the whole design turns on. Alongside it you want hit rate split by tier, the distribution of similarity scores on hits rather than the mean, miss reasons broken out by scope-versus-threshold-versus-expiry, hit rate crashing to zero right after a prompt bump, and the age distribution of what’s being served. Each one is a different way of catching the same failure before a user does — a mean of 0.94 hides the cluster sitting just above threshold, and a version bump that doesn’t zero out the hit rate means the version isn’t actually in the key.
Host: And if someone’s interviewing for a role where they’d own this, what question actually separates the person who’s built one from the person who’s read about one?
Guest: Ask what threshold eliminates both false hits and missed paraphrases, and the right answer is none, because the populations are inverted. And ask whether their eval suite runs through the cache — because if it does, it’s grading the cache, not the model, and that’s the same blind spot as the false-hit question wearing a different hat.
Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.
Problem
Section titled “Problem”Model calls are the dominant cost and the dominant latency in most AI products, and a large share of them are repeats. An exact-match cache captures almost none of that, because people do not ask questions the same way twice — “how do I reset my password” and “how can I reset my password” are one question and two cache keys.
Semantic caching closes that gap by matching on embedding similarity instead of string equality. It is one of the highest-leverage cost levers available, and it is the one that most quietly changes what your system returns.
The reason is a category difference the architecture has to take seriously: an exact-match cache can only be slow or empty; a semantic cache can be wrong. A miss costs a model call. A false hit serves a confidently incorrect answer, fast, in a form nothing downstream can distinguish from a correct one — not the caller, not the model, not the logs, not the evaluation suite. The engineering problem is not how to raise the hit rate. It is how to know what the hit rate is costing.
Requirements
Section titled “Requirements”- An exact tier before the semantic tier, because a string match is free, certain, and answers a large share of traffic on its own.
- A scope key that makes a hit valid — at minimum model, prompt version, and tenant. An answer is an answer from a particular model, under a particular prompt, for a particular tenant.
- Scope filtered before similarity search, not after.
- False-hit rate measured, not just hit rate, with a mechanism that keeps measuring it in production.
- Invalidation on the things that change answers — prompt edits, model changes — rather than waiting for a TTL.
- A bypass path, so evaluation and debugging can reach the model rather than the cache.
Constraints
Section titled “Constraints”- The threshold cannot be inherited. It is a joint property of your corpus and your embedding model. A number that works in someone’s blog post, or in this architecture’s backing lab, carries no information about yours.
- False hits have no downstream symptom. Nothing in the response distinguishes one. The only detection is deliberately sampling cached answers against fresh ones, which almost nobody does.
- Entity swaps are a known hard case. The token carrying the semantic difference —
EUversusUS,2025versus2026,freeversuspaid— is short and low-weight, and it is swamped by the long shared framing around it. Meanwhile paraphrases differ in function words, which are a larger fraction of a short question. The pairs that must never be conflated can be more similar than the pairs that should hit. - Prompts change far more often than models do. Weekly edits are normal, and every edit that changes behaviour keeps serving the old behaviour until something invalidates it.
- The cache sits underneath your evaluation. An eval suite running through the cache measures the cache, and will keep reporting yesterday’s scores after a model change.
Request Flow
Section titled “Request Flow”flowchart TD
Q["Incoming question"] --> Scope{"Exact key hit?<br/>tenant + model + prompt_version + text"}
Scope -->|Yes, unexpired| Exact["EXACT hit<br/>always correct"]
Scope -->|No| Filter["Filter candidates by scope<br/>BEFORE similarity, never after"]
Filter --> Near["Nearest neighbour<br/>cosine similarity"]
Near --> Cut{"score >= threshold?"}
Cut -->|No| Miss["MISS<br/>call the model<br/>similarity still reported, so<br/>the threshold can be tuned"]
Cut -->|Yes| Risk{"Is it the same question?"}
Risk -->|"Genuine paraphrase<br/>0.45 - 0.83 here"| True["TRUE hit<br/>cost saved"]
Risk -->|"Entity swap<br/>EU/US · 2025/2026 · free/paid<br/>0.60 - 0.91 here"| False["FALSE hit<br/>wrong answer, served fast<br/>indistinguishable downstream"]
True --> Note["The ranges overlap.<br/>Tightening to exclude the right column<br/>also excludes the left one."]
False --> NoteTwo orderings carry the design.
Exact before semantic, because it costs a hash lookup, has no false-hit risk at all, and takes a meaningful share of traffic off the expensive path before any similarity question is asked.
Scope filter before similarity search, which is the one that matters most and is easiest to get backwards. Filtering after the search means the nearest neighbour was already read out of another tenant’s data before you discarded it. The vector store’s ANN index and the tenant boundary have to be the same query, not two steps — the Vector DB lookup’s first gotcha, arriving here as a cache.
Note what the diagram makes explicit: a semantic hit has two outcomes, and the system cannot tell them apart at serve time. Everything below follows from that.
Failure Modes
Section titled “Failure Modes”Reporting hit rate without false-hit rate
Hit rate is the number that gets a cache approved, and on its own it is an argument for lowering the threshold until the cache is wrong most of the time. The two curves have to be read together, because raising one raises the other, and only one of them appears on the cost dashboard.
Treating a tighter threshold as the fix
Tightening trades false hits for misses. It does not separate the two populations, and where they overlap badly there is no setting that gets one without the other. Tuning cannot fix a threshold problem whose real cause is that similarity is not equivalence.
Filtering scope after the similarity search
A post-filter means the search already crossed the tenant boundary and you are discarding the evidence. It usually passes tests, because the filter does remove the wrong-tenant result — right up until the day a result is returned before the filter runs, or a bug drops the filter entirely. The isolation was never structural.
Prompt version missing from the scope key
The most common real-world failure, because prompts are edited weekly and models change rarely. An edit that changes the answer keeps serving the previous behaviour for the whole TTL, and the symptom is a change that “did not deploy” — with a deploy log saying it did.
Evaluation running through the cache
The eval suite reports the cache’s answers, so scores stay flat across a model upgrade and a regression is invisible until it reaches users. Evaluation needs a bypass, and the bypass needs to be the default in that path rather than a flag someone remembers.
TTL as the only invalidation
A TTL bounds staleness by time, which is unrelated to what actually changed. It answers “how old is this” and never “is this still the answer” — so a prompt fix waits out an interval chosen for unrelated reasons.
Scaling
Section titled “Scaling”- Linear scan is fine for longer than people expect, and then suddenly is not. Below a few thousand entries per scope a brute-force cosine scan beats the operational cost of an index. Above that, an ANN index becomes necessary — and it changes the failure surface, because approximate search means the nearest neighbour is now itself approximate.
- Scope keys partition naturally, which is what keeps the search space small: the candidate set is one tenant’s entries under one prompt version, not the whole corpus. This is a scaling property and an isolation property at once, which is unusually convenient.
- Every lookup pays an embedding call. The cache does not save a model call for free; it trades a small model call for a large one. That trade is excellent, and it means the embedding endpoint inherits the request rate of the entire product.
- Cache misses stampede. A popular question that expires sends every concurrent caller to the model at once. Single-flight the miss path, or the cache amplifies the load it exists to reduce.
- Write volume is proportional to misses, so the cache is busiest writing exactly when it is least useful, which is the first hour after a prompt change invalidates everything.
Security
Section titled “Security”- Cross-tenant leakage is the headline risk, and the mitigation is structural rather than procedural: the scope filter belongs inside the query the vector store executes.
- A cache is a store of answers, therefore a store of whatever was in them. It inherits the data classification of the responses it holds, and usually gets a retention policy written for performance infrastructure.
- The prompt version in the key is a security control, not only a correctness one, when the prompt change was the fix for a jailbreak or a data-exposure bug.
- Timing is an oracle. A fast answer means a cache hit, which means someone asked something similar before; in a multi-tenant deployment that leaks the existence of other tenants’ queries. Usually acceptable, occasionally not, and worth deciding rather than discovering.
Trade-offs
Section titled “Trade-offs”Threshold: false hits vs. missed savings
The threshold is a dial between two costs, and the honest position is that it cannot be set without a number for what a wrong answer costs. Set it by hit-rate targets and you have chosen a false-hit rate by accident. On corpora where the two populations invert, the correct choice may be to serve only the exact tier and take the smaller saving.
Exact-only vs. semantic
Exact matching cannot be wrong and captures far less. Semantic matching captures the paraphrases — which is most of the repeat traffic — and introduces a correctness surface that did not previously exist. For a support assistant, that trade is usually worth making with care; for anything whose answers turn on entities the user supplies, exact-only is a defensible final answer rather than a failure to optimise.
TTL vs. versioned scope keys
A TTL is one line and bounds staleness by time. Versioned keys invalidate on the thing that actually changed, cost nothing at read time, and require the discipline of bumping a version on every prompt edit. Running both is the practical answer: versions for correctness, a TTL as the backstop for the edit somebody forgot to version.
Detecting false hits vs. paying for them
Sampling a percentage of cache hits against a fresh model call gives you a real false-hit rate and spends a fraction of the savings to get it. Skipping it keeps the full saving and leaves you unable to answer how often the cache is wrong. Teams almost always choose the second, then discover the rate during an incident.
- The saving is real and large — a hit replaces a full model call with an embedding call and a vector lookup, typically two orders of magnitude cheaper and faster.
- The embedding call is the cost floor, paid on every lookup including misses, so a cache with a low hit rate is a small tax rather than a saving.
- Storage is the cheap part; responses are text, and TTLs keep the working set bounded.
- The dominant cost is the one that never appears on an invoice. A false hit costs whatever a wrong answer costs in your product — a support deflection that misinforms, a policy quoted for the wrong region. Until that has a number, every threshold discussion is a debate about a dashboard.
A threshold from someone else's system is not evidence about yours
The backing lab swept thresholds on its corpus and found no setting where the cache served a genuine paraphrase before it served a wrong answer — at every threshold that served anything at all, precision was below 100%. Two things about that result should travel and one should not.
Travelling: tightening the threshold is not the fix, and a threshold cannot be inherited — it is a joint property of a corpus and an embedding model, and the only way to know yours is to build a labelled set of near-duplicates and measure both curves.
Not travelling: the specific numbers, and the strength of the effect. That lab’s embedding model is a deterministic lexical stand-in, not a learned one, and a real dense model would separate these populations better. The lab says so in its own tests rather than in a footnote — one exists specifically so a reader reproducing the result with a real model knows which assertion is expected to weaken.
Observability
Section titled “Observability”- False-hit rate from sampling, produced by re-running a small percentage of hits against the model and comparing. This is the only metric here that cannot be derived from the cache’s own behaviour, and it is the one the design turns on.
- Hit rate split by tier — exact versus semantic. They have different risk profiles and averaging them hides which one is growing.
- The distribution of similarity scores on hits, not the mean. A cluster of hits sitting just above the threshold is where the false hits live, and a mean of 0.94 says nothing about it.
- Miss reasons — no scope match, below threshold, expired. A rise in scope misses after a deploy is a prompt version bump working correctly; a rise with no deploy is a key being computed inconsistently.
- Hit rate immediately after a prompt version bump, which should drop to zero and recover. If it does not drop, the version is not in the key.
- Age distribution of served entries, which is how staleness becomes visible before someone reports it.
Production Deployment
Section titled “Production Deployment”Before real traffic
- The exact tier runs before the semantic tier.
- The scope key includes model identifier, prompt version, and tenant, with a test that changing any one of them misses.
- The scope filter is applied inside the vector query, not after the results return.
- The threshold was chosen against a labelled set of near-duplicates from your own corpus, with both curves measured — not inherited from a blog post or from this page.
- False-hit rate is sampled continuously in production, and someone owns the number.
- There is a bypass path, and evaluation uses it by default.
- Prompt edits bump the version as part of the deploy, not as a manual step.
- The miss path is single-flighted so an expiry does not stampede the model.
- The cache’s retention and data classification match the responses it stores.
- What a wrong answer costs has been written down, because the threshold encodes an answer to it either way.
Hands-on Lab
A running implementation: a two-tier lookup, scope keys covering model, prompt version, and tenant, scope filtering ahead of the similarity search, and a threshold sweep reporting hit rate, false-hit rate, and precision together. Its embedding model is a lexical stand-in, and the lab is explicit about how much of the headline result that explains — including a test written so the assertion expected to weaken under a real model is named. Read the lab documentation →
labs/semantic-cacheproduction-shaped
Interview Questions
Section titled “Interview Questions”What is in the cache key for an LLM response, and why?
The normalised request, plus the model identifier, the prompt version, and the tenant. A cached answer is an answer from a particular model under a particular prompt for a particular tenant, so each of those changing must miss. Prompt version is the one people leave out, and it is the one that changes weekly — omit it and every prompt fix keeps serving the old behaviour until the TTL expires.
Your semantic cache has a 60% hit rate. What do you need to know before shipping it?
The false-hit rate. Hit rate on its own is an argument for lowering the threshold until the cache is usually wrong, because both curves move together and only one is on the cost dashboard. And since a false hit is downstream-indistinguishable from a correct answer, the number does not exist unless you build it — by sampling hits against fresh model calls, continuously, in production.
Paraphrases score 0.45–0.83 and near-duplicates with different answers score 0.60–0.91. What threshold do you pick?
None — that is the point of the question. The populations are inverted, so every threshold that serves a paraphrase also serves a wrong answer, and tightening removes the right answers before the wrong ones. The available moves are to serve only the exact tier, or to detect that two questions differ on an entity that changes the answer. Threshold tuning trades one error for the other; it cannot eliminate both.
Why does the tenant filter have to be inside the vector query?
Because filtering afterwards means the nearest-neighbour search already read another tenant’s entries and you are discarding the result rather than never producing it. The isolation is procedural instead of structural, so it holds only as long as every code path remembers to apply it. Inside the query, the boundary is a property of the search.
Your evaluation scores did not move after a model upgrade. What would you check first?
Whether the eval ran through the cache. If it did, it measured the cache and will keep reporting the previous model’s answers indefinitely — which looks exactly like a model change that did nothing. Evaluation needs a bypass path, and it should be the default in that path rather than a flag someone has to set.
Where did you get the threshold?
From a labelled set of near-duplicates drawn from the product’s own traffic, swept with both hit rate and false-hit rate reported. A threshold is a joint property of a corpus and an embedding model, so a number taken from anywhere else — a vendor default, a paper, another team — is not evidence about this system. Answering with a number and no measurement behind it is the failure mode this question is looking for.