Skip to content

Vector DB

Listen to this page7:34
Read the transcript

1. It’s not storage — it’s a search category

Host: So let’s start with the phrase everyone throws around: vector database. I think most people hear that and picture a database that stores embeddings, basically a fancy filing cabinet for vectors. Is that actually right?

Guest: It’s the most common misconception, and it undersells the whole thing. A vector DB is a search system for high-dimensional vectors — given a query vector, it returns the nearest ones fast, approximately. Storage is the least interesting part of it, and honestly hasn’t been the differentiator in a long time. And here’s the other thing: there’s no single ‘vector DB version’ to point to. It’s a category, not a product — index implementations differ enough between vendors that what sounds like a portable claim often isn’t. We’ll use Pinecone’s 2026-04 API as a concrete anchor when we need one, but the concepts are vendor-neutral.

Host: Okay so if it’s a search category, what’s actually being searched? I keep hearing about dense vectors and lexical search like they’re competing approaches.

Guest: They used to be treated as separate paths — embeddings go into a vector index for ANN search, while raw text goes into a lexical or full-text index for keyword search. But the category is converging: platforms increasingly fuse both into hybrid retrieval, then rerank before handing off context. Pinecone’s full-text search preview is a good signal of that — in July 2026 it picked up fuzzy matching and n-gram substring search, which tells you where this is headed: the vector store is becoming the retrieval layer, not just the vector layer.

2. The vocabulary that actually determines behavior

Host: Okay, so before we go further, let’s nail down the vocabulary, because I think this is where people’s mental models get fuzzy. Start with the basics — dense versus sparse, and then this ANN thing everyone name-drops.

Guest: Dense means an embedding, a vector — similarity there is geometric, so meaning-adjacent text scores highly even if the words don’t match. Sparse or lexical is term-based, so exact identifiers, codes, and rare words that embeddings tend to blur actually live there. ANN, approximate nearest-neighbor, is the index structure that makes searching millions of dense vectors fast — it trades perfect recall for speed, and the tuning knobs for that tradeoff vary a lot by vendor. On top of that you’ve got two things fixed at index creation time that people don’t realize are permanent: the similarity metric — cosine, dot product, or Euclidean, which has to match how the embedding model was actually trained — and the embedding dimension, so if you swap models later, you’re not updating an index, you’re building a new one.

Host: So those two are basically load-bearing walls you can’t move after the fact. What about the knobs people actually touch at query time — top-k, filtering, namespaces, reranking?

Guest: Top-k is how many candidates come back, and it’s really a recall ceiling — if the right document isn’t in that returned set, no amount of clever reranking downstream saves you. Metadata filtering restricts candidates by tenant, ACL, source, or recency, and critically it’s evaluated as part of the search itself, not as a filter bolted on after. Namespace is your partition within an index, usually the unit of tenant isolation, and reranking is a stronger model re-scoring that shortlist afterward — often served by the platform itself now rather than something you bolt on separately.

3. The numbers that bite: versioning, fixed choices, and latency

Host: Let’s get concrete, because I feel like this is where teams get burned in production. You mentioned Pinecone has date-based API versioning — walk me through why that’s not just a footnote.

Guest: So Pinecone ships a new stable API version quarterly, and each one is supported for at least twelve months, which gives you roughly nine months of overlap to migrate. Current stable is 2026-04, and you’re supposed to send that explicitly as a header. The trap is what happens if you don’t: an unversioned call doesn’t default to the newest version, it falls back to the oldest supported stable one — so silently, you’re pinned to the past, not the present.

Host: That’s a nasty default. What are the other choices that quietly lock you in, ones you don’t get a do-over on?

Guest: Embedding dimension and similarity metric are both fixed the moment you create the index. Change your embedding model down the line and that’s not a config tweak, it’s a full reindex of your entire corpus. And if the similarity metric doesn’t match what the embedding model was trained for, you don’t get an error — you get results that are wrong but not obviously wrong, which is worse. On top of that, don’t assume index type either, since ‘all vector databases use HNSW’ is false and tuning advice doesn’t port between vendors. And the latency gotcha nobody budgets for is reranking — ANN search plus filtering plus network is often smaller than that final rerank pass.

4. Where teams actually get burned

Host: Let’s get concrete about the war stories. What’s the failure mode that keeps showing up when teams put multi-tenant data into one of these systems?

Guest: The filtering-after-retrieval trap. If the tenant filter isn’t part of the actual ANN query, the system has already read another tenant’s documents in the traversal and is just deciding afterward whether to show them to you. That’s not a filter, that’s a leak that happened to not surface this time — and restrictive filters have their own version of this problem, where they interact badly with the approximate index and quietly return way fewer than k results, or worse ones, unless you actually measure recall with your real filters applied.

Host: And the deletes issue — walk me through why that one’s so insidious compared to a normal bug.

Guest: Because it doesn’t fail loudly, it fails convincingly. A document gets removed at the source, nobody removes it from the index, and it keeps getting retrieved and cited as if it’s current — the system isn’t broken, it’s confidently wrong. Pair that with people assuming ‘approximate’ is just a performance knob rather than a stated tradeoff, or copying another vendor’s tuning advice wholesale, and you get the same pattern every time: it looks like it’s working until someone checks.

5. The store inside the bigger system

Host: So if the vector DB isn’t the whole story, where does it actually sit? Because everything we’ve talked about — the ANN math, the fixed choices, the fragile syncing — feels like it’s solving one piece of a much bigger puzzle.

Guest: It’s the retrieval half of a RAG pipeline, and the crucial thing people miss is that retrieval quality and answer quality get measured separately — a wrong answer should be traceable to a document, not a mystery. The vector store is doing candidate retrieval; there’s still chunking, hybrid search, reranking, and access control sitting around it, and those are architecture decisions, not store features. So ‘just add a vector DB’ undersells the job — you’re really building a retrieval system, and the store is just the piece that happens to do the ANN search.

Not covered

The planner wanted these and found nothing in the source to support them:

  • Benchmark comparisons between specific vector DB vendors’ throughput or cost
  • A walkthrough of setting up a live Pinecone index or writing queries against it

Generated from this page by Claude Sonnet 5 on , spoken by Kokoro-82M running locally. Two synthetic voices, not a recorded conversation. Every claim is drawn from this page — where it differs from the text above, the text is correct.

A search system for high-dimensional vectors: given a query vector, return the nearest ones, fast, approximately. “A database that stores embeddings” undersells it — storage is the least interesting part, and it has not been the differentiator for some time.

There is no “Vector DB version” — it is a category, not a product, and index implementations differ enough between vendors that portable claims are narrower than they look. This page stays vendor-neutral and uses Pinecone’s 2026-07 API as the versioned anchor wherever a concrete version is needed.

The two retrieval paths, and where they meet:

Embedding → dense vector → vector index → ANN search ─┐
├→ hybrid fusion → reranking → context
Text → lexical / full-text index → keyword search ────┘

Modern platforms increasingly serve both sides of that diagram. Pinecone’s full-text search gained fuzzy matching and n-gram substring search in public preview in July 2026, and went generally available on 2026-09-02 — BM25 ranking, Lucene query syntax, and vector similarity in one index. That is the category direction in one data point: the vector store is becoming the retrieval layer, not just the vector layer.

Concept What it does
Dense vector An embedding. Similarity is geometric, so meaning-adjacent text scores highly
Sparse / lexical index Term-based matching. Exact identifiers, codes, and rare words live here
Hybrid search Both paths, fused into one candidate list
ANN index Approximate nearest-neighbor structure. Trades exact recall for query speed — implementation and tuning knobs vary by vendor
Similarity metric Cosine, dot product, or Euclidean. Fixed at index creation and must match how the embedding model was trained
Embedding dimension Fixed per index. Changing the model means a new index
Top-k How many candidates come back. A recall ceiling for everything downstream
Metadata filtering Restricting candidates by tenant, ACL, source, or recency — evaluated as part of the search, not after it
Namespace Partition within an index. The usual unit of tenant isolation
Reranking A stronger model re-scoring the shortlist, often served by the platform itself
Index update / upsert Writing new or changed vectors. Visibility is not necessarily immediate
Deletion Removing vectors. The step people forget when a source document is deleted or a tenant leaves
Backup and restore Point-in-time recovery of an index — including, increasingly, across regions
RBAC / SSO Role-based access, with SAML/SCIM-driven provisioning on managed platforms
Data residency Which region the vectors physically live in. A contractual constraint, not a preference
Quantity Value Why it matters
The category’s version None “Vector DB” is a category. Pin a vendor’s API version instead
Pinecone’s current stable API 2026-07 Date-based versioning; send X-Pinecone-Api-Version: 2026-07 explicitly. It replaced 2026-04 this quarter
Index creation on 2026-07 Schema-only dimension, metric, and spec move into a schema. Hand-written REST calls break; SDK create_index does not
Pinecone stable API support At least twelve months New stable version quarterly, so you get ~9 months to migrate off the one you are on
Unversioned API calls Fall back to the oldest supported stable version Not the newest. Omitting the header pins you to the past
Embedding dimension Fixed at index creation A model change is a reindex of the entire corpus
Similarity metric Fixed at index creation Mismatched with the embedding model, results are wrong but not obviously wrong
Index type Vendor-specific “All vector databases use HNSW” is false; do not port tuning advice between vendors
Isolation unit Namespace, usually Per-tenant indexes cost more and isolate harder — pick deliberately
Latency contributors ANN search + filtering + rerank + network The rerank stage is frequently the largest and the least budgeted
  • Filtering after retrieval is a leak, not a filter. If the filter is not part of the query, the system has already read another tenant’s documents and is deciding what to do about it afterwards.
  • Changing the embedding model invalidates the whole index. Old and new vectors are not comparable. This is a migration with a backfill, not a config change.
  • Deletes are the forgotten half of ingestion. A document removed at the source but not from the index keeps being retrieved and cited — the most convincing kind of wrong answer.
  • Metadata filters can quietly wreck recall. A restrictive filter over an approximate index can return far fewer than k results, or worse ones, because the filter and the ANN traversal interact. Measure recall with your real filters applied.
  • Approximate means approximate. ANN search misses true nearest neighbors by design. If a use case needs exactness, that is a requirement to state, not a knob to turn up.
  • Vendor tuning advice is not portable. Index structures, parameters, and their failure behaviors differ. Advice about one vendor’s knobs may be actively wrong for another’s.
  • Omitting the API version does not mean “latest”. Pinecone resolves an unversioned call to the oldest supported stable API, so the safe-looking option is the stale one.
  • Bumping the API version is a code change, not a header change. 2026-07 made index creation schema-only, so a hand-written POST /indexes that works on 2026-04 fails on 2026-07. Read the version’s breaking changes before moving the header — the twelve-month window exists for exactly this.
  • Upserts are not immediately visible everywhere. Write-then-read-your-own-write is not a guarantee to assume. Test it before building a workflow that depends on it.
  • Semantic search will not find an error code. This is not a tuning failure; it is the wrong index. Use the lexical path, or hybrid.
  • Cost tracks vectors, dimensions, and replicas, not documents. A chunking change that triples chunk count triples the bill without any new content.
  • RAG — the pipeline this sits inside, and why retrieval quality is measured separately from answer quality.
  • Module 8: RAG — chunking, hybrid retrieval, re-ranking, and retrieval-time access control in full.
  • Enterprise RAG Platform — tenancy, freshness, and access control as architecture requirements rather than store features.
  • hybrid-retrieval — dense and BM25 retrieval fused by reciprocal rank fusion, with reranking and a groundedness check, as running code.
  • Redis — the other data store in most of these systems, doing caching, rate limiting, and leases.