Trust is earned, not given

A different perspective

2023-11-16 · Projects

AI Frontiers, part 6: Vector databases β€” the missing infrastructure layer

Part 6from the AI Frontiers series · 65 parts in all

Part 5 ended with a claim: the capability that matters lives in systems, not weights. This entry is about the piece of infrastructure that claim conjured into existence. Between March 2023 and November 2023, the vector database went from specialist curiosity to funded category β€” Pinecone raised a $100M Series B at a reported $700M+ valuation in April, Weaviate took $50M in April, Qdrant, Chroma, and Milvus/Zilliz all raised through the year β€” and every serious LLM stack grew a vector store as its neighbor. This entry is the engineer's tour: what the thing actually computes, why the algorithmic core is older than the hype, and my honest read on whether the category survives its own success.

The problem, stated small

You have ten million documents, each embedded as a 1,536-float vector. A query arrives; embed it the same way; return the forty nearest vectors. Brute force β€” compute 10 million cosine distances, sort, take top-k β€” is embarrassingly simple and costs O(NΒ·d) per query. At 10 million vectors and 1,536 dimensions, one query is billions of multiply-adds: tens of milliseconds per core, hundreds of queries per second needed, and now your retrieval bill exceeds your model bill. The entire category exists because brute force does not survive contact with production scale and latency budgets.

The escape is approximate nearest-neighbor (ANN) search: give up the guarantee of the exact top-k in exchange for orders-of-magnitude speed. Recall@k β€” how often the approximation returns the true top-k β€” becomes a tunable knob, and 95–99% recall at 100Γ— speedup is the usual sweet spot. For RAG, approximate is fine: the goal is fishing relevant context out of a corpus, not solving a metric-space optimization.

The algorithms behind the products

Three algorithm families dominate, and every product conversation in 2023 was secretly one of these three wearing a brand.

HNSW β€” hierarchical navigable small-world graphs. Build a multi-layer proximity graph over the vectors; search descends from a sparse top layer through progressively denser ones, greedily hopping toward the query. Concretely: a query on 10M vectors touches a few thousand nodes, not ten million. Search is O(log N)-ish with excellent recall/speed curves; the costs are memory (the graph roughly doubles footprint) and build time (deletes and updates are awkward β€” tombstones pile up until compaction). Malkov and Yashunin's paper is the reference (Malkov and Yashunin). Most hosted products made HNSW the default index by 2023 because its latency tails are the best behaved.

IVF β€” inverted file indexes with quantization. Cluster vectors into thousands of cells with k-means; at query time, probe only the nearest few cells, optionally compressing the vectors inside each cell with product quantization (PQ) so distance computations run on 8-bit codes instead of 32-bit floats. IVF-PQ trades recall for memory more aggressively than HNSW and shines when the corpus is huge and RAM is the binding constraint. FAISS β€” Meta's library, the academic and industrial workhorse underneath many products β€” is where these pieces live in their sharpest form (Johnson et al.). Faiss's combination of IVF, PQ, and GPU kernels set the performance ceiling the products marketed against.

Disk-native indexes. The late-2023 twist: Turbopuffer and others betting that object storage plus a smart index beats RAM economics; Cloudflare's Vectorize entered beta in September with a similar pitch. If your corpus is billions of vectors but your queries-per-second are modest, paying to keep all of it hot in RAM is the wrong bill. This is the classic storage-hierarchy argument re-fought with vectors, and my default prediction is that it ends the way those arguments usually end: cold data on cheap tiers, hot shards in memory, nobody remembering it was ever a controversy.

Alongside indexes, the 2023 product surface converged on the same feature set: metadata filtering (pre-filter the candidate set by tenant, date, ACL before ANN runs β€” mandatory for multi-tenant anything), hybrid retrieval (BM25 and dense fused, because rare terms and exact identifiers still defeat embeddings), and the operational layer β€” replication, snapshots, namespace-per-customer. The smarter products realized that the index is a commodity and the metadata layer is the moat.

Embeddings: the load-bearing assumption

Everything above assumes the embeddings are good β€” that geometric proximity means classessemantic relevance. Worth pausing on, because the entire category inherits this assumption and 2023 gave it a stress test. The embedding models of early 2023 were decent at topical similarity and unreliable at anything requiring precision: negation ("ships to Europe" vs. "does not ship to Europe" embed nearly identically), exact identifiers (SKU codes, part numbers β€” tokenization eats them), recency ("current price" has no stable vector), and rare-domain vocabulary. The MTEB leaderboard β€” Hugging Face's massive text embedding benchmark β€” became the reference for picking models, and its churn through 2023 (top model dethroned every few months, first by OpenAI's ada-002 lock-in, then by open models like BGE and E5) made embedding choice a quarterly decision rather than an annual one (Muennighoff et al.).

The production consequences were three. Hybrid retrieval stopped being optional β€” every serious deployment fused BM25 with dense search because lexical matching covers exactly what embeddings miss. Re-embedding became a planned migration with a budget line, which is why storing raw text beside vectors went from best practice to moral imperative. And customer-specific models β€” fine-tuned embedders trained on your domain's click or feedback data β€” emerged as the quiet advantage of teams with real usage volume, since retrieval quality is ultimately a measurement of your corpus against your users' actual queries, not a public benchmark.

One more failure mode belongs in the permanent record because it consumed so many weekends: chunk-to-query mismatch. The embedding of a 500-token passage is a lossy compression of that passage; a 10-word user query compresses to a vector that may sit closer to a passage's topic than to the specific fact needed. The mitigations that worked were all forms of making the comparison finer-grained: smaller chunks with neighborhood expansion, multiple embeddings per document (one per section, one per claim), query rewriting into the corpus's own vocabulary, and the cross-encoder reranking described in part 5. The general lesson: the vector is a lossy index, and mature systems treat it like one β€” as the first filter, never the final judge.

Do you even need one?

The honest engineering answer, as of this month: often, no. Postgres gained pgvector with HNSW support moving from patch to mainstream through 2023 (pgvector 0.5.0 shipped HNSW in August); Elasticsearch and OpenSearch shipped dense-vector kNN; Redis added vector similarity. If your corpus fits in the low millions of vectors and your team already runs a real database, putting embeddings next to your relational data β€” one query for joins and filters, one backup story, one permission system β€” beats a second datastore on every operational axis I care about. The dedicated stores earn their keep at genuine scale, at extreme latency requirements, or when managed-metadata-features are the point.

My default in production designs, then and now: start with pgvector in the database you already operate; graduate to a dedicated store when you hit its walls β€” not when the architecture diagram would look cooler. The counter-pressure is real: the category raised serious money and will out-market Postgres forever. But infrastructure decisions age better when made on operational grounds, and "one less system" is an operational ground.

Twelve months, one procurement guide

Close the category tour with the question every engineering manager actually asked in 2023: which one do we buy? Strip the marketing and the decision reduces to five questions, so here they are as a field guide. How big, really? Under a few million vectors and modest QPS, pgvector inside your existing Postgres is the correct answer, and the end of the analysis. What is the freshness contract? If documents update hourly and queries expect current data, disk-native and tombstone-heavy designs need scrutiny; if the corpus is rebuilt nightly anyway, almost anything works. Who owns the metadata layer? ACL filtering, tenant namespaces, and hybrid BM25 fusion are where the dedicated products earn their keep β€” if you need all three at scale, that is their market. What happens at the migration? Embedding-model changes are certain; ask each vendor how re-embedding ten million vectors is priced and parallelized before you have ten million vectors. And what does the exit look like? Vectors plus raw text plus metadata export in open formats (Parquet for the corpus, standard JSON for metadata) keeps every vendor honest, including the one you chose.

The pattern underneath the guide is the same one that governed databases for forty years: the managed service wins the demo, the boring embeddable wins the decade. Which is not a prediction that the dedicated vector stores fail β€” their metadata and hybrid retrieval layers are genuinely good products, and corpus scale genuinely splits them from Postgres β€” but a reminder that the center of gravity in any data infrastructure market settles on operational simplicity, not benchmark leadership. Watch the 2024 releases in this category: every serious player added filtering, hybrid search, and export. None of those show up in the ANN benchmarks; all of them show up in renewals.

Operations: the part the demos skip

Every product comparison ends at recall@10 on a benchmark; every incident begins somewhere else entirely. The operational lessons of 2023's first wave of vector-store productions are worth their own section, because they repeat across teams with eerie uniformity.

Rebuild is a when, not an if. Index format changes, embedding model changes, chunking strategy changes β€” a mature pipeline gets rebuilt several times a year. Whatever you choose, keep the raw text, the chunking config, and the embedding model version in a manifest so the rebuild is a rerun, not an archaeology project. Teams that stored only vectors in the store, source documents elsewhere, and the chunking script in someone's head spent their winter rebuilding by hand.

Deletion is harder than insertion. The right-to-delete is a legal requirement in some deployments and an operational nuisance in all of them. HNSW deletes are tombstones until compaction; IVF cell membership makes purging partial; and the failure mode that actually bites is forgetting that the source documents and the derived chunks must be deleted transactionally β€” leave one orphan chunk and your GDPR request has a lingering tail. The clean pattern: derive everything from a source of record you control, delete there, and let an idempotent reindex propagate.

Multitenancy is a data-plane decision, not a query filter. Adding "WHERE tenant_id = X" to a vector query is necessary and nowhere near sufficient: noisy-neighbor effects on shared indexes, per-tenant embedding-version drift, and per-tenant corpus update cadences all push toward namespace-per-tenant or index-per-tenant at the top end. The hosted products' namespace features existed for exactly this; using them from day one is cheaper than retrofitting isolation after the first cross-tenant leak, which is the sort of incident that ends pilots.

Finally, a monitoring note that generalizes beyond this layer: track retrieval quality trends, not just system health. Latency and uptime dashboards were standard on every deployment; a weekly sample of (query, retrieved passages, judge score) was not β€” and the teams that had it caught embedding-model regressions in days instead of quarters. The vector store made retrieval observable; somebody still has to look.

Why this layer matters to the series arc

Zoom out. The transformer collapsed model architecture into a uniform stack (part 1); RAG moved knowledge out of weights into retrievable corpora (part 5); the vector store is the physical embodiment of that move β€” the place where "what the model knows" became "what the system indexes." From here the series follows the infrastructure consequences: context windows (part 7) attack the same problem by brute-forcing everything into the prompt; mixture-of-experts (part 8) restructures the weights themselves; small models (part 9) ask how little weight is enough when the system around the model is this good. The pattern to carry forward: every capability question in 2024 is secretly a memory-hierarchy question.

Works Cited

"ANN Search: 100x Faster Than Brute Force." Faiss Wiki (Facebook Research), 2023, github.com/facebookresearch/faiss/wiki. Accessed 16 Nov. 2023.

Johnson, Jeff, Matthijs Douze, and HervΓ© JΓ©gou. "Billion-Scale Similarity Search with GPUs." IEEE Transactions on Big Data, vol. 7, no. 3, 2021, pp. 535–47, ieeexplore.ieee.org/document/7878393. Accessed 16 Nov. 2023.

Malkov, Yu A., and Dmitry A. Yashunin. "Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs." IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 42, no. 4, 2020, pp. 824–36, ieeexplore.ieee.org/document/8581719. Accessed 16 Nov. 2023.

Muennighoff, Niklas, et al. "MTEB: Massive Text Embedding Benchmark." Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, ACL, 2023, aclanthology.org/2023.eacl-main.148. Accessed 16 Nov. 2023.

"Pinecone Raises $100M to Fuel the AI Revolution." Pinecone, 6 Apr. 2023, www.pinecone.io/learn/pinecone-funding/. Accessed 16 Nov. 2023.

"pgvector: Open-Source Vector Similarity Search for Postgres." GitHub, 2023, github.com/pgvector/pgvector. Accessed 16 Nov. 2023.

Robertson, Stephen, and Hugo Zaragoza. "The Probabilistic Relevance Framework: BM25 and Beyond." Foundations and Trends in Information Retrieval, vol. 3, no. 4, 2009, pp. 333–89, doi:10.1561/1500000019. Accessed 16 Nov. 2023.