Trust is earned, not given

A different perspective

2024-09-05 · Projects

AI Frontiers, part 43: Embeddings as infrastructure β€” search, clustering, and drift

Part 43from the AI Frontiers series · 65 parts in all

There is a component in almost every AI system I have shipped that nobody puts on a slide deck: the embedding model. It sits between raw text and everything interesting β€” retrieval, deduplication, clustering, classification, anomaly detection, recommendations, the "related posts" box, the ticket router. In part 6 I argued that vector databases were the missing infrastructure layer of 2023; a year later I think I put the emphasis in the wrong place. The database is a commodity now β€” three good open-source options and every cloud vendor's managed service. The part that actually decides whether your system works is the embedding model and the maintenance regime around it, and that part is treated as a constant when it is anything but.

This entry is the case for treating embeddings as infrastructure: something you version, monitor, budget for, and migrate β€” not something you import from a library and forget. It is also an argument about why dense similarity is a poor proxy for relevance, and what to do about it.

The training recipe, and why it matters to users

Modern sentence embeddings descend from a simple observation in the Sentence-BERT paper: take a pretrained transformer, run two sentences through it with shared weights, pool the token representations into one vector each, and train with a contrastive objective that pulls matching pairs together and pushes mismatched pairs apart (Reimers and Gurevych, "Sentence-BERT"). SimCSE showed that the positive pairs could be almost free β€” the same sentence passed through the encoder twice with different dropout masks β€” and that this alone produced strong embeddings (Gao et al.). The field then industrialized the negative-sampling side: in-batch negatives scaled to thousands per step, hard negatives mined from an existing retriever, and eventually instruction-prefixed objectives that let one model serve many tasks (Wang et al., "Text Embeddings by Weakly-Supervised Contrastive Pre-Training").

Two consequences of that recipe matter if you are the person operating the system. First, embedding quality is domain-conditional. These models are trained on general web text; the contrastive task they solved was "which of these web passages answers this web query." Your corpus is not web text, and your queries are not web queries. The BEIR benchmark made this concrete in 2021 by evaluating retrievers across eighteen zero-shot datasets: dense models that dominated in-domain benchmarks lost to plain BM25 on several out-of-domain collections (Thakur et al.). That result has aged well. If you have never measured your embedding model on your own corpus, you do not know whether it beats keyword search β€” and repeatedly, it does not.

Second, the training objective is baked into the geometry. Instruction-tuned embedders from 2024 expect an instruction prefix on the query and none on the document; getting that backwards degrades recall silently rather than loudly. Models trained with Matryoshka representation learning let you truncate the vector to a shorter prefix and keep most of the quality (Kusupati et al.), which is genuinely useful for cost control and easy to get wrong if you truncate a model that was not trained that way. The practical rule: read the model card, and treat the preprocessing (prefixes, truncation, normalization) as part of the model version, because it is.

Dense similarity is not relevance

Cosine similarity between two vectors measures a learned notion of topical overlap. It does not measure authority, recency, specificity, or whether the passage answers the question. Four failure modes recur, and none of them are fixed by a better vector database.

Low-dimensional collapse. Reimers and Gurevych followed up their own model with a paper showing that dense retrieval degrades as index size grows: the low-dimensional space simply cannot separate a very large number of documents, and the problem worsens with dimension (Reimers and Gurevych, "The Curse of Dense Low-Dimensional Information Retrieval"). The 384-dimensional vectors that are convenient to store are precisely the ones most exposed. A 100,000-document index and a 50-million-document index are different engineering problems even if the code is identical.

Semantic drift toward the generic. Embeddings cluster around frequent topics. Ask for "the refund window for annual enterprise plans" and the nearest neighbors are all about refunds generally, once you get past the handful of exact matches. Lexical signals β€” rare terms, identifiers, error codes, SKUs β€” are where dense models are weakest and where BM25 is still unbeaten. This is why hybrid retrieval won: run both, fuse with reciprocal rank fusion or a weighted score, and let a reranker sort the shortlist. SPLADE made the learned-sparse variant practical by combining term expansion with inverted-index efficiency (Formal et al.).

Embedding-space shortcuts. A model trained contrastively learns whatever correlates with the training pairs. If your corpus has boilerplate headers above the content, the encoder will happily key on them. GPL showed that you can fix a lot of this cheaply with domain adaptation: generate synthetic queries from your own documents with a language model, label them, and fine-tune the encoder on that (Wang et al., "GPL"). An afternoon of fine-tuning on a few thousand synthetic pairs routinely beats upgrading to a bigger general model.

Stale assumptions about what the user typed. Queries got shorter over the past two years β€” half the traffic in systems I have looked at is now pasted code, an error string, or a sentence an assistant wrote. Each of those has a different distribution than the question-shaped queries these models were trained on, and each is worth its own evaluation slice.

The parts that break at scale

An exact nearest-neighbor scan over tens of millions of 768-dimensional vectors is hundreds of billions of floating-point operations per query. That is why every production index is approximate, and why the approximate index is where the interesting failures live. HNSW β€” the hierarchical navigable small-world graph β€” is the default for good reason: it gives logarithmic-ish search with excellent recall, at the cost of a graph structure that is expensive to build and awkward to update (Malkov and Yashunin). FAISS established the GPU-accelerated, billion-scale toolkit that most of the ecosystem still borrows from (Johnson et al.).

The operational trap is that recall is a tunable you must actually tune. An HNSW build with a small graph degree and an aggressive search parameter can lose 10–20% of true neighbors, which looks exactly like a retrieval-quality problem and gets debugged as one. Quantization introduces the same class of silent loss: product quantization and its successors compress vectors by an order of magnitude, and the compressed distance is monotone in the true distance but not identical to it. Two disciplines catch this: keep a frozen query set with known relevant documents, and measure recall@k of the index separately from recall@k of the pipeline. When they diverge, the index is the suspect.

Drift: the failure that arrives without a commit

Software breaks when you change it. Embedding-based systems break when the world changes. Three distinct drifts get conflated, and separating them is most of the diagnostic work.

Corpus drift. The documents change while the model stays fixed. A new product line, a migrated documentation platform, a batch import of scanned PDFs β€” any of these shifts the document distribution so the neighborhoods that used to be meaningful no longer are. Query drift. Users change what they ask and how. A feature that changes the entry point into search can move the query distribution overnight. Model drift is the one you cause yourself: you upgrade the encoder, and every stored vector is now in a different coordinate system. There is no incremental fix. Re-embedding the whole corpus is the only correct move, and the reason embedding-model choice is a multi-year commitment rather than a config value.

The monitoring that catches this is unglamorous. Log the query distribution's aggregate statistics (median length, top terms, the fraction that hit the index at all). Track the fraction of queries whose top result comes from a document older than the query (Webb et al. formalize concept drift detection generally; the retrieval-specific version is mostly plumbing you have to write). Track "zero-result, low-max-similarity" rate as a first-class alert: a rising floor of weak matches means the corpus moved. Above all, keep a golden set of a few hundred real query-relevance judgments. It is a day of annotation, it never rots entirely, and it is the only instrument that can tell you whether Tuesday's change helped.

Clustering: the supervised work nobody wants to fund

The unadvertised benefit of a good embedding space is that it converts unlabeled data into structure. k-means over embeddings is the fastest way I know to find out what is actually in a support inbox: cluster, sample ten documents per cluster, read the samples, name the clusters. That is a taxonomy for the price of an afternoon, and it is usually more honest than the taxonomy in the ticketing system because it is derived from the text rather than from a form field someone abandoned in 2021. HDBSCAN's density-based clustering is the better default when cluster count is unknown and outliers matter β€” the "-1" bucket is often the most interesting output, because it is where the unusual, expensive, or novel cases live.

Two warnings. Clustering results are only as stable as the embedding model, so re-cluster after a model change and expect the labels to move. And clustering is not classification: cluster IDs are arbitrary, so the moment you wire them into a product feature you have built a dependency on an artifact with no backward compatibility. Materialize the human-readable labels, not the IDs.

What I would put in the plan

An embedding system that survives contact with production has six things in it. A pinned model version with its preprocessing contract documented. A frozen evaluation set with graded relevance judgments, refreshed quarterly from real traffic. Hybrid retrieval with a reranker, because pure dense retrieval is leaving recall on the floor. An index-recall measurement distinct from end-to-end quality. A re-embedding budget and runbook, because you will do it more than once. And an explicit owner, because "the embeddings" is not a system until someone is accountable for them.

None of this is exciting, which is exactly why it is a moat. The teams whose search quietly works are the teams that treated the vector layer as infrastructure β€” versioned it, measured it, and knew what would break first. The model is the easy part; the maintenance is the product.

The boring decisions that decide the outcome

Above the model choice sit five mundane decisions that account for more of the variance in retrieval quality than the encoder does, and they are worth naming because they are usually made by accident.

Chunking. Splitting a document every 500 tokens produces fragments that begin mid-sentence and lose the heading that gave them meaning. Splitting on structure β€” headings, paragraphs, table boundaries β€” produces fewer, larger, more coherent units, and coherence is what an embedding measures. My working rule is to make the chunk a unit a human would recognize as self-contained, add the heading path to the text before embedding, and accept the uneven sizes. Fixed-size chunking exists because it is easy to implement in a vector store, not because it is good.

Filtering before versus after. Metadata filters (tenant, date, permission, document type) applied after the vector search produce a perverse effect: the model returns the nearest k neighbors globally, then you discard most of them, and recall drops silently for exactly the queries where the filter is most restrictive. Pre-filtering with an index that supports it is the correct design, and it is the most common reason a search system that works in staging collapses for one tenant in production.

Asymmetry. Queries and documents are not the same kind of text, and an encoder trained on a symmetric similarity task will do worse than one trained on query-to-document pairs. If your model card mentions an instruction prefix, that prefix exists to encode the asymmetry. Dropping it costs recall quietly.

Re-embedding cost. Every vector you store is a promise that you know how to recompute it. For a corpus of a few million documents, a full re-embedding pass is a budgeted, multi-hour, rate-limited operation with a failure mode β€” partial re-embedding β€” that produces a subtly broken index. Write the runbook, including how you detect a half-finished pass, before you need it.

The test harness. Twenty queries with known relevant documents, run on every index or model change, telling you recall@5 for the retriever and for the whole pipeline separately. This is the same argument as building your own evaluation set, applied to the retrieval layer, and it is the difference between iterating on a system and re-rolling dice.

Works Cited

Formal, Thibault, et al. "SPLADE v2: Sparse Lexical and Expansion Model for Information Retrieval." arXiv, 2021, arxiv.org/abs/2109.10086. Accessed 5 Sept. 2024.

Gao, Tianyu, Xingcheng Yao, and Danqi Chen. "SimCSE: Simple Contrastive Learning of Sentence Embeddings." arXiv, 2021, arxiv.org/abs/2104.08821. Accessed 5 Sept. 2024.

Johnson, Jeff, Matthijs Douze, and HervΓ© JΓ©gou. "Billion-Scale Similarity Search with GPUs." arXiv, 2017, arxiv.org/abs/1702.08734. Accessed 5 Sept. 2024.

Kusupati, Aditya, et al. "Matryoshka Representation Learning." arXiv, 2022, arxiv.org/abs/2205.13147. Accessed 5 Sept. 2024.

Malkov, Yury A., and Dmitry A. Yashunin. "Efficient and Robust Approximate Nearest Neighbor Search Using Hierarchical Navigable Small World Graphs." arXiv, 2016, arxiv.org/abs/1603.09320. Accessed 5 Sept. 2024.

Neelakantan, Arvind, et al. "Text and Code Embeddings by Contrastive Pre-Training." arXiv, 2022, arxiv.org/abs/2201.10005. Accessed 5 Sept. 2024.

Reimers, Nils, and Iryna Gurevych. "Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks." arXiv, 2019, arxiv.org/abs/1908.10084. Accessed 5 Sept. 2024.

Reimers, Nils, and Iryna Gurevych. "The Curse of Dense Low-Dimensional Information Retrieval for Large Index Sizes." arXiv, 2020, arxiv.org/abs/2012.14210. Accessed 5 Sept. 2024.

Thakur, Nandan, et al. "BEIR: A Heterogeneous Benchmark for Zero-Shot Evaluation of Information Retrieval Models." arXiv, 2021, arxiv.org/abs/2104.08663. Accessed 5 Sept. 2024.

Wang, Liang, et al. "Text Embeddings by Weakly-Supervised Contrastive Pre-Training." arXiv, 2022, arxiv.org/abs/2212.03533. Accessed 5 Sept. 2024.

Wang, Kexin, Nandan Thakur, Nils Reimers, and Iryna Gurevych. "GPL: Generative Pseudo Labeling for Unsupervised Domain Adaptation of Dense Retrieval." arXiv, 2021, arxiv.org/abs/2112.07577. Accessed 5 Sept. 2024.

Webb, Geoffrey I., et al. "Characterizing Concept Drift." Data Mining and Knowledge Discovery, 2016, arxiv.org/abs/1511.04353. Accessed 5 Sept. 2024.