Trust is earned, not given

A different perspective

2023-09-21 · Projects

AI Frontiers, part 5: Retrieval-augmented generation before it was everywhere

Part 5from the AI Frontiers series · 65 parts in all

Of all the patterns this series will cover, retrieval-augmented generation has the highest ratio of practical importance to paper-page count. The core idea fits in one sentence — fetch relevant text, paste it into the prompt, let the model answer from it — and by late 2023 that sentence had become the default architecture for every enterprise that wanted an LLM to know things the model was never trained on. Like most "sudden" ideas, it was not sudden. This entry traces where RAG came from, why it beat the obvious alternative, and what the first year of production deployments got wrong.

Before the acronym: three separate threads

Thread one: open-domain question answering. DrQA, from Facebook Research in 2017, answered Wikipedia questions by retrieving paragraphs with bigram hashing plus TF-IDF and reading them with a trained neural reader — retrieval and reading as two cooperating modules, no end-to-end training of the retriever (Chen et al.). Thread two: dense retrieval. DPR, in 2020, replaced TF-IDF with dual BERT encoders trained so that relevant question-passage pairs land close in vector space, outperforming BM25 with orders-of-magnitude fewer dimensions of engineering (Karpukhin et al.). Thread three: the generative turn. When GPT-3 demonstrated in-context learning in 2020 — the model using information placed in the prompt, zero-shot (Brown et al.) — the missing piece fell into place: if a model can use retrieved text presented in its context, the "reader" needs no training at all.

The RAG paper proper — Lewis et al., 2020, also Facebook — combined the threads: a DPR retriever and a BART generator, trained end-to-end so that the retriever learns to fetch what the generator can use, with the marginal likelihood marginalizing over retrieved documents (Lewis et al.). Reading it in 2023, the surprise is how much of the production playbook already appears: a non-parametric memory you can update without retraining, cited provenance for outputs, and the explicit framing of retrieval as compensation for what weights alone cannot carry. The papers that followed — FiD's fusion-in-decoder scoring passages independently then jointly in the decoder (Izacard and Grave), RETRO's trillion-token retrieval database at training time (Borgeaud et al.) — pushed scale and integration, but the 2020 architecture is recognizable inside nearly every 2023 deployment.

Why it beat the alternative

The obvious competing idea was: teach the model your documents with fine-tuning. Part 3 explained why that fails structurally — LoRA teaches behavior, not facts; and full fine-tuning is an expensive way to encourage a model to parrot rather than understand. But the deeper argument for retrieval is operational, and it is the one that won the enterprise argument. Knowledge in weights is frozen at training time, mixed irreversibly with everything else the model knows, versionless, and un-attributable. Knowledge in a retrieval corpus is a database: versioned, permissioned, auditable, correctable in minutes, and deletable on demand — which matters in ways that a single EU court ruling about a "right to be forgotten" applied to model weights makes vivid. When a model must cite its sources, retrieval gives you the citation for free; fine-tuning cannot even formulate the promise.

Cost arithmetic finished the argument. Chunk, embed, and index a few hundred thousand documents once — dollars, not thousands — and swap the model underneath without redoing the knowledge work. Every alternative (fine-tuning, or the long-context designs discussed later in this series) re-prices that knowledge work onto the model side, where it compounds with every model change. In 2023, roughly every enterprise LLM proof-of-concept that survived its first demo did so by adding retrieval; the pattern was so universal that "LLM app" and "RAG pipeline" were nearly co-extensive by autumn.

What the first production year got wrong

Now the useful part: what broke. The 2023 deployments failed in a small number of stereotyped ways, and if you are building one of these, each failure mode below is worth an hour of design attention before the first demo.

Chunking is the whole game, and everyone did it badly. Fixed 512-token chunks with naive overlap cut paragraphs mid-argument, orphaned tables from their headers, and split procedures from their preconditions. Retrieval then faithfully surfaced the wrong units. The teams that did well treated chunking as document-structure parsing — section-aware splits, table extraction, heading metadata — not as a tokenizer parameter.

The retriever was blamed for the generator's sins, and vice versa. Two distinct failure classes hide under "the bot is wrong": retrieved-nothing-relevant (a retrieval problem — fix embeddings, chunking, or the query) and retrieved-the-right-thing-but-misread-it (a generation problem — fix prompting, or accept it). Without logging retrieved passages alongside answers, teams spent weeks tuning the wrong component. The single highest-leverage debugging artifact of 2023 was a UI panel showing what was retrieved for each answer.

Embeddings aged like milk. Models released in 2022 were superseded within quarters; re-embedding a corpus meant re-encoding everything, and index compatibility became a migration problem. Prudent teams stored raw text alongside vectors precisely to make re-embedding survivable.

And the deepest confusion: top-k is not truth. Retrieval guarantees similarity, not correctness. A confident wrong passage retrieves perfectly against a confident wrong query — garbage in, fluent garbage out. The mitigation that actually worked was verification: force the model to quote its sources, check the quote exists, grade the quote-answer entailment. It is the same muscle as the cross-checking habit — and it is why the production systems that earned trust in 2023 were, without exception, the ones that showed their sources.

The architecture that actually shipped

Set the failure modes aside and look at what the successful 2023 deployments converged on, because the consensus design is worth writing down. An ingestion pipeline: parse documents structure-aware, chunk by section with metadata attached (source, date, ACL, document type), embed with a current model, store vectors beside the raw text — never vectors alone. A query path: rewrite the user's raw question into a retrieval query (often with the LLM itself — the first mainstream use of the model-on-model pattern), retrieve top-k with hybrid dense plus keyword fusion, optionally rerank the candidates with a cross-encoder, assemble the prompt with quoted, attributed passages, and generate with instructions to cite. A feedback loop: log retrieved passages next to answers, capture thumbs-down with the retrieval state, and review weekly. None of this is glamorous; all of it is why two teams with the same model and the same corpus got wildly different systems.

The reranking step deserves its own sentence, because it was 2023's quiet quality unlock. Bi-encoder embeddings retrieve fast but coarsely — query and passage are embedded independently, so their vectors must compress everything into one fixed representation. A cross-encoder reranker reads query and passage together and scores relevance directly, which is slower per candidate but dramatically more accurate — so the standard architecture retrieves 100 candidates cheaply and spends a small model's attention scoring the top of them. The pattern is generic and worth memorizing: cheap recall, expensive precision, in that order. The same shape recurs in this series later — cheap generation, expensive verification — and it is, I think, the single most transferable design idea the RAG era contributed to systems engineering.

Evaluation, or the lack of it

The uncomfortable truth about RAG in 2023: most teams had no systematic way to answer "did our retrieval get better?" The end-to-end metric — does the final answer satisfy the user — conflates retrieval quality, generation quality, and prompt assembly, so teams flew by vibes and demo anecdotes. The evaluation tooling that matured late in the year fixed this by decomposing the pipeline: retrieval metrics (does the gold passage appear in the top-k? at what rank?) measured against a labeled or LLM-graded set; faithfulness metrics asking whether every claim in the answer is entailed by retrieved passages — which attacks hallucination directly rather than hoping fluency correlates with truth; and answer relevance scored independently. The decomposition framing — context precision, context recall, faithfulness, answer relevance — became the de-facto four-way scoreboard (Es et al.). None of it replaces your own task-specific eval set; all of it makes the weekly pipeline regression-testable instead of aesthetic.

There is a deeper point hiding in the tooling, and it connects to part 2: RAG systems are systems, with components whose failure modes compose. A 2023 survey of production failures found most bad answers traced to retrieval misses, not model hallucination — the opposite of what teams assumed. Measurement is how you learn which layer to fix. This is the LLM-era version of the oldest lesson in operations engineering: when something breaks, the interesting question is never "is the software bad" but "which layer, measured how."

The day the demo dies: a field taxonomy of retrieval failures

Because this series is written from the building side, it is worth cataloguing the failure taxonomy that 2023 produced — the five ways a RAG pipeline actually breaks, as opposed to the way it breaks in the architecture diagram. One: the fact is not in the corpus. The model answers from parameters or declines; teams discover this when the demo asks about something real. No retrieval stack fixes a missing document; ingestion coverage is a requirement, not a phase-two. Two: the fact is in the corpus but not in any chunk's neighborhood — split across a page boundary, buried in a table's footnote, phrased in the complement ("the refund does not apply when"). Structure-aware chunking and claims-level extraction are the partial answers. Three: the chunk is retrieved and the model ignores it — the lost-in-the-middle pattern, or a prompt whose instructions drown the evidence. Four: the chunk is stale — the corpus says last year's policy, and the model confidently narrates history. Freshness is an ingestion-pipeline SLA, never a model property. Five: the answer requires synthesizing across chunks — comparing three documents, totaling five line items — and top-k passage lists are the wrong shape for the job. Multi-hop retrieval, agentic loops, or structured tool calls are the 2024 answers; in 2023 the honest answer was "don't promise this yet."

The taxonomy matters because each failure has a different owner and a different fix, and because every one of them produces the same user-visible symptom: a confident, fluent, wrong answer. Fluency is a red herring in evaluation — it correlates with the pre-training corpus, not with your pipeline's correctness. The only evaluation that tracks reality is decomposed: retrieval hit rate, citation faithfulness, answer relevance, freshness lag, each with its own dashboard and its own owner. The teams that built that scoreboard in 2023 are the ones whose systems you now trust; the correlation was that clean.

The permission problem nobody demos

Here is the question that separates a chatbot prototype from a system you can deploy inside a company: when the finance team's assistant retrieves a passage, does it respect the fact that the passage's source document is confidential? In 2023 this became known as the permission problem, and it is the reason RAG deployments have security reviews while fine-tuning deployments do not. The weights model leaks nothing it was not trained on; a retrieval system is a search index over everything you fed it, and search indexes over corporate content inherit corporate access control — or fail audits.

The design space is small and worth memorizing. Index-per-audience: separate collections per team or clearance level, simplest to reason about, costly at high tenant counts. Filter-first retrieval: one index, ACL metadata on every chunk, and the filter applied before the ANN search — never after, because post-filtering a top-k that already leaked headlines is a lawsuit with extra steps. Identity-scoped queries: pass the user's entitlements into the retrieval call as a first-class parameter, testable per user. And the subtle one that bit several public deployments: summarization and generation can only see what retrieval fetched, so a correctly-filtered retrieval layer makes the whole system access-safe — which means the filter is the security boundary, and it belongs in the data path, not in the prompt. "Please only use documents Bob may see" in the system prompt is not a security control; it is a suggestion.

The 2023 state of practice, honestly stated: most pilots ignored all of this and got away with it because the corpus was public documentation. The deployments that graduated to production discovered that access control, not model quality, was the long pole — and that the vector database's metadata-filtering features (part 6) existed precisely for this requirement. The lesson generalizes across this entire series: the LLM is the easy 20% of the system; the boring enterprise requirements are the 80% that determine whether anything ships.

Where this sits in the series

RAG's 2020 paper and 2023 production reality bracket a claim I will keep returning to: capability that lives in systems beats capability that lives in weights for anything that must be maintained. The next two parts follow that thread into infrastructure — vector databases, the context-window economics that shape how much retrieval you can afford, and why the 128K-context releases of late 2023 changed the design space without retiring it.

Works Cited

Borgeaud, Sebastian, et al. "Improving Language Models by Retrieving from Trillions of Tokens." Proceedings of the 39th International Conference on Machine Learning, PMLR, 2022, proceedings.mlr.press/v162/borgeaud22a.html. Accessed 21 Sept. 2023.

Brown, Tom B., et al. "Language Models are Few-Shot Learners." Advances in Neural Information Processing Systems 33, 2020, proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html. Accessed 21 Sept. 2023.

"TruLens: Evaluation for LLM and RAG Applications." TruEra Open Source, 2023, github.com/truera/trulens. Accessed 21 Sept. 2023.

Chen, Danqi, et al. "Reading Wikipedia to Answer Open-Domain Questions." Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL, 2017, aclanthology.org/P17-1171. Accessed 21 Sept. 2023.

Izacard, Gautier, and Edouard Grave. "Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering." Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, ACL, 2021, aclanthology.org/2021.eacl-main.116. Accessed 21 Sept. 2023.

Karpukhin, Vladimir, et al. "Dense Passage Retrieval for Open-Domain Question Answering." Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, ACL, 2020, aclanthology.org/2020.emnlp-main.550. Accessed 21 Sept. 2023.

Lewis, Patrick, et al. "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." Advances in Neural Information Processing Systems 33, 2020, proceedings.neurips.cc/paper/2020/hash/6b49323088f684c8b4b1b9b4c0eaf4d5-Abstract.html. Accessed 21 Sept. 2023.