AI Frontiers, part 20: The RAG-to-agentic-RAG evolution
Part 20from the AI Frontiers series · 65 parts in all
In part 5 the retrieval pipeline was three moving parts: embed the query, find the nearest vectors, paste the results into the prompt. That architecture is still the right starting point for most systems, and it is still what a majority of production deployments actually run. But the interesting work moved on, and the direction it moved in is easy to describe and hard to get right: retrieval stopped being a pre-processing step and became a decision the model makes, repeatedly, while it works on a problem.
The forced-single-shot design had a failure mode that everyone hit in 2023 and nobody could fix from inside. A question like "which of our vendors changed their SLA in a way that conflicts with the contract we signed" cannot be answered by one similarity search, because the search terms it would need are not words that appear in either document. The retrieval step had to know the answer to construct the query. Everything that follows is an attempt to break that circle.
The static pipeline and why it plateaus
The baseline deserves a fair statement before it is criticized, because a well-tuned static RAG pipeline is genuinely good at a large class of problems: question answering over documentation, support deflection, codebase search, policy lookup. Chunk the corpus sensibly, embed with a good model, retrieve the top k, rerank with a cross-encoder, put the best passages in the prompt with instructions to cite them. Cache the prefix. That system is cheap, predictable, auditable, and fast, and I would build it before anything more elaborate.
Where it plateaus is on questions with structure. Multi-hop questions that require chaining facts across documents (HotpotQA, Yang et al., remains the canonical formulation). Questions where the query vocabulary differs from the document vocabulary (HyDE, Gao et al., addresses this directly by generating a hypothetical answer and embedding that). Questions that require aggregating over a set rather than finding a passage — "how many of our suppliers are in countries with new data-residency rules" — where no single chunk contains the answer. Questions where the right retrieval strategy depends on what the first retrieval returned.
Measurement lagged the problem badly for a while. RAGAS (Es et al.) was an early attempt to score a pipeline on faithfulness and answer relevance without ground-truth labels, and RGB (Chen et al.) built a benchmark specifically to separate a model's retrieval ability from its noise-robustness. The general finding in both cases was the same and remains underappreciated: most RAG failures are retrieval failures, not generation failures, and teams spend their effort tuning the wrong half.
Making retrieval a decision instead of a step
The research path out of the plateau is a sequence of papers that all share one idea — let the model participate in retrieval — and differ in how aggressively.
Interleaving. IRCoT (Trivedi et al.) alternated a reasoning step with a retrieval step, using each intermediate thought as a query for the next search. FLARE (Jiang et al.) went further and made retrieval conditional: generate ahead, and only when the model's own confidence drops — measured as low probability on the next tokens — pause and fetch supporting documents. This is a genuinely different contract from static RAG: instead of deciding what to retrieve before knowing anything, the system retrieves when it is stuck.
Self-assessment. Self-RAG (Asai et al.) trained a single model to emit reflection tokens: decide whether retrieval is needed at all, critique each retrieved passage for relevance and support, and critique its own final answer for groundedness. That produces a system with the property every compliance department eventually asks for, which is a traceable statement of which retrieved evidence supports which claim.
Query transformation as a first-class step. Rewriting a conversational question into a standalone search query (Ma et al.), decomposing a compound question into sub-queries, and generating multiple paraphrases for recall are all cheap, boring operations that disproportionately improve results. In my experience this is the highest-yield single change available to an existing pipeline, and it requires no new model.
Optimization over the pipeline. DSPy (Khattab et al.) treats prompts and pipeline structure as parameters to be compiled against a metric rather than strings to hand edit. Whether or not you adopt the framework, the framing is valuable: a retrieval pipeline has many interacting knobs — chunk size, k, reranker depth, query rewriting strategy — and tuning them by intuition in production is how teams end up with a system nobody can explain.
Structure, which is the part people skip
Chunks are a lossy representation. A contract clause means something because of the document it sits in; a function means something because of what calls it. Two responses to that loss matter.
The first is context enrichment: before embedding a chunk, prepend a short generated summary of the document it came from, so the chunk's vector carries information about its setting. Anthropic's contextual retrieval work reported large reductions in retrieval failure rate from exactly this technique, which is an unglamorous win with a clear mechanism — the embedding sees more of the meaning.
The second is graph structure. GraphRAG (Edge et al.) builds an entity-and-relation graph over a corpus, assembles community summaries, and answers questions by traversing that graph rather than by nearest-neighbor search. For the aggregation class of question — "what are the themes across these three hundred incident reports" — graph summarization answers questions that flat retrieval structurally cannot, because no passage contains the answer. It is more expensive to build and harder to keep fresh, and for those specific question types it is the only thing that works.
There is a third structural question that is really the important one: what is a unit of retrieval in your domain? A paragraph is the default because it is easy, not because it is right. In code it is a symbol plus its callers. In legal work it is a clause plus its definitions section. In incident data it might be a timeline. Getting this right beats every embedding model upgrade.
The boring knobs that still matter most
Before adding a planner, it is worth exhausting the levers that do not require one, because they are cheaper and easier to reason about.
Chunking is a semantic decision disguised as a parameter. Fixed-size splits with overlap are the default and they slice sentences in half, which damages both the embedding and the passage handed to the model. Splitting on structure — headings, function boundaries, table rows, transcript turns — and keeping a small amount of neighbouring context attached to each chunk improves results more than swapping embedding models. The failure mode to watch for is a chunk that contains the answer but not the question it answers, which is why prepending a short summary of the source document is so effective.
Retrieve more than you use, then rerank. Dense retrieval is a coarse filter; a cross-encoder reranker that scores each candidate against the actual query is far more accurate and far more expensive per document, which is exactly the right shape for a two-stage pipeline. Retrieve fifty, rerank, keep five. This is the highest-yield architectural change available to a static pipeline, and it is also the step most often skipped because the first stage appears to work well enough.
Combine lexical and semantic search. Embeddings are bad at exact strings — part numbers, error codes, proper nouns, rare identifiers — and those are frequently what people search for. BM25 or another sparse retriever handles them almost perfectly at negligible cost. Hybrid retrieval with reciprocal rank fusion is not sophisticated and it fixes a whole class of embarrassing failures, which is why every mature search system has used both signals for decades.
Freshness, which is the real operations problem
Retrieval systems are usually launched against a snapshot and then maintained against a moving target, and the move is where most of the long-run cost hides.
An index has to handle insertions, updates and deletions, and deletion is the one people forget. Removing a vector from an approximate-nearest-neighbour structure is nontrivial, so most implementations tombstone and rebuild, which means a document that was deleted for legal reasons can remain retrievable until the next rebuild unless the pipeline explicitly filters them at query time. That is a compliance problem dressed as an implementation detail.
Re-embedding is the other cost. Changing the embedding model invalidates every vector in the index, so the choice is a migration decision, not a configuration change — which is a strong argument for versioning embeddings, storing them separately from the documents, and being willing to run two indexes side by side during a transition rather than committing to a flag day.
And freshness itself is a product decision. Some corpora must reflect reality within seconds; others are perfectly useful a week stale. Deciding which you have determines whether you need streaming ingestion, incremental indexing and cache invalidation, or whether a nightly job is sufficient. Teams that skip this question end up building the expensive version by accident.
The agentic version, and what it costs
Agentic RAG is what you get when the loop is genuinely open-ended: a planner decomposes the question, retrieval is one of several tools, each attempt is evaluated, failed attempts inform the next query, and the process terminates when a sufficiency condition is met or a budget is exhausted. It is the same architecture as the coding agents of part 14, applied to knowledge instead of code, and it inherits the same economics.
Those economics are the reason to be skeptical of it as a default. A single retrieval call is tens of milliseconds and a few thousand tokens. A planner with a verification loop can become eight to fifteen model calls, a variable number of searches, and a latency measured in tens of seconds. That is fine for a research task where a human waits, and unacceptable for a chat interface where a user expects an answer before they lose interest. The honest framing is that agentic RAG is a capability upgrade purchased with latency and cost, and it should be reserved for questions that static retrieval demonstrably fails on — which you can only know if you have built the evaluation set to tell you.
There is also a failure mode that the agentic literature understates. Verification loops can converge on a plausible wrong answer, because the verifier is the same model that proposed it, and the more iterations it runs the more confidently it can rationalize. Bounding the number of retries and surfacing uncertainty is not a limitation to apologize for; it is the difference between a system that fails visibly and one that fails invisibly.
Long context did not replace retrieval
The obvious challenge to this whole area is that if a model can hold a million tokens, why retrieve at all? The question is fairer than the confident answers on either side.
The clearest empirical treatment compared retrieval-augmented models against long-context models head to head (Li et al.) and found that long context wins on average quality when the corpus genuinely fits and the hardware allows it, while retrieval remains dramatically cheaper and still wins on some tasks. That is consistent with the memory analysis in part 19: holding a corpus in a KV cache costs memory per request, and cost is linear in tokens, so long context is a luxury good that scales badly.
The synthesis I would recommend, and the one the good production systems have converged on, is a hierarchy. Retrieve to narrow a large corpus to a small candidate set, then put the candidates fully in context so the model can attend over them without chunk boundaries getting in the way. Retrieval handles scale; long context handles precision within the retrieved set. Neither alone is sufficient, and the split point between them is an engineering decision you should make by measurement rather than by taste.
What has actually changed since 2023 is not that retrieval got replaced. It is that the retrieval step stopped being a lookup and became part of the reasoning, with all the costs that implies. That is a real advance and it is also a genuinely harder system to operate, which is why the teams doing it best are the ones with a real evaluation set and a working static pipeline to fall back on.
Works Cited
Asai, Akari, et al. "Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection." arXiv, 2023, arxiv.org/abs/2310.11511. Accessed 19 Mar. 2026.
Chen, Jiawei, et al. "Benchmarking Large Language Models in Retrieval-Augmented Generation." arXiv, 2023, arxiv.org/abs/2309.01431. Accessed 19 Mar. 2026.
Edge, Darren, et al. "From Local to Global: A Graph RAG Approach to Query-Focused Summarization." arXiv, 2024, arxiv.org/abs/2404.16130. Accessed 19 Mar. 2026.
Es, Shahul, et al. "RAGAS: Automated Evaluation of Retrieval Augmented Generation." arXiv, 2023, arxiv.org/abs/2309.15217. Accessed 19 Mar. 2026.
Gao, Luyu, et al. "Precise Zero-Shot Dense Retrieval without Relevance Labels." arXiv, 2022, arxiv.org/abs/2212.10496. Accessed 19 Mar. 2026.
Jiang, Zhengbao, et al. "Active Retrieval Augmented Generation." arXiv, 2023, arxiv.org/abs/2305.06983. Accessed 19 Mar. 2026.
Khattab, Omar, et al. "DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines." arXiv, 2023, arxiv.org/abs/2310.03714. Accessed 19 Mar. 2026.
Lewis, Patrick, et al. "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." arXiv, 2020, arxiv.org/abs/2005.11401. Accessed 19 Mar. 2026.
Li, Zhuowan, et al. "Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach." arXiv, 2024, arxiv.org/abs/2407.16833. Accessed 19 Mar. 2026.
Ma, Xinbei, et al. "Query Rewriting for Retrieval-Augmented Large Language Models." arXiv, 2023, arxiv.org/abs/2305.14283. Accessed 19 Mar. 2026.
Trivedi, Harsh, et al. "Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions." arXiv, 2022, arxiv.org/abs/2212.10509. Accessed 19 Mar. 2026.
Yang, Zhilin, et al. "HotpotQA: A Dataset for Diverse, Explainable Multi-Hop Question Answering." arXiv, 2018, arxiv.org/abs/1809.09600. Accessed 19 Mar. 2026.
Yu, Hao, et al. "Evaluation of Retrieval-Augmented Generation: A Survey." arXiv, 2024, arxiv.org/abs/2405.07437. Accessed 19 Mar. 2026.