AI Frontiers, part 52: Long-context retrieval — position, recency, and what gets lost
Part 52from the AI Frontiers series · 65 parts in all
Every time a model ships a bigger context window, someone announces that retrieval is dead. The reasoning is seductive: if the model can read a hundred thousand tokens at once, why build a retriever, an index, a reranker, and a freshness pipeline? Just put the documents in the prompt. This has now been predicted at every window size from 4K through a million tokens (part 7 covered the early rounds), and it has been wrong every time — not because long context is a gimmick, but because reading capacity and retrieval quality are different problems that happen to share a resource.
By mid-2025 the picture is clearer than it was, and it is more interesting than either "retrieval is dead" or "retrieval solves everything." Long context changed what a retriever should return, changed where in the prompt information should go, and made a set of previously theoretical failure modes — positional bias, middle-of-input neglect, distractor sensitivity — into the everyday concerns of anyone tuning a RAG pipeline. This entry is about those failure modes, why they exist, and the architecture that works.
What long context actually costs
Two costs scale differently from what people assume, and both are documented elsewhere in this series but bear repeating here because they explain the architecture. Prefill compute grows roughly quadratically with input length for a dense attention pass, which is why long inputs are slow to start answering and why prompt caching (part 50) is so valuable for repeated prefixes. And the KV cache grows linearly with the tokens held, consuming memory for the entire generation (part 19). A hundred-thousand-token conversation is not a free operation that happens to be long; it is a different resource profile.
This is the first reason retrieval survives: it is a selection mechanism, and selection is how you spend a fixed budget well. Sending everything is a strategy only when everything fits comfortably and the model can find the needle. Neither condition is usually true.
The lost-in-the-middle result, and why it endures
The paper that framed this whole discussion showed that language models retrieve information from the beginning and the end of a long input much more reliably than from the middle, and that performance can fall below the closed-book baseline when the relevant document is buried in the center of a long context (Liu et al., "Lost in the Middle"). The effect is robust across model families and has survived several generations of scaling. It has an intuitive explanation in terms of how models are trained — most pretraining documents place salient information early, and positional encodings combined with causal attention give the extremes outsized influence — but the important thing for engineering is that the curve is real and roughly U-shaped.
Follow-up work sharpened the diagnosis in a way that is more actionable than the original result. The Michelangelo study evaluated long-context models on synthetic tasks designed to isolate specific capabilities and found that the models are poor at reasoning over multiple pieces of information distributed through a long input — the failure is not merely finding a fact, but relating two or three facts that are far apart (Vodrahalli et al.). That distinction matters enormously for how you design a pipeline: a single-fact lookup in a long context works surprisingly well, and multi-hop synthesis over a long context does not, no matter how large the window.
There is also a training-side explanation that produces a concrete engineering instruction. Models are typically trained on packed sequences created by concatenating documents and truncating at the boundary; a model whose whole experience of long inputs is truncated documents never learns that salient information can appear anywhere, and it behaves accordingly. Training with fewer truncations and more coherent long documents measurably improves utilization of the middle of the window (Ding et al.), and approaches like information-intensive fine-tuning explicitly rebalance attention across the whole input (An et al.). The practical corollary for a RAG builder is uncomfortable: the model's middle-neglect is partly an artifact of how it was trained, which means it varies by model and by model version, which means the placement strategy that worked last quarter needs re-verifying.
How windows actually got big
The engineering behind context extension is a small, well-understood body of work, and knowing it prevents several categories of production surprise.
Rotary position embeddings use a rotation whose angle is a function of the position index, which gives the model a continuous positional signal but does not extrapolate gracefully beyond the lengths seen in training (Su et al.). Position interpolation addressed this by squashing position indices into the trained range, effectively trading position resolution for length (Chen et al., "Extending Context Window via Positional Interpolation"). YaRN improved the scheme with frequency-dependent scaling and a small amount of continued training, achieving strong extension with modest compute (Peng et al.), and LongRoPE pushed the idea toward very large windows with an evolutionary search over scaling factors (Jiang et al.).
The critical detail for anyone using a long-context model is that effective context length is a learned property, not a configuration parameter. A model advertised at 128K tokens has different reliability at 16K, 64K, and 128K, and the degradation curve is task-dependent. The RULER benchmark was built precisely to measure this, with synthetic tasks of controlled length and complexity, and its finding is the one to internalize: almost every model's effective context is substantially shorter than its nominal window, and the gap widens as tasks require more reasoning per fact (Hsieh et al.). BABILong made the analogous point using tasks requiring reasoning over chains of facts embedded in long noise (Kuratov et al.). When someone quotes a context window as a capability, the honest response is to ask for the effective-length curve.
Which context benchmarks mean anything
Long-context evaluation had a credibility problem early on: several popular measures could be gamed by not reading the context at all, because the answer was inferable from the question or recoverable from the model's parametric memory. LongBench aggregated a diverse set of long-context tasks and became the standard reference (Zhang et al.), with LongBench v2 adding harder, more realistic tasks and human-validated difficulty (Bai et al.). RULER and BABILong contribute something more valuable than a score: controlled difficulty, so you can see where a model breaks rather than only that it did.
The engineering instruction is to build your own version of the controlled sweep. Take twenty real questions from your domain, place the supporting evidence at five different positions in prompts of three different lengths, add plausible distractors, and measure answer accuracy across the grid. It is an afternoon of work and it will tell you more about your effective context than any leaderboard. Distractor sensitivity in particular is under-measured: models that answer correctly in a clean context frequently fail when ten topically similar but irrelevant documents are present, which is the normal condition of a real corpus.
The pattern that works
The architecture that has survived all of this is not "retrieval instead of long context" or the reverse. It is coarse retrieval, then generous long context — sometimes labeled long-context RAG. Instead of retrieving the twenty most similar 300-token chunks, retrieve a handful of whole documents or large sections and let the model handle the reading. LongRAG formalized this and found that it improves performance on long-context question answering while reducing the number of retrieval units the system has to manage, because the model's reading capacity absorbs the precision that the retriever was being asked to provide (Li et al., "LongRAG"). The complementary result is that retrieval and long context compose rather than substitute: combining a modest retriever with a long window outperformed either alone on mixed question sets (Xu et al., "Retrieval Meets Long Context").
Four practices follow from the failure modes above and are worth spelling out.
Place the highest-value evidence at the edges. Given the U-shaped attention curve, the single most relevant passage belongs at the beginning or the very end of the context, not the middle. This is free, and it is the most consistently effective trick in the whole pipeline.
Never fill the window because it is available. Accuracy does not increase monotonically with context length once distractors enter. The right test is a curve, not a ceiling: sweep context sizes and pick the point where quality flattens.
Design for single-fact lookups at long length; design multi-hop tasks differently. If a question requires synthesizing facts that are far apart in the input, explicitly retrieve and surface each constituent fact near each other, or decompose the question into separate calls. Do not rely on the model to bridge a long input.
Make structure do work. Headings, delimiters, and per-document source labels all reduce the ambiguity of a long prompt; a well-labeled context is a structured context, and structured prompts measurably improve the model's ability to cite and relate. Pair this with requiring citations back to source IDs, which converts an unverifiable long-context answer into a checkable one.
The window will keep growing, and that is a good thing — it makes the coarse-retrieval pattern viable in more places. But the window is a resource, not a strategy. The question that never changes is which tokens deserve to be there, and that question is exactly the one retrieval was invented to answer.
Memory is not the same problem as context
Everything above concerns a single request: how to assemble one prompt well. Production systems also need the other thing, which is state that survives across requests — what the user asked last week, which documents were already reviewed, what the system already tried and failed. A long window does not solve this, because it is empty at the start of every call and because stuffing that history into the prompt is exactly the anti-pattern the middle-neglect result warns about.
The research framing that took hold treats memory as a management problem rather than a storage problem, modeled explicitly on operating systems: a small fast context as the working set, a large external store as the backing store, and an agent that pages information in and out and decides what to evict (Packer et al., "MemGPT"). Earlier work on memory streams added the complementary idea that retrieval over a remembered history should weight recency, importance, and relevance rather than similarity alone (Park et al.), and MemoryBank demonstrated the summarization-on-eviction pattern — compress old interactions rather than keeping them verbatim (Zhong et al.).
What that means in an application, minus the operating-system metaphor: keep a structured profile of durable facts rather than a transcript, summarize past sessions with the model and store the summaries alongside their sources, retrieve memory using the same hybrid machinery you built for documents, and put an explicit budget on how many tokens of history a request may carry. The failure mode to avoid is the one that is easiest to build: appending every turn to the prompt until the window fills, which produces a system whose cost grows with conversation length and whose accuracy degrades at the same rate.
Where this leaves retrieval
Read in sequence, the retrieval entries in this series describe a job that has changed twice. Part 5 was about fetching the passage that answers the question. Part 6 was about the storage layer underneath it. Part 20 described the shift to retrieval as a tool an agent decides when to call. This entry adds the last piece: the retriever's output is no longer a passage, it is a context assembly — a curated working set with a known position profile, a known token cost, and citations that let the model's use of it be audited.
That reframing is what makes the long-context era good news for retrieval rather than bad. When the window was small, the retriever had to be a precision instrument and its errors were unrecoverable. When the window is generous, the retriever can afford to be coarse, and the model can absorb the imprecision — provided the tokens are well chosen, well ordered, and well labeled. The selection problem did not disappear when the window grew. It got a bigger budget and a higher standard.
Works Cited
An, Chenxin, et al. "Make Your LLM Fully Utilize the Context." arXiv, 2024, arxiv.org/abs/2404.16811. Accessed 5 June 2025.
Bai, Yushi, et al. "LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-Context Multitasks." arXiv, 2024, arxiv.org/abs/2412.15204. Accessed 5 June 2025.
Chen, Shouyuan, et al. "Extending Context Window of Large Language Models via Positional Interpolation." arXiv, 2023, arxiv.org/abs/2306.15595. Accessed 5 June 2025.
Ding, Hantian, et al. "Fewer Truncations Improve Language Modeling." arXiv, 2024, arxiv.org/abs/2404.10830. Accessed 5 June 2025.
Hsieh, Cheng-Ping, et al. "RULER: What's the Real Context Size of Your Long-Context Language Models?" arXiv, 2024, arxiv.org/abs/2404.06654. Accessed 5 June 2025.
Jiang, Huiqiang, et al. "LongRoPE: Extending LLM Context Window beyond 2 Million Tokens." arXiv, 2024, arxiv.org/abs/2402.13753. Accessed 5 June 2025.
Kuratov, Yuri, et al. "BABILong: Testing the Limits of LLMs with Long Context Reasoning-in-a-Haystack." arXiv, 2024, arxiv.org/abs/2406.10149. Accessed 5 June 2025.
Li, Zhuowan, et al. "LongRAG: Enhancing Retrieval-Augmented Generation with Long-Context LLMs." arXiv, 2024, arxiv.org/abs/2410.05983. Accessed 5 June 2025.
Liu, Nelson F., et al. "Lost in the Middle: How Language Models Use Long Contexts." arXiv, 2023, arxiv.org/abs/2307.03172. Accessed 5 June 2025.
Packer, Charles, et al. "MemGPT: Towards LLMs as Operating Systems." arXiv, 2023, arxiv.org/abs/2310.08560. Accessed 5 June 2025.
Park, Joon Sung, et al. "Generative Agents: Interactive Simulacra of Human Behavior." arXiv, 2023, arxiv.org/abs/2304.03442. Accessed 5 June 2025.
Peng, Bowen, et al. "YaRN: Efficient Context Window Extension of Large Language Models." arXiv, 2023, arxiv.org/abs/2309.00071. Accessed 5 June 2025.
Su, Jianlin, et al. "RoFormer: Enhanced Transformer with Rotary Position Embedding." arXiv, 2021, arxiv.org/abs/2104.09864. Accessed 5 June 2025.
Vodrahalli, Kiran, et al. "Michelangelo: A Natively Multimodal Foundation Model for Long-Context Understanding." arXiv, 2024, arxiv.org/abs/2409.12640. Accessed 5 June 2025.
Xu, Peng, et al. "Retrieval Meets Long Context Large Language Models." arXiv, 2023, arxiv.org/abs/2310.03025. Accessed 5 June 2025.
Zhang, Jiajie, et al. "LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding." arXiv, 2023, arxiv.org/abs/2308.14508. Accessed 5 June 2025.
Zhong, Wanjun, et al. "MemoryBank: Enhancing Large Language Models with Long-Term Memory." arXiv, 2023, arxiv.org/abs/2305.10250. Accessed 5 June 2025.