Trust is earned, not given

A different perspective

2024-01-18 · Projects

AI Frontiers, part 7: The context window wars β€” from 4K to 200K (and the road to a million)

Part 7from the AI Frontiers series · 65 parts in all

GPT-2 could hold about a page in its head. GPT-3, three pages. GPT-3.5-turbo, twenty. Claude 2 β€” nineteen months after GPT-3.5 β€” held an entire novella. Then GPT-4 Turbo shipped in November 2023 with a 128,000-token window, and Anthropic answered with 200K. Multiply the 2020 number by 100 and you have the state of the art as I write this in January 2024. No other dimension of model capability moved this fast. This entry is about what made the growth possible, what it costs, and what it does β€” and does not β€” obsolete.

Why windows were small: the two walls

Two hard walls capped early context lengths. Wall one is architectural, and part 1 planted the flag: self-attention compares every pair of positions, so cost grows with the square of sequence length. Quadruple the window, sixteen times the attention compute and memory. Wall two is positional: the original transformer told the model where each token sits with fixed sinusoidal encodings, and β€” as predicted β€” they did not extrapolate past training length. A model trained at 4K literally had no notion of position 100,001. Most position-representation research of 2021–2023 is the story of knocking down wall two so that scaling engineering could attack wall one.

Transformer-XL introduced relative encodings β€” positions expressed as offsets between tokens, which generalize across lengths (Dai et al.). T5 made biases learned per attention-bucket (Raffel et al.). ALiBi discarded learned positions entirely, penalizing attention in proportion to distance with a fixed slope β€” and trained at 1K attended gracefully to 10K (Press, Smith, and Lewis). RoPE β€” rotary embeddings, rotating query and key vectors by position-dependent angles so relative offsets emerge from dot products β€” became the open-weights default through LLaMA (Su et al.). Two more inventions made wall one climbable: FlashAttention's exact-but-I/O-aware kernels cut activation memory from quadratic to linear in sequence length for the attention matrix (Dao et al.) β€” the single most consequential systems paper of the era, in my estimation β€” and grouped-query attention made the inference-time KV cache affordable, which Llama 2 70B shipped with (part 4).

Extension: how 4K models became 128K models

The 2023 trick that unlocked everything was position-interpolation: rather than asking a 4K-trained model to handle positions it never saw, rescale the positions so 32K fits into the 4K range the model already understands β€” compress the ruler, not the model. Meta's paper showed the rescaled model recovers long-context capability witha trivial amount of continued fine-tuning (Chen et al.). RoPE variants went further: NTK-aware scaling adjusts the interpolation non-uniformly across frequency components, and "YaRN" refined the recipe to where 128Γ— extension needed only hundreds of fine-tuning steps (Peng et al.). CodeLlama shipped 16K contexts out of the box with RoPE scaling, and Mistral 7B (September 2023) paired an 8K window with sliding-window attention β€” bounded attention past 8K back, stacked in layers, long-range information carried forward through the residual stream β€” completing the practitioner's toolkit.

By late 2023 the recipes were standard enough that frontier labs shipped 128K (GPT-4 Turbo) and 200K (Claude 2.1) as product features. Claude 2.1's release was also the honest moment of the year: Anthropic reported that recall degraded on very long inputs β€” the "lost in the middle" phenomenon, named months earlier by Stanford researchers who showed that models reliably attend best to the beginning and end of long contexts while mid-document facts get lost (Liu et al.) β€” and set about improving it. The pricing told the same truth: long contexts bill proportionally, because the compute really is quadratic at training and linear-but-huge at inference.

A 2024 design worksheet

Synthesize the entry into the checklist I now apply to any system that touches long context, in the order the questions should be asked. What is the working set? Distinguish what the model must see every call (instructions, schemas, few-shot anchors) from what varies per call (the retrieved evidence, the conversation). The first kind belongs in a cached, stable prompt prefix; the second kind is the only part worth paying fresh attention for. What is the attention budget? Long documents self-servingly believe they deserve whole-context treatment; your job is to decide which fragments actually drive the decision, which is retrieval's whole purpose β€” a 128K window with 120K of padding is a 128K bill for an 8K answer. What is the freshness contract? Long windows invite pasting stale dumps; if the underlying facts move, define the refresh path before launch. What is the failure posture? When the model misses a mid-context fact, does the system retry with reordering, escalate to a bigger window, or fail loudly? Silent quality loss is the worst of the three. And what is the exit? Context sizes, prices, and model versions all change quarterly; encode prompts and retrieval as data, not as strings in code.

Run any real 2023-era long-context project against that list and the gaps fall out in minutes β€” which is, in miniature, this series' whole method: the papers move monthly, the engineering questions move annually, and the disciplines that survive are the ones that turn novelty into checklists.

The first battle: why 32K was hard in 2023

Before the extension tricks, there was a year of failed experiments, and the failures teach more than the successes. The obvious attack on wall one β€” replace quadratic attention with something linear β€” produced a whole family of efficient-attention papers from 2021: Performer's kernel-based approximations, Linformer's low-rank projections, Longformer and BigBird's sparse patterns. All work; none won. The reason, learned the hard way, is that exact attention's failure mode (quadratic cost) is predictable and engineerable β€” throw FlashAttention-class kernels at it β€” while linear approximations' failure mode (recall degradation on precise retrieval tasks) is subtle, task-dependent, and disqualifying for exactly the long-document work the big windows exist to serve. The field's conclusion, now embedded in every frontier model: keep exact attention, attack the constant factors with systems engineering, and solve length via position-encoding surgery plus data. The linear-attention research program was not wasted β€” it re-entered through state-space models (Mamba, December 2023) whose selective-copying results finally made sub-quadratic architecture credible again β€” but it lost the 2023 round decisively.

The data lesson matters as much as the architecture one: long-context capability is curated, not emergent. Models pre-trained at 4K did not develop 32K competence by accident; it took continued pre-training on genuinely long documents with matched objectives β€” which is why the 2023 recipe reads like fine-tuning with a data problem attached, and why datasets of books, repositories, and full-thread conversations became a quiet competitive asset. The labs that treated long context as a data engineering program shipped it; the labs that treated it as a scaling knob shipped demos that degraded at halves.

Needle in a haystack: reading the demos honestly

Every long-context launch since mid-2023 has shipped with the same demo: bury one fact β€” a phrase, a magic number β€” somewhere in a maximal-length document and show the model retrieving it. The demos are real, and they prove much less than they appear to. Retrieval of a planted needle is the easiest long-context task: the needle is distinctive, the query matches it lexically, and the answer requires no synthesis. The failures start one step up the difficulty ladder β€” multi-needle variants (several facts to combine), adversarial needles (facts that contradict context elsewhere), and reasoning over all of the haystack rather than spotting one straw, where mid-context degradation (Liu et al.) compounds with task complexity. A model that finds one needle in 200K tokens may still miss that two of your hundred documents disagree.

The practical consequence for system builders is a debugging discipline: when a long-context call misbehaves, first determine whether the needed facts were present, retrieved into attention, or used β€” three different failures with three different fixes. Present-but-unused usually resolves with prompt restructuring (edges over middle, per the literature); genuinely-absent facts need retrieval, not a bigger window; used-but-misweighted is a reasoning failure that neither window nor retrieval fixes. The worst production incidents I have watched came from conflating the three and shipping a bigger context as the fix for what was actually a retrieval gap.

The inference bill: what a big window costs at runtime

Architecture papers celebrate training, but the context window's real cost lands at inference, and it is worth doing the arithmetic once because it explains half the product decisions of the following year. Every token in the window carries a key-value pair at every layer, and generation re-reads all of it for every new token: the KV cache for a 13B-class model at 128K tokens is on the order of tens of gigabytes in float16 β€” several high-end GPUs' worth of HBM for one user's session if held naively. This is why grouped-query attention (sharing key/value heads across query heads, cutting the cache by the group factor β€” Llama 2 70B shipped it) and quantized KV caches went from obscure optimizations to load-bearing infrastructure during 2023. It is also why providers price long-context calls super-linearly in tokens: the memory is reserved for the session's lifetime, whether or not your next token needs all of it.

Three consequences for system design. First, per-token pricing means context hygiene is a budget line: the 200K window is not an invitation to paste your entire wiki; it is headroom for what retrieval selected. Second, latency: time-to-first-token grows with prompt length (prefill is parallel but not free), so interactive systems cache prefixes aggressively β€” conversation history that never changes should never be re-prefilled β€” the API providers who add prefix caching (and someone will, this year) will reward exactly that structure. Third, quality, per the lost-in-the-middle literature: ordering matters, so put the material you most want reasoned about at the edges of the prompt, and put instructions where they will be re-read β€” the frontier models' instruction following at position zero is better than at position 90K. None of this appears on a spec sheet; all of it appears on your invoice and in your eval scores.

What long context changes β€” and what it merely re-prices

What it changes: document-scale tasks become single-prompt tasks. Contract review across an entire agreement; codebase comprehension across dozens of files; multi-document synthesis where cross-references matter. These were multi-step RAG orchestration problems in March 2023 and became "paste it in" problems by November. For interactive work β€” a session where the model accumulates context over hours β€” big windows are not a convenience but an enabler of a genuinely new interaction pattern.

Meanwhile the frontier's own next move is already visible on the horizon as this entry is written. Google's Gemini launch of December 2023 shipped 32K contexts with a very different compute recipe, and the rumors out of Mountain View β€” like the rumors attached to every lab β€” point toward contexts measured in millions of tokens within the year, most plausibly via mixture-of-experts routing (the subject of part 8) that attends selectively rather than quadratically. If that lands, the next battle shifts from window size to window cost: a million tokens that bill like ten thousand are a different product than a million tokens that bill like a million. The current stable fact remains the 128K/200K tier β€” and the design disciplines below are written for it.

What it re-prices rather than retires: RAG. Every token in the window costs money and attention quality, every call. A 10M-token corpus pasted at 128K context per query costs per-query what a retrieval-substrate approach costs per-month. The 2024 design I keep converging on: retrieval to select, long context to hold, and the "needle in a haystack" demos carefully discounted β€” retrieving one planted needle is the easy case; reasoning over the whole haystack degrades measurably in the middle (Liu et al.). And the KV-cache arithmetic behind all of this β€” why a 200K window costs real RAM per session even with GQA β€” is where this series is headed in part 19, because cache management is quietly becoming the load-bearing infrastructure question of the long-context era.

Works Cited

Chen, Shouyuan, et al. "Extending Context Window of Large Language Models via Positional Interpolation." arXiv.org, 2023, arxiv.org/abs/2306.15595. Accessed 18 Jan. 2024.

Dai, Zihang, et al. "Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context." Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, ACL, 2019, aclanthology.org/P19-1167. Accessed 18 Jan. 2024.

Dao, Tri, et al. "FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness." Advances in Neural Information Processing Systems 35, 2022, proceedings.neurips.cc/paper_files/paper/2022/hash/67d57c32e20fd0da7aab1f214747f2b6-Abstract-Conference.html. Accessed 18 Jan. 2024.

Liu, Nelson F., et al. "Lost in the Middle: How Language Models Use Long Contexts." TMLR, 2023, openreview.net/forum?id=iHv2MscYSC. Accessed 18 Jan. 2024.

Peng, Bowen, et al. "YaRN: Efficient Context Window Extension of Large Language Models." arXiv.org, 2023, arxiv.org/abs/2309.00071. Accessed 18 Jan. 2024.

Press, Ofir, Noah Smith, and Mike Lewis. "Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation." arXiv.org, 2021, arxiv.org/abs/2108.12409. Accessed 18 Jan. 2024.

Raffel, Colin, et al. "Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer." Journal of Machine Learning Research, vol. 21, no. 140, 2020, jmlr.org/papers/v21/20-074.html. Accessed 18 Jan. 2024.

Rozière, Baptiste, et al. "Code Llama: Open Foundation Models for Code." arXiv.org, 2023, arxiv.org/abs/2308.12950. Accessed 18 Jan. 2024.

Su, Jianlin, et al. "RoFormer: Enhanced Transformer with Rotary Position Embedding." arXiv.org, 2021, arxiv.org/abs/2104.09864. Accessed 18 Jan. 2024.