AI Frontiers, part 2: GPT-4 and the emergent abilities debate
Part 2from the AI Frontiers series · 65 parts in all
In part 1 I argued that the transformer paper's real legacy was clearing the road for a different question: not "what architecture?" but "what happens when you scale?" This entry is about the sharpest formulation of that question to date, and the sharpest recent challenge to it. On March 14, 2023, OpenAI released GPT-4 (OpenAI, "GPT-4 Technical Report"). Six weeks later, a Stanford-centered team posted a paper with a thesis that, if correct, deflates much of what made GPT-4's release feel historic: that "emergent abilities" — the supposed discontinuous leaps in capability that appear only in large models — are a mirage, an artifact of the metrics we use rather than a property of the models (Schaeffer et al.). Between those two publications sits the most practically important argument in applied AI right now. I want to work through both sides, because which side you believe changes how you plan products, budgets, and benchmarks.
What GPT-4 actually claimed
The GPT-4 technical report is, by the standards of machine learning publications, a strange document: ninety-ish pages that read like a product datasheet written by people forbidden to say anything useful. No architecture details, no dataset composition, no parameter count, no training cost. What it does contain is a battery of benchmark results, human-preference comparisons, and a set of stress tests — professional exams mostly — where GPT-4 lands in or near the top decile of human test-takers: the bar exam, the GRE, AP Art History, sommelier examinations. The report is also candid about failures: the model can be confidently wrong, has a limited context window, does not learn from experience, and hallucinates facts — though markedly less than GPT-3.5 on internal factual-evaluation prompts (OpenAI, "GPT-4 Technical Report").
Two things in the report deserve more attention than they got. The first is the claim, made in one quiet paragraph, that a next-token predictor trained at scale develops utility "despite its capabilities appearing to stem from process," and that the report's authors could not fully predict the model's behavior after training — the sentence responsible for half the discourse about AI risk in the following months. The second is the plot showing performance on a MMLU variant in languages from English through Urdu to Telugu, with a curve that declines gracefully rather than collapsing — evidence that the model's capabilities are not an artifact of any single language's data. Whatever else is true, the report documents a model whose breadth broke the benchmarking conventions of the previous three years.
For builders, the release mattered in a more mundane way: it moved the goalposts on what a "prompt" could accomplish. Chain-of-thought prompting — the finding, from a 2022 Google paper, that asking a model to write out intermediate reasoning steps substantially improves arithmetic, commonsense, and symbolic reasoning (Wei et al.) — went from research curiosity to standard practice within weeks of GPT-4's release, because the model was finally reliable enough at instruction-following to make the technique dependable in production. The engineering pattern of early 2023 — decompose the task, ask for intermediate artifacts, validate each one — is chain-of-thought with guardrails, and it runs on capabilities that either emerged or were unlocked in this generation of models, depending on which side of the debate you take.
The mirage argument, stated fairly
The Stanford paper's argument is about measurement, and it is a good argument. "Emergence" had acquired a specific meaning in the 2022 paper that named the phenomenon: abilities are emergent if they are absent in smaller models but present in larger ones, with performance that is flat at random-chance level until some scale threshold, after which it climbs steeply — a discontinuity (Wei et al., "Emergent Abilities"). The mirage paper's counterclaim is that many celebrated discontinuities are produced by non-linear or discontinuous scoring rules. Their canonical example is multi-digit multiplication: graded by exact string match, a 7-digit product is pass/fail, so small models score ~0 and the metric looks like a cliff. Measure instead the per-digit accuracy — a continuous, token-level metric — and the same models show smooth, steady improvement with scale. No cliff, no miracle: just a continuous underlying capability that a binary metric was chopping into a step function (Schaeffer et al.).
They apply the same analysis to chess piece-move legality, again finding smooth gains under a soft metric. And they make a sharper methodological point: the more benchmarks a model is tuned to pass, the more any apparent emergence reflects the benchmark rather than the phenomenon. Their conclusion is not that large models are not better — they plainly are — but that "emergence" as commonly used conflates a property of the world with a property of our rulers. The paper closes with a pointed line about how such claims inflame both hype and alarm, which, given the paper's timing six weeks after GPT-4, reads as a deliberate intervention in the discourse.
The rebuttal, stated fairly
The strongest response — and I find it more convincing than its proponents initially received credit for — is that the mirage argument proves too much. Yes, exact-match multiplication is a discontinuous metric. But some downstream tasks are inherently discontinuous: the code either compiles or it does not; the SQL either returns the right rows or it does not; the JSON either parses or it does not. If a capability only matters to users through a discrete success criterion, then a smooth interior metric is a soft-hearted comfort, not a refutation. The per-digit accuracy of a multiplication model is cold comfort if the final answer is wrong. The authors of the original emergence paper, along with other researchers, made versions of this argument in the weeks that followed, noting that for discontinuous real-world tasks, the "mirage" framing itself is the artifact.
There is a subtler version too. Whether an ability's growth curve looks smooth depends on the resolution of your measurement. Neural network scaling curves studied carefully over several orders of magnitude — the scaling-laws work of the previous three years — are smooth when you fit power laws to aggregate loss, but aggregate loss is precisely the kind of averaged, continuous metric that can hide reallocated capability: a model can be trading skill A for skill B smoothly while any given user-facing skill looks like it arrived suddenly. Loss is continuous; useful competence is not. Both things are true at once, and the debate's heat comes from people arguing about different objects — loss curves versus product-grade tasks.
What the exams actually test — and what they do not
It is worth being precise about what the exam-results framing demonstrates, because both boosters and skeptics talk past it. A bar exam or a GRE section is a multiple-choice test of reading speed, domain vocabulary, and test-taking strategy, with answer selection as the output format. A next-token predictor with strong instruction following is, structurally, a near-ideal machine for that format: the task is a pattern match over written knowledge with a closed set of options. So the exams establish something real — breadth of retained knowledge and coherence across domains, at a level consistent with competent human test-takers — while establishing nothing about the things practice actually stresses: acting on information over time, verifying sources, bearing consequences for being wrong, or knowing what you do not know. A sommelier exam does not check whether the model can smell the wine. The honest reading is that 2023's benchmarks measure a real and new thing — fluent, broad, mostly reliable textual competence — while the public conversation kept sliding from "passed the bar exam" to "can replace lawyers," a leap the measurements never supported.
The mirror-image problem is saturation. MMLU, the workhorse benchmark behind these comparisons, was built in 2020 from questions scraped from free online courses and practice sets; by the time GPT-4 scored in the mid-80s on it, the benchmark was within a few points of its estimated human ceiling and — more corrosively — had been included in enough public text that no one can fully exclude contamination from pre-training corpora. This is the quiet methodological crisis underneath the GPT-4 launch: the field's measuring sticks are being consumed by the very models they measure. The mirage debate of this entry and the saturation problem are the same disease seen from two angles — when a ruler is both discontinuous and compromised, the numbers a launch press release cites deserve exactly the scrutiny the Stanford paper applied.
What it means for people building on these models
I have been writing software professionally for a long time, and this debate matters to me for one reason: capacity planning. If abilities are truly emergent — thresholded, unpredictable, arrive-unannounced — then a model vendor cannot promise that next quarter's model will not trivially break your product's moat, and you cannot budget for capabilities you cannot forecast. If abilities are instead smooth-but-binarized, then smaller models are steadily closing the gap on whatever your product does, and you should assume continuous competitive pressure rather than a one-time discontinuity. These are different business risks with different mitigations, and it is worth knowing which theory you are implicitly running on.
My working position, for what it is worth, is a hybrid. Take the mirage paper seriously as a claim about measurement: any single benchmark discontinuity, especially on a heavily tuned-for benchmark, deserves suspicion, and continuous proxies (partial credit, per-token metrics, calibration curves) are better instruments for tracking model families over time. Take emergence seriously as a claim about products: for discrete, verifiable tasks — code that runs, answers that check out, tools that succeed — user-perceived capability really does flip from useless to useful, and from the product's point of view the flip is the reality. The practical synthesis: expect smooth science and lumpy products. Plan for the lump.
Whichever theory you hold, the engineering response is the same, and it is worth spelling out because almost nobody was doing it in May 2023: build your own evaluation set before you need it. Twenty to two hundred real prompts from your actual workload — including the pathological ones your users actually type — scored consistently, tracked over time, gives you the smooth interior view that public benchmarks cannot. It converts the emergence debate from a forecast into a measurement: when the next model ships, you re-run your set and see whether your task's score moved smoothly or jumped. Teams that did this in early 2023 discovered things the benchmarks could not tell them — that their legal-summarization task gained little from GPT-4 while their extraction task doubled in reliability, or vice versa. The public debate is about the model; your eval set is about your problem, and only the second one bills hours.
There is also a caution here about how we know what we know. GPT-4's report withheld every architectural detail, which means the entire world — including the researchers arguing about emergence — is reasoning about a black box's exterior measurements. The mirage debate is accordingly a debate about graphs in papers, not about models. That epistemic situation is new for our field: we are used to arguing about code we can read. It is worth holding both papers' conclusions a little loosely, because the ground truth they argue about is proprietary, and the arguments are young.
Where this leaves the series
The emergent-abilities debate is the first time in the series where the engineering and the epistemology collide head-on, and it will not be the last. The next entry returns to firmer ground: LoRA and QLoRA, the fine-tuning techniques that made adapting these large, inscrutable models something a person with a single GPU can actually do — and which, fittingly, work precisely because the models' capabilities turn out to live in surprisingly few parameters that matter.
A closing note on method. Both papers cited here are three to fourteen months old as I write this; both will be argued about for years; and by the time this series reaches its later parts, at least one of their central claims will probably look quaint. That is not a reason to skip careful reading — it is the reason for it. In a field that moves this fast, the only durable skill is the ability to read the argument, weigh the measurement, and decide what to do on Monday. That is the muscle this series is training.
Works Cited
OpenAI. "GPT-4 Technical Report." arXiv.org, 2023, arxiv.org/abs/2303.08774. Accessed 4 May 2023.
Schaeffer, Rylan, Brando Miranda, and Sanmi Koyejo. "Are Emergent Abilities of Large Language Models a Mirage?" arXiv.org, 2023, arxiv.org/abs/2304.15004. Accessed 4 May 2023.
Wei, Jason, et al. "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models." arXiv.org, 2022, arxiv.org/abs/2201.11903. Accessed 4 May 2023.
Wei, Jason, et al. "Emergent Abilities of Large Language Models." Transactions on Machine Learning Research, 2022, openreview.net/forum?id=yzkSU5zdwD. Accessed 4 May 2023.