AI Frontiers, part 16: Test-time scaling and the reasoning models
Part 16from the AI Frontiers series · 65 parts in all
When OpenAI released o1 in September 2024, the mechanism was described in a blog post and the recipe was not. By July 2025 that had changed completely, and the change is the most consequential thing that happened in the intervening ten months. DeepSeek published R1 in January with a technical report describing, in detail, how to train a reasoning model with reinforcement learning on verifiable rewards — including the observation that reasoning behavior emerges from pure reinforcement learning without any supervised chain-of-thought data to start from (DeepSeek-AI, "DeepSeek-R1"). Within weeks, "reasoning model" had become a product category rather than a research result, and the techniques for building one were public, cheap enough to replicate, and already being distilled into small models.
So this entry is less about whether inference-time compute works — part 11 made that case — and more about what the field learned when the recipe became public: which parts of the pipeline matter, how the training signal is constructed, what gets distilled, and where the whole approach runs into a wall made of arithmetic.
The recipe, as far as anyone has published it
The R1 report's central claim is that a base model, given reinforcement learning against verifiable rewards, will learn to produce longer and more structured reasoning — that "aha" moments of self-correction appear without being trained for. Verifiable rewards mean exactly what they sound like: problems where you can check the answer mechanically. Math with a numeric answer. Code that either passes the tests or does not. Formal proofs. Multiple-choice questions. In R1's case, the mixture was math, code and logic problems with rule-based reward functions, plus a format reward that encouraged the model to put its reasoning inside an explicit scratchpad.
Two technical choices in the report were widely copied. The first was GRPO, the group-relative policy optimization loss that the DeepSeekMath paper had introduced a year earlier — a variant of policy-gradient RL that compares candidates within a sampled group rather than training a separate value network, which cuts memory substantially and removes a large source of instability. The second was distillation. Once R1 existed, the team generated reasoning traces from it and fine-tuned smaller dense models on those traces, producing 1.5B to 70B models that beat their own base versions on reasoning benchmarks by large margins. That result is the one with the longest shadow, because it means the expensive part — reinforcement learning at scale — has to be done once, at the top of the pyramid, and can then be propagated downward cheaply.
The academic follow-up came fast and pushed in the same direction with less. The s1 paper (Muennighoff et al.) showed that a thousand carefully selected reasoning examples plus a simple decoding-time trick — "budget forcing", where you cut the model off or extend it to control how long it thinks — produced surprisingly strong reasoning from a small model, trained for a reported cost in the tens of dollars. The budget-forcing detail is the interesting one: it showed that the amount of thinking is a controllable knob at inference time rather than a fixed property of the model, which is the practical foundation for everything that followed.
What the frontier did with it
By the middle of 2025 every major lab had a reasoning mode, and the differences between them were revealing. Claude 3.7 Sonnet shipped in February with "extended thinking" as a toggle, exposing a summarized reasoning trace and letting the caller set a budget (Anthropic). OpenAI's o3-mini added low, medium and high reasoning-effort settings, and o3 followed in April (OpenAI, "Introducing OpenAI o3 and o4-mini"). Gemini 2.5 arrived with thinking on by default. The common thread is not the models — it is that the user or the application now decides how much the model thinks, which is exactly the routing problem part 11 predicted, appearing in product interfaces within six months.
The benchmark story also got a second act. o3 set records on ARC-AGI, a benchmark designed to be resistant to memorization, but the demos that circulated paired the score with a cost figure that made eyes water — on the order of thousands of dollars for the full evaluation set at high reasoning effort (ARC Prize Foundation). FrontierMath, a benchmark of research-level mathematics designed with Epoch AI, likewise showed that the same model at different compute budgets behaves like different products (Glazer et al.). That pairing of capability and cost is the honest way to present these models, and it took most of a year for the presentations to adopt it.
Where the compute actually goes
It is worth being precise about the cost structure, because the word "reasoning" obscures it. In a standard language model, generation is memory-bandwidth bound: the expensive part is reading the weights for each token, so throughput is roughly flat in batch size and cost scales linearly with output tokens. A reasoning model produces many more output tokens per answer — often ten to fifty times more — and those tokens are not free. But there is a second-order effect that matters as much: the reasoning tokens are generated, then read back into context, which means the prompt grows monotonically with every step and the attention cost grows with it. Long chains of thought are quadratic-ish in a dimension that part 7 and part 8 both treated as the bottleneck, which is why KV-cache management is the next thing this series has to talk about.
Then there is the diminishing return, which is the most practically important finding of the year. More thinking helps up to a point and then stops, and the point depends on the question. The overthinking literature made this concrete in a paper whose title is a runnable experiment: "Do NOT Think That Much for 2+3=?" (Chen et al.). Providing a long reasoning budget for a trivially easy question does not merely waste tokens — it can degrade accuracy, because the model invents complications that were not there. The same paper's counterpart observation is that hard problems genuinely benefit from long traces, which is why the routing heuristics matter and why a monolithic "always think hard" policy is a losing strategy for anyone paying per token.
The evaluation problem, again, differently
Verifiable rewards are powerful and narrow. Everything the reasoning-model wave has achieved sits inside domains where correctness can be checked by a rule: arithmetic, unit tests, formal proofs, multiple-choice keys. That has two consequences worth stating plainly.
First, the training data problem is now a verification problem. The bottleneck is not producing answers, it is producing a way to tell a good answer from a bad one at scale. Where you cannot write a verifier you cannot use this method, and a great deal of valuable work — "is this analysis actually sound?", "is this contract clause a risk?" — does not have a checker that runs in a few milliseconds.
Second, benchmarks in verifiable domains saturate. AIME went from a hard contest for models to near-solved in about eighteen months. That does not mean the underlying capability generalized; it means the measurement instrument maxed out. When a benchmark becomes a training target, it stops being a benchmark, and the field's habit of announcing that reasoning has been solved on the strength of a test that the models were trained on is the least impressive thing in the whole area.
What the distilled models changed
The second-order effect of publishing R1 was the creation of a reasoning-trace economy. Because the report described distillation explicitly, and because the model weights were released under permissive terms, a large number of open-weight models appeared within weeks that had been fine-tuned on reasoning traces generated by a stronger model. Many of them were sized between 7B and 32B — the range that fits on a workstation — and several outperformed substantially larger non-reasoning models on math and code benchmarks.
Three things follow. First, the capability curve for open models flattened onto the frontier's reasoning behavior much faster than the underlying training work would suggest, because tracing is a form of copying. Second, the evaluation question got harder: a distilled model that scores 90% on AIME may be reproducing the teacher's problem-solving style rather than the teacher's underlying competence, and the way to tell the difference is to test out-of-distribution. Third, the research community started asking whether you can skip the supervised warm-up entirely — the Zero paper (Sun et al.) is an attempt to do reinforcement learning from nothing but self-proposed tasks and a code executor, which is the logical endpoint of "verifiable rewards are the whole game."
I would expect this pattern to keep repeating for as long as a frontier exists: capability is invented at the top, demonstrated publicly, and absorbed downward within months. The practical implication for anyone building products is that being early to a capability is worth less than being early to the problem it solves, because the capability itself is a commodity on a short clock.
The horizon, which is the more interesting measure
If you want a number that is harder to overfit to, the best candidate came from METR: rather than measuring accuracy on problems, they measured the length of task a model can complete autonomously at a given success rate, and estimated how that horizon grows over time. Their March 2025 analysis put the doubling period at roughly seven months, with the caveat that the estimate rests on a fairly small number of long-horizon evaluations (Kwa et al.).
I find that framing much more useful than a percentage on a contest, because it translates directly into engineering. If an agent can reliably handle a twenty-minute task today, the question "what would we delegate?" has an answer, and it is different from the answer for a two-hour task. The reasoning-model literature is, in effect, buying task length with compute. The interesting part of the next year is not whether the horizon grows — it is whether it grows in the kinds of tasks that businesses actually have, which are full of ambiguity, implicit requirements and checkers that do not exist.
The distillation result tells us what to expect for the products. The expensive capability gets built once at the frontier with reinforcement learning, and then it flows downward into small models via synthetic traces, where it becomes cheap, private and fast. That is the same pattern as part 9, with reasoning instead of instruction-following. It also means the interesting frontier capability is, for a while, always available below the frontier — which is a good deal for anyone building products and a difficult situation for anyone whose business model is being the only place a capability exists.
Operating one of these things
If you run a reasoning model in production, four operational facts shape the design.
Latency is now variable and user-visible. A request that thinks for thirty seconds is a request a user will abandon unless the interface does something. The standard answer is to stream the reasoning summary so progress is visible, but that creates its own problem: a user who reads a wrong intermediate step loses confidence in an answer that turns out to be right. The interfaces that work show high-level progress — searching, checking, verifying — rather than the raw trace.
Cost caps are a correctness requirement, not a nicety. Because the token count is model-chosen, a single pathological request can burn an entire day's budget. Every deployment needs a hard ceiling on thinking tokens per request, and a fallback answer when the ceiling is hit.
Caching is worth more here than anywhere else. The reasoning prefix is generated once and reused within a session; if your stack does not reuse it, you are paying for the same thinking twice. This is the strongest practical argument yet for the prefix-caching machinery in modern serving stacks.
Route by difficulty, and measure the router. The overthinking result is the empirical justification, but the engineering reason is more basic: if easy requests cost the same as hard ones you have made a bad product, and if you cannot tell which is which you have made an unmeasurable one. Start by logging the thinking-token count per request class. The distribution will surprise you, and the surprise will be worth more than any benchmark score.
Works Cited
Anthropic. "Claude 3.7 Sonnet and Claude Code." Anthropic, 24 Feb. 2025, www.anthropic.com/news/claude-3-7-sonnet. Accessed 17 July 2025.
ARC Prize Foundation. "OpenAI o3 Breakthrough High Score on ARC-AGI-Pub." ARC Prize, Dec. 2024, arcprize.org/blog/oai-o3-pub-breakthrough. Accessed 17 July 2025.
Chen, Xingyu, et al. "Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs." arXiv, 2024, arxiv.org/abs/2412.21187. Accessed 17 July 2025.
DeepSeek-AI. "DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning." arXiv, 2025, arxiv.org/abs/2501.12948. Accessed 17 July 2025.
---. "DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models." arXiv, 2024, arxiv.org/abs/2402.03300. Accessed 17 July 2025.
Glazer, Elliot, et al. "FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI." arXiv, 2024, arxiv.org/abs/2411.04872. Accessed 17 July 2025.
Kwa, Thomas, et al. "Measuring AI Ability to Complete Long Tasks." arXiv, 2025, arxiv.org/abs/2503.14499. Accessed 17 July 2025.
Muennighoff, Niklas, et al. "s1: Simple Test-Time Scaling." arXiv, 2025, arxiv.org/abs/2501.19393. Accessed 17 July 2025.
OpenAI. "Introducing OpenAI o3 and o4-mini." OpenAI, 16 Apr. 2025, openai.com/index/introducing-o3-and-o4-mini/. Accessed 17 July 2025.
---. "OpenAI o3-mini." OpenAI, 31 Jan. 2025, openai.com/index/openai-o3-mini/. Accessed 17 July 2025.
Rein, David, et al. "GPQA: A Graduate-Level Google-Proof Q&A Benchmark." arXiv, 2023, arxiv.org/abs/2311.12022. Accessed 17 July 2025.
Snell, Charlie, et al. "Scaling LLM Test-Time Compute Optimally Can Be More Effective Than Scaling Model Parameters." arXiv, 2024, arxiv.org/abs/2408.03314. Accessed 17 July 2025.
Sun, Yutao, et al. "Absolute Zero: Reinforced Self-Play Reasoning with Zero Data." arXiv, 2025, arxiv.org/abs/2505.03335. Accessed 17 July 2025.