Trust is earned, not given

A different perspective

2024-09-19 · Projects

AI Frontiers, part 11: OpenAI o1 and inference-time compute

Part 11from the AI Frontiers series · 65 parts in all

On 12 September 2024 OpenAI released o1, a model whose distinguishing feature is not what it knows but how long it thinks. The reported numbers were startling in exactly the way the field had stopped expecting: 83% on the 2024 AIME exam versus 13% for GPT-4o, 78% on GPQA Diamond, a 94.8% score on MATH-500, and a Codeforces rating placed around the 89th percentile of human competitors (OpenAI, "Learning to Reason with LLMs"). The architecture was not published, but the description was unusually specific: o1 is trained with reinforcement learning to produce long internal chains of reasoning, learning to search, backtrack, and check its own work, and the amount of that internal computation grows at inference time rather than being fixed when training ends.

That last clause is the whole story. For five years the field had a simple recipe — train a bigger model on more data, then serve it and pay a fixed per-token cost forever. o1 said out loud what a handful of research papers had been arguing all year: you can also buy capability by letting the model spend more thinking when the question deserves it. This entry is about where that idea came from, what is genuinely new in it, and why I think the practical consequence is not "smarter chatbots" but a new kind of routing problem in every production system.

The wall that turned out to be a door

By late 2023 the pretraining conversation had gone slightly sour. Scaling laws (Kaplan et al.) had been repeatedly confirmed — Chinchilla (Hoffmann et al.) corrected the parameters-versus-tokens ratio and made everyone retrain — but the returns on successive frontier models looked increasingly like a matter of polish rather than category change. Data exhaustion was the standing worry: the supply of high-quality human text is finite, and the models were consuming it faster than it could be produced.

The reason this matters is that it framed every downstream decision. If pretraining returns were flattening, then the interesting questions moved from "how much compute do we throw at training?" to "how do we use the compute we have, and how do we use the compute we pay for at run time?" o1 is the first major product built on the second question. And it is not a repudiation of scaling — training o1 presumably consumed enormous compute. It is a second axis of scaling that had been sitting there under-exploited, because the industry had standardized on a serving pattern that assumed thinking takes zero time.

Chain of thought was never just a prompting trick

The intellectual lineage starts with a paper about prompting. "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" (Wei et al.) showed that giving a model worked examples with intermediate steps — not just question and answer — unlocked arithmetic and commonsense reasoning that plain prompting failed at, and that the effect appeared only past a certain model scale. Kojima et al. then showed you did not need worked examples at all: append a single sentence inviting the model to think step by step and the same capabilities surfaced. For about eighteen months, chain of thought was treated by many practitioners as prompt engineering — a formatting convention.

The research community treated it as something else, which is why the later results were predictable in hindsight. If the model's reasoning is a token sequence, then that sequence is searchable. Self-consistency (Wang et al.) sampled many independent chains and took the majority answer, and the accuracy gains were large enough to matter. Tree of Thoughts (Yao et al.) made the search explicit, branching, evaluating partial states and backtracking. These were inference-time algorithms that spent more compute without touching a single weight. What o1 does, on the evidence available, is make that behavior a property of the model rather than a scaffold bolted around it.

And it was never purely a language-model idea. Meta's CICERO, the system that reached human-level play in Diplomacy, was a language model wrapped around a planning module that searched over dialogue intentions before speaking (Meta Fundamental AI Research Diplomacy Team). That is test-time compute with a different vocabulary, and it beat approaches that just made the model bigger.

Rewarding the process, not the answer

The missing piece was a training signal for good reasoning. Supervised fine-tuning on correct solutions teaches a model to imitate the shape of worked examples; outcome-based reinforcement learning tells the model whether the final answer was right and lets it discover its own intermediates, which is exactly what o1 is described as doing. But the sharpest research result on this axis came from process supervision.

"Let's Verify Step by Step" (Lightman et al.) trained two kinds of reward models: an outcome reward model scoring only the final answer, and a process reward model scoring every individual step. Then it used those to rank candidate solutions. The process model won, and won where it mattered most — on problems from a challenging math subset where the outcome model was weakest. The mechanism is intuitive once stated: an outcome model gives a correct final answer credit for a wrong derivation, because it cannot see the difference, and it punishes a correct derivation that ended in an arithmetic slip. Searching over reasoning trajectories requires a scorer that can tell a promising partial solution from a hopeless one. Uesato et al. had reached a compatible conclusion a year earlier about process-based feedback for math word problems.

Put the pieces in order and o1 stops looking like a surprise: chain of thought establishes that reasoning is a token sequence; process reward models establish that intermediate steps can be scored; reinforcement learning (Schulman et al.) is the machinery for optimizing against such a signal; and sampling-and-ranking results establish that spending more compute per query produces better answers. The novelty is doing all of it at frontier scale and shipping it.

What test-time compute actually buys

Two papers published in the months before o1 gave the quantitative version of the claim. "Large Language Monkeys" (Brown et al.) found that repeatedly sampling from a model and selecting the best answer — plain repeated sampling with a verifier — improved pass rates smoothly with the number of attempts, and in some settings overtook a substantially larger model. "Scaling LLM Test-Time Compute Optimally" (Snell et al.) went further and asked where the marginal dollar is best spent, and found the answer depends on the difficulty of the prompt: for easier questions, a small model sampling many times beats a bigger model sampling once; for hard questions, the ordering reverses. That is a routing rule, not a benchmark curiosity.

It also explains the shape of the o1 lineup. o1-mini is reported to reach roughly 70% of the flagship's coding performance at a fraction of the cost — the economic point being that you do not want your best reasoner answering every request. And it explains why the most interesting engineering question around o1 was never "is it good?" but "which requests deserve to think?" A reasoning model with a large thinking budget is a scarce resource in exactly the way a database index is: cheap to use indiscriminately, expensive to use everywhere.

Exam scores are not work

The benchmarks chosen to introduce o1 deserve a skeptical paragraph. AIME is a competition for high-school students; GPQA is designed to be "Google-proof" for PhD-level science questions; Codeforces is competitive programming under a clock. All three are well-constructed, hard to game, and share one property that flatters reasoning models: the answer is checkable. When the answer is checkable, you can train against it, sample against it, and score candidates against it. The whole test-time compute economy rests on the existence of a verifier.

Most of the work people want done does not have a verifier. "Is this migration plan sound?" and "which of these three architectures will age better?" are exactly the questions o1 would be most valuable for and exactly the questions no benchmark measures, because nobody can grade the answer automatically. This is the same gap that part 2 ran into from the other direction: the benchmarks saturate, the capabilities that matter resist measurement, and the distance between the two is where most of the arguing happens.

Context for how fast the verifiable frontier moved: in July 2024, Google DeepMind's AlphaProof and AlphaGeometry 2 reached silver-medal performance at the International Mathematical Olympiad, solving four of six problems with a system built around formal verification in Lean (Google DeepMind). That is a different technical path — formal proofs rather than free-form chains — but it makes the same point about the destination. The tasks that fall first are the ones where correctness can be mechanically decided. The tasks that fall last are the ones where a human has to decide what "correct" even means.

The opacity problem

OpenAI made one decision that I expect to be litigated for years: the raw reasoning tokens are not shown. Users get a summary of the thinking, not the thinking. The stated rationale is competitive and safety-related — the chain of thought is a training artifact worthy of protection, and might contain content the model would not say in its answer.

There is a third cost, less discussed. Reasoning models make evaluation harder because they move the failure from the answer to the process. With a text model you can read the output and judge it. With a reasoning model you get a conclusion and a summary of the summary, and you cannot tell whether the model reasoned its way there or pattern-matched to a plausible answer and then constructed supporting steps afterwards. That distinction is the difference between a model you can improve with better data and one you can only improve with more compute. It also means the summary is a liability in regulated contexts, where "show your work" is not a stylistic preference but a requirement.

Both arguments are real, and both have costs. The chain of thought had become the single best debugging surface practitioners had. When a model got a question wrong, reading the reasoning told you whether it misread the question, misremembered a fact, or made an arithmetic slip — a distinction with entirely different fixes. Hide it and you go back to treating the model as a black box whose failures are anecdotes. The safety argument also cuts the other way: the o1 system card devotes substantial space to deception evaluations, because a model that plans internally is harder to supervise than one that narrates every step. Redacting the narration does not make the planning go away.

Why this changes architecture, not just benchmarks

The practical consequence I keep arriving at is boring, which is how you know it is probably right: reasoning models reintroduce variable cost and variable latency into inference, and most systems designed in 2023 assumed both were constants. When every request costs the same handful of tokens and returns in a second, you can build a synchronous API, log costs as a function of request count, and set timeouts with confidence. When a request can spend ten or a hundred times longer thinking, all three assumptions break at once.

So the systems that will age well are the ones that treat "how much should we think about this?" as an explicit, observable decision — a classifier, a heuristic, a cheap first pass that escalates to an expensive second pass. That is a return to a very old pattern in computing: cheap speculation with expensive verification, which is what branch prediction and speculative execution have done in CPUs for decades. It is also, not coincidentally, a pattern that the mixture-of-experts architecture applies inside a single forward pass. Routing keeps showing up as the answer.

The unsettled question is what happens when models are trained to think longer and the thinking is invisible. Every deployment I have seen that works has an evaluation harness, a review queue, and a human who can explain a failure. o1's headline numbers are the best evidence yet that inference-time compute is a real second axis of capability. Its hidden reasoning is a reminder that capability without visibility is a harder thing to operate, and operating is where all the value lives.

Works Cited

Brown, Bradley, et al. "Large Language Monkeys: Scaling Inference Compute with Repeated Sampling." arXiv, 2024, arxiv.org/abs/2407.21787. Accessed 19 Sept. 2024.

Cobbe, Karl, et al. "Training Verifiers to Solve Math Word Problems." arXiv, 2021, arxiv.org/abs/2110.14168. Accessed 19 Sept. 2024.

Hendrycks, Dan, et al. "Measuring Mathematical Problem Solving With the MATH Dataset." arXiv, 2021, arxiv.org/abs/2103.03874. Accessed 19 Sept. 2024.

Google DeepMind. "AI Achieves Silver-Medal Standard Solving International Mathematical Olympiad Problems." Google DeepMind, 25 July 2024, deepmind.google/discover/blog/ai-solves-imo-problems-at-silver-medal-level/. Accessed 19 Sept. 2024.

Hoffmann, Jordan, et al. "Training Compute-Optimal Large Language Models." arXiv, 2022, arxiv.org/abs/2203.15556. Accessed 19 Sept. 2024.

Kaplan, Jared, et al. "Scaling Laws for Neural Language Models." arXiv, 2020, arxiv.org/abs/2001.08361. Accessed 19 Sept. 2024.

Kojima, Takeshi, et al. "Large Language Models are Zero-Shot Reasoners." arXiv, 2022, arxiv.org/abs/2205.11916. Accessed 19 Sept. 2024.

Lewkowycz, Aitor, et al. "Solving Quantitative Reasoning Problems with Language Models." arXiv, 2022, arxiv.org/abs/2206.14858. Accessed 19 Sept. 2024.

Lightman, Hunter, et al. "Let's Verify Step by Step." arXiv, 2023, arxiv.org/abs/2305.20050. Accessed 19 Sept. 2024.

Meta Fundamental AI Research Diplomacy Team. "Human-Level Play in the Game of Diplomacy by Combining Language Models with Strategic Reasoning." Science, vol. 378, no. 6624, 2022, arxiv.org/abs/2210.05492. Accessed 19 Sept. 2024.

OpenAI. "Learning to Reason with LLMs." OpenAI, 12 Sept. 2024, openai.com/index/learning-to-reason-with-llms/. Accessed 19 Sept. 2024.

---. "OpenAI o1 System Card." OpenAI, 12 Sept. 2024, cdn.openai.com/o1-system-card-20240917.pdf. Accessed 19 Sept. 2024.

Rein, David, et al. "GPQA: A Graduate-Level Google-Proof Q&A Benchmark." arXiv, 2023, arxiv.org/abs/2311.12022. Accessed 19 Sept. 2024.

Schulman, John, et al. "Proximal Policy Optimization Algorithms." arXiv, 2017, arxiv.org/abs/1707.06347. Accessed 19 Sept. 2024.

Snell, Charlie, et al. "Scaling LLM Test-Time Compute Optimally Can Be More Effective Than Scaling Model Parameters." arXiv, 2024, arxiv.org/abs/2408.03314. Accessed 19 Sept. 2024.

Uesato, Jonathan, et al. "Solving Math Word Problems with Process- and Outcome-Based Feedback." arXiv, 2022, arxiv.org/abs/2211.14275. Accessed 19 Sept. 2024.

Wang, Xuezhi, et al. "Self-Consistency Improves Chain of Thought Reasoning in Language Models." arXiv, 2022, arxiv.org/abs/2203.11171. Accessed 19 Sept. 2024.

Wei, Jason, et al. "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models." arXiv, 2022, arxiv.org/abs/2201.11903. Accessed 19 Sept. 2024.

Yao, Shunyu, et al. "Tree of Thoughts: Deliberate Problem Solving with Large Language Models." arXiv, 2023, arxiv.org/abs/2305.10601. Accessed 19 Sept. 2024.