AI Frontiers, part 23: Where the frontier goes next
Part 23from the AI Frontiers series · 65 parts in all
Twenty-three entries is enough distance to see the shape of the thing. The series started with a paper about attention and ends here, and the arc in between has a clear structure: a period where capability followed compute and data on a predictable curve, a period where the second axis of inference-time compute opened up a great deal of headroom nobody had priced in, and a period that is still running in which the practical constraint has moved from what models can do to whether anyone can tell what they are doing.
I want to use this last entry to argue that the frontier is now bounded by four resources, only one of which is compute, and that the binding one is the newest and the least discussed.
Resource one: compute, which is no longer the interesting constraint
Scaling laws held. That is the least controversial statement in the field, and it should still be said plainly, because the predictions of their imminent failure have been made every year for six years and have been wrong every year (Kaplan et al.; Hoffmann et al.). The chinchilla correction in 2022 was not a refutation of the law — it was a better specification of it, and it is why models trained after it use far more tokens per parameter than their predecessors.
What changed is not the validity of scaling but its economics. Training runs consume capital at a scale that concentrates the frontier in a handful of organizations, the energy footprint of training and serving has become a real constraint rather than a rounding error (de Vries; International Energy Agency), and the marginal return on another order of magnitude of pretraining compute is increasingly the sort of improvement that shows up as a better product rather than a new capability. None of that means scaling stopped working. It means the next order of magnitude is now a business decision rather than a technical one, and the industry has decided to spend the compute at inference time instead, where the returns are more immediate and the capital outlay is amortized over customers (part 11, part 16).
Resource two: data, and the wall everyone can see coming
The supply of high-quality human text is finite and the models have consumed most of the usable part of it. Estimates of when the stock runs out disagree on the date and agree on the shape: the highest-quality sources are exhausted first, and continued scaling requires either lower-quality data or synthetic generation (Villalobos et al.).
The two responses are both advanced and both incomplete. Synthetic data works — the evidence in part 21 is solid — as long as you accumulate real data, verify what you generate, and accept that density of reasoning is traded against breadth of knowledge. And data quality improvements have proven worth more than volume: the Phi line, the FineWeb pipeline (Penedo et al.), and every serious deduplication study all say the same thing, which is that a small clean corpus beats a large dirty one by a margin that surprises people who think of data as a commodity.
The part of the data problem that gets less attention is the one that will outlast the scraping debate: the highest-value data is not on the open web. It is inside organizations, in tickets and contracts and diagnostic logs and institutional memory, and it is siloed, badly labelled and often legally constrained. Which means the durable competitive advantage in applied AI is not a better model — it is better access to a domain's data and the willingness to instrument it. That has been true since before any of this and it has not stopped being true.
Resource three: energy, which is now on the critical path
For most of the last decade the energy cost of machine learning was a footnote in a sustainability report. That changed. Data center electricity demand grew fast enough that utility planning, grid interconnection queues and local permitting became practical constraints on where training clusters can go, and the largest facilities are now large enough that their siting is a matter of public interest (International Energy Agency).
The engineering response is the part I find most interesting, because it collides with the serving optimizations of the last three entries in a satisfying way. Quantization, prefix caching, speculative decoding and batching are all, viewed from a different angle, energy efficiency measures: the work not done is the power not drawn. The recent history of the field is largely a story of algorithmic efficiency catching up with the demand that capability created, and there is no particular reason to think that has stopped. The pessimistic case is that demand grows faster; the optimistic case is that a factor of ten of headroom sits in the serving stack. Both have been true simultaneously for several years.
Resource four: verification, which is the real frontier
Here is the argument, and I think it will be the organizing idea of the next few years. Every meaningful capability gain of the last two years came from a domain where correctness could be checked mechanically: unit tests, arithmetic, formal proofs, multiple-choice keys, executable code. Reinforcement learning needs a reward, and a verifier is a reward. The domains with verifiers are now being solved faster than anyone expected, and the domains without them are not moving at the same rate.
Which produces a strange split. A model can solve graduate-level physics problems and fail at reconciling two invoices, and the reason is not intelligence. It is that one task has an oracle and the other does not. The bottleneck on applied progress is going to be how well we can construct checks for the things we actually care about: does this contract create a risk, is this migration plan sound, is this diagnosis consistent with the evidence, does this customer's complaint reflect a real bug. Constructing those checks is domain work, it is uncooperative, it does not generalize across industries, and it is where the value is.
The measurement literature already reflects this shift. Agent benchmarks that ignore cost optimize for a world nobody deploys in (Kapoor et al.). Task-length horizons — how long a task a model can complete autonomously, not how many questions it answers correctly — predict better than accuracy scores do (Kwa et al.). Long-context claims collapse under harder retrieval tests (Hsieh et al.). In each case the better metric is closer to the actual work and harder to game. The field is being forced, by its own success, into more honest evaluation.
The evaluation crisis, stated bluntly
A benchmark that becomes a training target stops being a benchmark. This has already happened to nearly every public evaluation set of the last five years, contamination is measurable and widespread (Sainz et al.), and the result is that a leaderboard number is now weak evidence about capability and strong evidence that someone optimized for it.
The only durable response is the one every serious lab has converged on privately: keep evaluation sets private, refresh them, prefer constructed tasks with automatic verification, and report capability alongside cost and latency because those are part of the honest description. For anyone building a product, the practical version is one afternoon of work that pays for itself forever — collect a few hundred examples of your actual task with known outcomes, and evaluate every model and prompt change against that instead of against anything published. You will make better decisions than the leaderboard would have produced, and you will have a defensible answer when someone asks why you chose a model.
Open weights, moats, and a practical prediction
Capability diffuses faster than anyone's business model would like. The pattern of the last three years is that a frontier capability is demonstrated, described publicly, and reproduced in open weights within roughly six to eighteen months — sometimes weeks, as with the reasoning models of part 16 — and then pushed down into small models through distillation. Training-efficiency results like DeepSeek-V3's (part 13) accelerate the diffusion further.
Two consequences follow. First, "we have a better model" is a weaker moat every quarter it is held. Second, the advantages that persist are the ones that are not models: proprietary data with a real feedback loop, a workflow that is genuinely hard to replicate, distribution, and the accumulated reliability of a system that has been running for years. This is the boring conclusion that every technology wave arrives at, and the AI wave is no exception.
What I would build, if I were starting today, is not another wrapper. It is the verification and instrumentation layer: the evaluation harness, the logging, the human review queue, the permission model, the cost accounting. Those are the components that decide whether a system can be trusted with real work, they are unglamorous enough that few teams build them well, and they get more valuable with every model improvement rather than less.
What would make me wrong
A prediction that cannot be falsified is a mood, so here is what would change my mind about the argument that verification is the binding constraint.
A general verification method. If someone finds a way to construct reliable checks for unstructured judgment — a debate protocol that works outside a laboratory, or a self-critique mechanism that demonstrably catches its own errors rather than rationalizing around them — then the split between verifiable and unverifiable domains collapses and progress resumes on the messy half. This is the outcome I am least confident is impossible, and it is the one I would most like to be wrong about.
Energy or hardware discontinuities. A step change in compute efficiency would make the inference-time scaling approach affordable enough to brute-force past reasoning difficulties rather than verifying them. It would not remove the constraint so much as change its price, and the history of this field suggests that when something becomes ten times cheaper, people find uses nobody predicted.
Data breakthroughs. If the practical limits of synthetic data turn out to be far more permissive than the current literature suggests — and the accumulation result in part 21 points that way — then the data wall is a speed bump rather than a wall, and pretraining returns could continue longer than the pessimistic reading implies.
Or a plateau that is not about any of this. The possibility I hold most loosely is that the curve simply flattens, that the gains from here are incremental, and that the interesting question stops being capability and becomes distribution. Many technologies follow that path, and the field's confidence in its own extrapolations has been wrong before in both directions.
Signals I would actually watch
Four things, in order of how much they would tell me.
Task-length horizons, not benchmark scores. A published, replicated measure of how long a task a system can complete autonomously is worth more than any accuracy percentage, because it maps directly to what can be delegated and it is hard to game without actually being able to do the work.
The cost curve for a fixed task. Track the price of automating one specific well-defined workflow over time. Falling cost curves are what turn capabilities into products, and the shape of the decline says more about adoption than the frontier's headline numbers.
Enterprise reliability metrics. How often do agentic systems run unattended without human intervention, and what is the incident rate when they do? These are the numbers that decide whether the technology escapes the demo, and they are conspicuously absent from most public discussion.
Independent evaluation capacity. Whether a credible third party can meaningfully test the systems that matter. If that capacity grows, the incentive structure improves; if it remains voluntary self-reporting, it does not. This is the slowest-moving signal and probably the most consequential.
A closing note on the series
The through-line across all twenty-three entries is not any particular technique. It is that each step forward came from someone finding the real constraint and working on it instead of the apparent one. The apparent constraint in 2023 was model intelligence; the real one was context. Then the real one was memory bandwidth, then integration, then verification. The people who did the useful work were the ones who could tell the difference.
I will end where the first entry started, with the observation that attention is a mechanism for deciding what matters. Twenty-three parts later, that is still a good description of the engineering problem: deciding what to attend to, what to store, what to forget, what to verify, and what to spend on. The models are remarkably good at the first of those and entirely dependent on us for the rest. Sutton's bitter lesson says we should prefer the methods that scale over the ones that encode our intuitions, and it has been right for seventy years of computing (Sutton). But someone still has to decide what to measure, and that judgment is not going to be automated away any time soon.
Works Cited
Amodei, Dario. "Machines of Loving Grace." darioamodei.com, Oct. 2024, darioamodei.com/machines-of-loving-grace. Accessed 17 Sept. 2026.
Altman, Sam. "The Intelligence Age." blog.samaltman.com, 23 Sept. 2024, blog.samaltman.com/the-intelligence-age. Accessed 17 Sept. 2026.
Anthropic. "Responsible Scaling Policy." Anthropic, 2023, www.anthropic.com/news/anthropics-responsible-scaling-policy. Accessed 17 Sept. 2026.
de Vries, Alex. "The Growing Energy Footprint of Artificial Intelligence." Joule, vol. 7, no. 10, 2023, pp. 2191–94. Accessed 17 Sept. 2026.
European Parliament and Council. "Regulation (EU) 2024/1689 Laying Down Harmonised Rules on Artificial Intelligence." Official Journal of the European Union, 2024. Accessed 17 Sept. 2026.
Hoffmann, Jordan, et al. "Training Compute-Optimal Large Language Models." arXiv, 2022, arxiv.org/abs/2203.15556. Accessed 17 Sept. 2026.
Hsieh, Cheng-Ping, et al. "RULER: What's the Real Context Size of Your Long-Context Language Models?" arXiv, 2024, arxiv.org/abs/2404.06654. Accessed 17 Sept. 2026.
International Energy Agency. "Electricity 2024: Analysis and Forecast to 2026." IEA, Jan. 2024, www.iea.org/reports/electricity-2024. Accessed 17 Sept. 2026.
Kaplan, Jared, et al. "Scaling Laws for Neural Language Models." arXiv, 2020, arxiv.org/abs/2001.08361. Accessed 17 Sept. 2026.
Kapoor, Sayash, et al. "AI Agents That Matter." arXiv, 2024, arxiv.org/abs/2407.01502. Accessed 17 Sept. 2026.
Kwa, Thomas, et al. "Measuring AI Ability to Complete Long Tasks." arXiv, 2025, arxiv.org/abs/2503.14499. Accessed 17 Sept. 2026.
National Institute of Standards and Technology. "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile." NIST AI 600-1, July 2024. Accessed 17 Sept. 2026.
OpenAI. "Preparedness Framework (Beta)." OpenAI, Dec. 2023, openai.com/safety/preparedness. Accessed 17 Sept. 2026.
Penedo, Guilherme, et al. "The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale." arXiv, 2024, arxiv.org/abs/2406.17557. Accessed 17 Sept. 2026.
Sainz, Oscar, et al. "NLP Evaluation in Trouble: On the Need to Measure LLM Data Contamination for Each Benchmark." arXiv, 2023, arxiv.org/abs/2310.18018. Accessed 17 Sept. 2026.
Sastry, Girish, et al. "Computing Power and the Governance of Artificial Intelligence." arXiv, 2024, arxiv.org/abs/2402.08797. Accessed 17 Sept. 2026.
Stanford Institute for Human-Centered AI. "Artificial Intelligence Index Report 2025." Stanford HAI, 2025, hai.stanford.edu/ai-index/2025-ai-index-report. Accessed 17 Sept. 2026.
Sutton, Rich. "The Bitter Lesson." Incomplete Ideas, 2019, www.incompleteideas.net/IncIdeas/BitterLesson.html. Accessed 17 Sept. 2026.
Villalobos, Pablo, et al. "Will We Run Out of Data? Limits of LLM Scaling Based on Human-Generated Data." arXiv, 2022, arxiv.org/abs/2211.04325. Accessed 17 Sept. 2026.