Trust is earned, not given

A different perspective

2025-10-09 · Projects

AI Frontiers, part 56: Science with models — from conjecture to verification

Part 56from the AI Frontiers series · 65 parts in all

The discussion about AI and science runs together two claims that need separating. The first is that models are useful tools inside scientific workflows — for prediction, simulation, literature synthesis, code, and instrument control. That claim is uncontroversially true and mostly boring. The second is that models can make discoveries: produce new knowledge that a scientist would not have found, or would have found much later. That claim is interesting, and it turns out to have a sharp boundary. Every convincing case of machine discovery shares one feature — a cheap, trustworthy evaluator that can tell a good candidate from a bad one without human judgment — and every disappointing case fails on exactly that point.

This entry is about that boundary. It follows part 55, which covered the field with the best existing case study, and it draws on the general picture of model-driven discovery that the field has been assembling for several years (Wang et al., "Scientific Discovery in the Age of Artificial Intelligence").

Three classes of use, with very different reliability

Surrogates. A model that approximates an expensive computation or measurement is about substitution: run the cheap approximation a million times, spend the expensive ground truth sparingly. Weather forecasting is the clearest success, where learned models trained on decades of reanalysis data now produce ten-day global forecasts in seconds and at accuracy comparable to or better than the operational numerical models that require supercomputers (Lam et al., "GraphCast"; Bi et al., "Pangu-Weather"). The scientific significance is not that the model is smarter than physics; it is that a forecast that costs seconds can be run in thousands of ensemble members, so the thing being improved is the decision, not the single prediction. The same pattern applies to molecular simulation, materials property prediction, and any domain where the exact computation is too slow to explore a design space.

Search with a verifier. A generator proposes candidates; an automatic evaluator scores them; a search procedure keeps the good ones. This is where machine discovery has actually delivered, and the reason is that the evaluator removes the need for human taste in the loop.

Hypothesis generation. Producing candidate explanations, mechanisms, or experimental designs for a human to evaluate. Useful, and irreducibly dependent on expert judgment, which is why it produces the most inflated headlines and the least verifiable results.

What the successful search results have in common

FunSearch paired a large language model with an automatic evaluator and searched over programs rather than over text, finding new constructions for combinatorial problems — including an improvement to a long-standing upper bound in the cap set problem — by having the model propose code and the evaluator score its behavior (Romera-Paredes et al.). The critical design decision was searching in a space where correctness is checkable by execution.

AlphaTensor searched over matrix multiplication algorithms, using correctness against exact arithmetic and speed as the score, and found algorithms that beat the best human-known procedures for specific matrix sizes (Fawzi et al.). AlphaGeometry combined a language model for conjecture construction with a symbolic deduction engine as the verifier, reaching expert-level performance on olympiad geometry problems whose solutions can be mechanically checked (Trinh et al.). GNoME generated candidate crystal structures and screened them with a graph network trained on density functional theory calculations, producing hundreds of thousands of plausible-stable materials, a subset of which were subsequently realized in the lab (Merchant et al.).

Read together, the recipe is consistent and unglamorous. Choose a domain where the objective is mechanically checkable — a proof that verifies, an algorithm that runs correctly, a structure whose stability can be evaluated by an accurate surrogate. Build the generator. Build the verifier. Then spend most of the effort on making the verifier trustworthy, because the search will exploit any weakness in it with relentless enthusiasm. It is exactly the reward-hacking dynamic from part 44, with a scientific objective substituted for a preference model.

The limitation is structural rather than technical: the method works where you can check, and science's most important questions are usually the ones where you cannot yet.

Closed loops: the autonomous lab

A second pattern puts the model inside the physical loop rather than in front of it. Coscientist demonstrated a language model system that designed and executed chemical experiments by driving robotic hardware and analysis instruments, including successful optimization of a palladium-catalyzed cross-coupling reaction (Boiko et al.). The A-Lab combined robotic synthesis with machine-learned structure models to iterate on inorganic materials candidates autonomously, reporting a majority success rate over a large batch of proposed compounds (Szymanski et al.).

The interesting engineering in both cases is not the model. It is the interface: the ability to express an experimental protocol as something a program can execute, to read instrument output back into a structured form, and to close the loop between result and next experiment. That is the same trace-and-replay discipline as part 49 and the same structured-interface work as part 45, applied to instruments. Where the loop closes, throughput improves by the ratio of human attention to machine attention, which is usually two to three orders of magnitude. Where it does not, the model is a suggestion engine and the bottleneck does not move.

The claims that did not survive scrutiny

End-to-end automation of the research process itself has been attempted and is instructive. The AI Scientist system generated complete papers — hypotheses, code, experiments, write-ups, and an automated review step (Lu et al.). Reviews of its output, including the authors' own flagging of limitations, identified the same failures that recur across every "AI scientist" demonstration: hallucinated or misattributed citations, experiments that do not test the stated hypothesis, results that are not reproducible from the artifact, and conclusions that overshoot the evidence. The system's strongest contribution was arguably the review component, which is a useful reminder that in science, machine-generated criticism has a better track record than machine-generated conclusions, because criticism is easier to check and its failures are cheaper.

The broader epistemological worry has been articulated well in the sociology of science literature: models that make individual tasks easier while making it harder to understand the whole picture can degrade the reliability of the scientific enterprise even as they accelerate its output (Messeri and Crockett). Every researcher who has watched a plausible-looking AI-generated summary displace a careful literature review knows the shape of that risk. The mitigation is the same one that runs through this entire series — keep a human-verifiable ground truth in the loop and measure the system against it, rather than against the fluency of its output.

What good practice looks like

For a research group or an industrial lab, five habits separate productive use from expensive theater.

Define the verifier before the generator. If you cannot say precisely how a candidate will be checked, you are not searching, you are sampling. Write the check first; it constrains the whole design.

Keep the surrogate honest. A learned surrogate is only as good as its validation on held-out ground truth, and its errors are not random — they concentrate in exactly the regions of space the search will be pushed toward, because that is what optimization does. Every surrogate needs a periodic re-validation schedule against real measurement, not a one-time accuracy number.

Log everything, including the failures. Negative results are the training data for the next loop, and in a closed-loop lab they are the majority of the data. Systems that discard failed experiments lose the most informative samples they produce.

Treat citations as a separate verification problem. Language models produce plausible references at a rate that makes manual checking mandatory. This is the failure mode most likely to embarrass an organization, because it is visible, mechanical, and entirely preventable.

Report the counterfactual. The honest claim about a model-assisted result is not that the model discovered something; it is that the discovery happened faster, or was found in a space a human would not have searched, relative to a stated baseline. Without that comparison, the result is a demonstration, not evidence.

My read, in one paragraph: models are already excellent instruments in science and only incidentally authors. The instruments that work are verifier-shaped, surrogate-shaped, or closed-loop-shaped, and all three reduce to the same idea — use the model where evaluation is cheap and reliable, and use humans where it is not. That is not a limitation we will outgrow by scaling. It is a description of where knowledge comes from, and the fields that internalize it early will accumulate results while everyone else accumulates drafts.

The workflow that actually helps

Setting aside discovery claims, there is a mundane set of practices that reliably makes a research group faster, and they are worth listing because they are where most of the value has been delivered so far.

Literature triage with verification. Models are excellent at finding candidate papers and terrible at remembering what they said. The workflow that works is: use the model to propose a reading list, retrieve the actual documents, and read the retrieved text rather than the model's summary of it. Every citation gets checked. Every claim about a paper gets traced to a passage. This is slower than trusting the summary and it is the difference between a useful assistant and a fabrication generator, and the failure mode — a confident citation to a paper that does not say the thing — is the one most likely to damage a career.

Code and analysis, with artifacts. Generating analysis scripts, test harnesses, plotting code, and data-cleaning pipelines is where the productivity evidence is strongest and least controversial, because the output is executable and therefore checkable. The discipline is to require that everything be runnable from a clean environment with pinned dependencies, which the model can also produce if you ask, and which makes the assistance auditable.

Surrogate-driven design loops. The pattern from the protein and materials work: propose many candidates with a model, screen them, and test a small number physically. The engineering investment is in the screening step, which must be validated against ground truth on a schedule, because optimization will find and exploit whatever error the surrogate has.

Adversarial review before human review. Use a model to attack a draft — check every citation, look for missing baselines, ask what evidence would falsify the claim, look for the statistical error. Criticism is cheaper to verify than generation, and reviewers report that papers arriving with this pass applied are materially better.

Instrument control and structured logging. Where a model drives an experiment, the set of instructions it can issue should be typed and bounded, and every issued instruction logged with its parameters. This is what makes the closed-loop patterns reproducible rather than anecdotal, and it is the same structured-output discipline applied to hardware.

The incentive problem in publications

One uncomfortable consequence deserves saying aloud, because it affects how much of the literature you should believe. Claims of machine discovery are unusually hard to evaluate, for three reasons that compound. The code and data are often incomplete, so reproduction is impossible; the baselines are often chosen by the authors, so improvements are flattering by construction; and the evaluation protocol is frequently retrospective, so contamination cannot be ruled out. Add the pressure to publish and the result is a literature in which a non-trivial fraction of the strong claims will not survive.

The practical response for a working scientist is to read with a checklist: is there an executable artifact, are the baselines the strongest available ones, was the evaluation prospective, is the claim the narrowest one the evidence supports, and does the paper report failures. Papers that pass all five are rare and worth reading twice. Papers that fail the first two are press releases with a methods section, and the fact that they may still contain a good idea does not change what can be concluded from them.

None of this is specific to AI — reproducibility debates are decades old in the empirical sciences — but language models make it easier to produce plausible work quickly, which raises the volume of both good and bad output. The communities that respond by tightening artifacts and prospective validation will accumulate reliable knowledge. The ones that respond by trusting fluency will spend the next decade arguing about results that were never solid.

Works Cited

Bi, Kaifeng, et al. "Accurate Medium-Range Global Weather Forecasting with 3D Neural Networks." Nature, vol. 619, 2023, pp. 533–538. Accessed 9 Oct. 2025.

Boiko, Daniil A., et al. "Autonomous Chemical Research with Large Language Models." Nature, vol. 624, 2023, pp. 570–578. Accessed 9 Oct. 2025.

Fawzi, Alhussein, et al. "Discovering Faster Matrix Multiplication Algorithms with Reinforcement Learning." Nature, vol. 610, 2022, pp. 47–53. Accessed 9 Oct. 2025.

Lam, Remi, et al. "Learning Skillful Medium-Range Global Weather Forecasting." Science, vol. 382, 2023, pp. 1416–1421. Accessed 9 Oct. 2025.

Lu, Chris, et al. "The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery." arXiv, 2024, arxiv.org/abs/2408.06292. Accessed 9 Oct. 2025.

Merchant, Amil, et al. "Scaling Deep Learning for Materials Discovery." Nature, vol. 624, 2023, pp. 80–85. Accessed 9 Oct. 2025.

Messeri, Lisa, and M. J. Crockett. "Artificial Intelligence and Illusions of Understanding in Scientific Research." Nature, vol. 627, 2024, pp. 49–58. Accessed 9 Oct. 2025.

Romera-Paredes, Bernardino, et al. "Mathematical Discoveries from Program Search with Large Language Models." Nature, vol. 625, 2024, pp. 468–475. Accessed 9 Oct. 2025.

Szymanski, Nathan J., et al. "An Autonomous Laboratory for the Accelerated Synthesis of Novel Materials." Nature, vol. 624, 2023, pp. 86–91. Accessed 9 Oct. 2025.

Trinh, Trieu H., et al. "Solving Olympiad Geometry without Human Demonstrations." Nature, vol. 625, 2024, pp. 476–482. Accessed 9 Oct. 2025.

Wang, Hanchen, et al. "Scientific Discovery in the Age of Artificial Intelligence." Nature, vol. 620, 2023, pp. 47–60. Accessed 9 Oct. 2025.