AI Frontiers, part 55: Biology and protein models — what AlphaFold changed
Part 55from the AI Frontiers series · 65 parts in all
Most claims about AI transforming science are made by people who have not watched a scientist work. Protein structure prediction is the exception, and it is worth studying carefully because it is the only case so far where a machine learning model changed the daily practice of an entire scientific discipline in a measurable way. When AlphaFold2 won CASP14 in 2020, the result was not a leaderboard improvement; it was a change in what a structural biologist could assume. Work that took a graduate student months of crystallography or years of NMR suddenly started with a downloaded coordinate file. The Nobel Prize in Chemistry in 2024 acknowledged the shift, which is about as official as a scientific reallocation of labor gets.
This entry is about that transition: what the models actually do, what the second generation added, where the protein language model line diverged from the structure prediction line, and why the field's practitioners remain considerably more careful than its press releases. It is also, unavoidably, a case study in the thing this series keeps returning to — the model is the easy part, and validation is the product.
Fifty years of a hard problem
Predicting a protein's three-dimensional structure from its amino acid sequence is a combinatorial problem with a physical objective: find the lowest free-energy conformation among an astronomically large space of possibilities. Physics-based simulation, molecular dynamics, and knowledge-based potentials made slow progress for decades. The community's measuring stick was CASP, a biennial blind assessment in which predictors submit structures for targets whose experimental structures are not yet public, and independent assessors score them. That infrastructure — a blind, adversarial, recurring benchmark with expert human judgment — is a large part of why progress here was real and verifiable rather than a collection of impressive demos, and it is a model that other scientific domains are still trying to copy.
The biennial comparison has been the field's connective tissue through every methodological turn, and the assessors' reports are the record of it (Kryshtafovych et al.). By CASP13 in 2018, deep learning approaches had made the first clear jump. At CASP14 in 2020, AlphaFold2 produced predictions for a majority of targets at accuracy comparable to experimental methods for the protein backbone, and the community's response was unusually united: this had changed (Jumper et al.).
What the architecture actually contributed
The technical story matters because it explains both the model's strengths and its blind spots. AlphaFold2 combined three ideas. Multiple sequence alignments supplied evolutionary covariation — residues that mutate together are probably close in space — giving the model a strong prior over contacts without any physics. An attention-based trunk (the Evoformer) iteratively refined pairwise representations of residue relationships together with per-residue representations, effectively learning to reason about the geometry of the whole chain at once. And an equivariant structure module produced 3D coordinates in a way that respects rotation and translation, so the network never had to relearn physics that does not depend on the coordinate frame.
Two consequences followed. First, the model is extraordinarily good at the parts of the problem where evolution and geometry are informative, which is most single-chain, well-folded proteins. Second, it is weakest exactly where the training signal is thin: proteins with few homologs, intrinsically disordered regions, and anything requiring a conformational ensemble rather than a single structure. AlphaFold2 predicts a structure, not the distribution of structures a flexible protein occupies, and much of a protein's biological function lives in that distribution. Practitioners learned to read the per-residue confidence measure as carefully as the coordinates themselves, which is the closest thing to a calibrated uncertainty communication this field has produced.
The non-technical contribution was just as decisive. The code was released openly, and the predictions themselves were published as a public database covering essentially every known protein. That is what converted a result into infrastructure: a model that only its authors can run changes a paper; a database that anyone can query changes a field.
The second generation: complexes, ligands, and the openness fight
AlphaFold3 extended the framework from single chains to complexes of proteins, nucleic acids, small-molecule ligands, and ions, replacing the structure module with a diffusion process that generates coordinates directly (Abramson et al.). Predicting how a drug candidate sits in a binding pocket is one of the most valuable computations in pharmaceutical research, and independent reproduction attempts converged on the conclusion that the approach works — with notable differences in accuracy across interaction types, and best-in-class performance on protein-ligand complexes.
The reception was also a lesson in scientific infrastructure politics. AlphaFold3 shipped with a server and code that permitted non-commercial use only, drawing criticism from researchers who argued that a tool with this much scientific leverage should not be gated on employment status. Within months,academic groups produced fully open reproductions — Boltz-1 (Wohlwend et al.) demonstrated that the architecture and data pipeline were reproducible under a permissive license — and the open ecosystem has kept pace since. The episode is worth remembering whenever a lab announces a breakthrough model without weights: the half-life of a capability advantage in open science is measured in months, and the resulting competition mostly benefits the science.
The other line: language models of protein and genome
Parallel to structure prediction, a second line of work treats biological sequences as language. Protein language models trained with masked-token objectives on hundreds of millions of sequences learn representations that capture structure, function, and evolutionary relationships without any explicit geometric supervision (Rives et al.). ESMFold showed that such a model could produce structure predictions at competitive accuracy with a single-sequence input and far lower compute than alignment-based methods — trading some accuracy for a hundredfold speed advantage and, crucially, removing the dependence on homolog search (Lin et al.). ESM3 extended this to generative protein design, with the striking demonstration that the model could generate a functional protein substantially diverged from known sequences (Hayes et al.).
The design side has its own genealogy. RFdiffusion applied diffusion models to protein backbone generation and, combined with sequence design tools, produced experimentally validated binders and symmetric assemblies (Watson et al., "De novo design of protein structure and function with RFdiffusion"). The practical payoff of all this is not "designing new life"; it is narrowing a search space. A wet lab can test a few hundred designs a week. A generative model that proposes a few hundred plausible candidates instead of an unranked space of possibilities changes the ratio of experiments to hits, and that ratio is what a discovery program actually optimizes.
The same techniques are now being applied up the scale of biological organization: nucleotide transformer models over DNA (Dalla-Torre et al.), and genome-scale sequence models such as Evo, trained across DNA, RNA, and protein, which demonstrated zero-shot prediction of molecular effects and generation of coherent sequences at length (Nguyen et al.). Whether genome-scale modeling repeats the AlphaFold story or stalls at the same wall may be the most consequential open question in the area.
What actually changed in practice
Four changes are visible from outside the field. Experimental structural biology got more targeted: instead of solving structures to discover them, labs increasingly solve structures to check a prediction or to capture a state the model cannot represent. Variant interpretation acquired a tool: AlphaMissense and related work score whether a missense variant is likely pathogenic, which turns an unbounded clinical question into a ranked list for follow-up (Cheng et al.). Drug discovery timelines compressed at the front end: target structure and binding-mode hypotheses that used to require months of crystallography are now available before a project formally starts. And a generation of computational scientists re-tooled, which is the change with the longest tail — a field that once mostly used simulation now mostly fine-tunes or queries models, and its methods sections have changed accordingly.
The limits, stated as plainly as the wins: predicted structures are not experimental structures, and the difference matters most precisely where the biology is interesting; conformational dynamics, allostery, and folding pathways remain largely out of reach of single-structure prediction; and the correlation between predicted binding affinity and measured activity is weak enough that medicinal chemists treat model scores as a filter rather than a decision. AlphaFold changed what you know before you go into the lab. It did not change what the lab does.
What other fields should take from this
Three transferable lessons, none of them about transformers.
A blind, recurring, independent benchmark is what made the progress credible. CASP's design — targets held back, assessors with no stake, biennial comparison — is why the AlphaFold result was accepted within weeks rather than argued about for years. Domains without that infrastructure get arguments about methodology instead of accumulation of knowledge, which is why the benchmark is often the highest-leverage contribution available.
Publishing the predictions was as valuable as publishing the model. The database turned a capability into a utility. If your field has a shared corpus, the same move is available to you.
Releasing only a server creates a competitor, not a moat. The AlphaFold3 licensing episode produced open reproductions within months. In science, the strategic choice is not between openness and advantage; it is between being the reference implementation and being one of several.
There is a fourth lesson that is less flattering and equally useful: the field's most durable contribution may turn out to be its habits rather than its models. A culture that expects prospective validation, publishes its misses, and treats a plausible answer as a hypothesis rather than a result is a culture that can absorb a hundred mediocre models without losing its way. Institutions that lack those habits can be damaged by a single impressive demo. Biology will keep being the most instructive applied domain in this series, because its ground truth is expensive, its feedback is slow, and its practitioners are unwilling to accept a plausible answer in place of a validated one. Whatever survives contact with that standard is worth borrowing everywhere else.
Evaluating a prediction you intend to act on
The reason biology is worth studying from a software perspective is that its practitioners have developed an unusually mature response to the central problem of applied AI: a model outputs a plausible answer, and someone has to decide whether to spend real resources on it. Four habits from that response transfer directly to any domain where predictions drive actions.
Grade confidence in the units of the decision. AlphaFold's per-residue confidence is useful not because it is a probability, but because the field learned roughly what score corresponds to "trust this for a design decision" versus "this is a hypothesis." Most applied AI systems ship a single number and no idea what threshold matters. Calibrating a score against decision outcomes — rather than against accuracy — is the work that makes a prediction operational.
Separate screening from confirmation. A cheap prediction narrows the space; an expensive measurement confirms the candidate. Every successful scientific deployment of these models has this shape, and it maps exactly onto the cascade pattern of part 59: spend the expensive verification only where the cheap screen says it is worth it, and measure the screen's false-negative rate as carefully as its precision, because the false negatives are invisible by construction.
Prospective validation or nothing. Retrospective evaluation on cases the model or its trainers could have encountered is how a field convinces itself of something that later fails. The blind, prospective protocol is what gives the CASP results their weight, and the AI literature has had to relearn the same lesson repeatedly through contaminated evaluations (part 36).
Publish the misses. Structural biology databases carry predictions that were subsequently contradicted by experiment, and that record is what allows the community to reason about the model's blind spots. Systems that log only their correct outputs are not evaluated; they are marketed.
The transferable summary is a discipline rather than a technique: define what action the prediction will trigger, measure the score against that action, validate prospectively, and keep the failures. A predictive model with those four properties is trustworthy in a way that a model with better benchmark numbers and none of them is not.
Works Cited
Abramson, Josh, et al. "Accurate Structure Prediction of Biomolecular Interactions with AlphaFold 3." Nature, vol. 630, 2024, pp. 493–500. Accessed 4 Sept. 2025.
Cheng, Jun, et al. "Accurate Proteome-Wide Missense Variant Effect Prediction with AlphaMissense." Science, vol. 381, 2023. Accessed 4 Sept. 2025.
Dalla-Torre, Hugo, et al. "The Nucleotide Transformer: Building and Evaluating Robust Foundation Models for Human Genomics." arXiv, 2023, arxiv.org/abs/2301.11280. Accessed 4 Sept. 2025.
Hayes, Thomas, et al. "Simulating 500 Million Years of Evolution with a Language Model." Science, 2025. Accessed 4 Sept. 2025.
Jumper, John, et al. "Highly Accurate Protein Structure Prediction with AlphaFold." Nature, vol. 596, 2021, pp. 583–589. Accessed 4 Sept. 2025.
Kryshtafovych, Andriy, et al. "Critical Assessment of Methods of Protein Structure Prediction (CASP)—Round XV." Proteins: Structure, Function, and Bioinformatics, vol. 91, 2023, pp. 1539–1549. Accessed 4 Sept. 2025.
Lin, Zeming, et al. "Evolutionary-Scale Prediction of Atomic-Level Protein Structure with a Language Model." Science, vol. 379, 2023, pp. 1123–1130. Accessed 4 Sept. 2025.
Nguyen, Eric, et al. "Sequence Modeling and Design from Molecular to Genome Scale with Evo." Science, vol. 386, 2024. Accessed 4 Sept. 2025.
Rives, Alexander, et al. "Biological Structure and Function Emerge from Scaling Unsupervised Learning to 250 Million Protein Sequences." Proceedings of the National Academy of Sciences, vol. 118, 2021. Accessed 4 Sept. 2025.
Watson, Joseph L., et al. "De Novo Design of Protein Structure and Function with RFdiffusion." Nature, vol. 620, 2023, pp. 1089–1100. Accessed 4 Sept. 2025.
Wohlwend, Jeremy, et al. "Boltz-1: Democratizing Biomolecular Interaction Modeling." bioRxiv, 2024. Accessed 4 Sept. 2025.