Beyond the Static Snapshot: AI Diffusion Models Learn to Predict How Proteins Move
Research teams at Cambridge and UCSF have independently built diffusion-based AI models that predict how proteins move between shapes over time, not just their resting structure — a step that could unlock drug targets long considered undruggable.
Since AlphaFold reshaped structural biology by predicting protein shapes with striking accuracy, one gap has remained stubbornly open: proteins are not frozen sculptures, they move. This month, separate research groups at the University of Cambridge and UC San Francisco published complementary approaches that push AI beyond static structure prediction and into the far harder territory of protein dynamics.
Key takeaways
- New diffusion-based AI models predict the range of shapes, or “conformations,” a protein moves through over time — not just a single resting structure.
- The underlying technique borrows directly from image-generation diffusion models, adapted for molecular motion.
- Many drug targets shift between multiple shapes, so a single static structure can miss the binding site that actually matters.
- Cambridge and UCSF teams arrived at complementary approaches independently, a sign the field is converging on diffusion modeling as the right tool for this problem.
The Limits of a Single Snapshot
When DeepMind’s AlphaFold first demonstrated it could predict a protein’s three-dimensional structure directly from its amino acid sequence, it collapsed a problem that had consumed careers into something closer to a computational lookup. Structural biology labs that once spent years crystallizing a single protein could suddenly generate a plausible structure in hours.
But a static structure is, in a real sense, a single frame pulled from a movie. Proteins are dynamic molecular machines. They flex, rotate, open, and close as they carry out their biological functions, and many of the proteins pharmaceutical researchers care most about — particularly those implicated in cancer, neurodegeneration, and infectious disease — are only biologically active in certain transient shapes. A drug designed against the wrong conformation can bind poorly, or not at all, even if the underlying static structure prediction was essentially correct.
This gap has quietly limited how useful structure-prediction tools can be for real drug discovery pipelines, where identifying a druggable binding site often depends on catching a protein in exactly the right transient shape.
Borrowing From Image Generation to Model Molecular Motion
The new wave of research tackles this by adapting diffusion models — the same class of generative AI technique behind modern image-generation systems — to model molecular dynamics instead of pixels. In image generation, a diffusion model learns to gradually transform random noise into a coherent picture. Applied to molecular structures, the same underlying mathematics can be repurposed to generate a distribution of plausible protein conformations, rather than a single best guess.
That distinction matters enormously for drug design. Instead of receiving one static structure and hoping it happens to represent the biologically relevant state, researchers using these new models receive a spread of likely conformations, weighted by probability, giving a far more complete picture of how a protein actually behaves in solution over time.
Both the Cambridge and UCSF teams built their systems around this diffusion-based approach, arriving at broadly similar strategies independently — itself a meaningful signal, since convergent independent discovery is often a sign that a field has found the right underlying tool for a problem, rather than a single group having gotten lucky with one approach.
Why Conformational Dynamics Matter So Much for Drug Design
Static protein structures have already proven enormously valuable for identifying binding sites — the pockets on a protein’s surface where a drug molecule can attach and alter its function. But many of the most medically important protein targets are not simple, rigid structures. They are conformationally flexible, shifting between multiple stable and semi-stable shapes depending on their biological context, binding partners, or cellular environment.
A drug candidate designed against a single static snapshot risks missing the shape that matters most, or worse, being optimized against a conformation the protein rarely actually adopts in a living cell. By modeling the full distribution of a protein’s likely shapes, these new AI systems give medicinal chemists a much richer map of where a drug might actually be able to intervene — and, just as importantly, where it likely cannot.
This has knock-on implications for an entire category of previously “undruggable” targets — proteins that have long been recognized as biologically important in disease processes but have resisted conventional drug design precisely because their relevant functional shape was poorly understood or highly transient.
From AlphaFold to a Fuller Picture of Cellular Machinery
It is worth situating this work in the broader arc of AI-driven structural biology. AlphaFold’s breakthrough answered the question “what shape is this protein?” with remarkable accuracy across a huge fraction of known proteins. The dynamics-focused work emerging this year answers a different, arguably harder, question: “what range of shapes can this protein take, and how does it move between them?”
That is a meaningfully more complex modeling challenge. Static structure prediction, however difficult, is still a single-answer problem. Dynamics prediction requires modeling a continuous, time-dependent process with many plausible trajectories, which is precisely the kind of problem diffusion-based generative approaches are suited for, since they are built from the ground up to represent distributions rather than single fixed outputs.
Where This Fits in the Drug Discovery Pipeline
Pharmaceutical R&D has already begun integrating AI structure prediction tools into early-stage target identification and lead compound design. The addition of dynamics prediction is likely to slot into that same pipeline, most immediately in the earliest stages of drug discovery, where researchers are trying to decide whether a given protein target is worth pursuing at all, and if so, which region of the protein is the most promising site for intervention.
Because both research teams have published their work rather than keeping it proprietary, the techniques are likely to be picked up quickly by both academic labs and industry drug-discovery groups, following the same rapid adoption pattern that AlphaFold itself saw after its initial release.
What to Watch Next
The obvious next test for these dynamics-prediction models is validation against experimentally measured protein motion, gathered through techniques such as nuclear magnetic resonance spectroscopy and cryo-electron microscopy, both of which can directly observe some aspects of protein flexibility in the lab. How closely the AI-predicted conformational distributions match these experimental measurements will determine how quickly the pharmaceutical industry moves from cautious interest to routine adoption.
If the correspondence holds up at scale, the practical effect could be significant: a meaningful fraction of the “undruggable” proteins currently sitting in pharmaceutical target databases may turn out to be druggable after all, simply because researchers can finally see the shape that was hiding in between the frames.
How Diffusion Models Differ From Earlier Molecular Dynamics Approaches
Before diffusion models entered this space, the standard computational method for studying protein motion was molecular dynamics simulation, which models the physical forces acting on every atom in a protein over time, one tiny time-step at a time. That approach is grounded in real physics and can, in principle, be highly accurate, but it is also punishingly expensive computationally. Simulating even microseconds of real biological motion can require days or weeks of supercomputer time, which puts full molecular dynamics simulation out of reach for the kind of rapid, iterative screening that early-stage drug discovery depends on.
Diffusion-based generative models take a fundamentally different approach. Rather than simulating the physical forces driving motion step by step, they learn, from training data, the statistical distribution of shapes a given protein family tends to adopt, and then generate samples from that distribution directly. This trades some of the first-principles physical grounding of full simulation for an enormous gain in speed, since generating a plausible conformation from a trained diffusion model can take seconds rather than days. For early-stage drug discovery, where researchers need to screen many candidate targets quickly before committing resources to the most promising ones, that speed advantage matters enormously, even if the resulting predictions require experimental confirmation before being fully trusted.
What Researchers Are Watching For Next
The most immediate open question is how well these models perform on protein families that are poorly represented in existing structural databases, since diffusion models, like other generative AI systems, tend to perform best on data that resembles what they were trained on. Proteins from less-studied organisms, or unusual structural families with few existing experimental examples, represent a genuine stress test for how far these techniques generalize beyond their training distribution.
There is also a practical infrastructure question: for these tools to see wide adoption inside pharmaceutical research pipelines, they will likely need to be integrated directly into the software platforms that medicinal chemists already use day to day, rather than existing as standalone academic research code. Given how quickly AlphaFold itself moved from an academic research breakthrough into a routinely used tool across the pharmaceutical industry, a similarly fast integration path for these dynamics-focused successors would not be surprising, particularly given how directly they address a limitation that drug discovery teams have been vocal about for years.
The Bigger Picture for Structural Biology
Taken together, static structure prediction and dynamics prediction increasingly look like two complementary halves of a single, more complete computational picture of how proteins actually behave inside living cells. Neither replaces careful experimental validation, and neither eliminates the need for skilled structural biologists to interpret what these models produce. What they do offer is a dramatically cheaper and faster way to generate strong initial hypotheses about a protein’s likely behavior, hypotheses that can then be prioritized for the kind of expensive, time-consuming experimental confirmation that remains the ultimate arbiter of truth in structural biology.
For an industry that has spent decades chasing a fuller understanding of how proteins actually move and function, rather than merely how they are shaped at rest, this convergence of independent research groups on diffusion-based dynamics modeling this year may end up marking as significant a turning point as the original static structure breakthrough did several years earlier.
