AI Tools and Scientific Innovation

Determining a protein’s structure—the three-dimensional shape that determines how it functions and how drugs might interact with it to influence disease pathways—has historically required months or even years of costly laboratory experiments. In July 2021, virtually overnight, Google DeepMind’s AlphaFold2 algorithm generated predicted structures for hundreds of thousands of proteins and made them freely available in a public database. This represented one of the most discrete and well-defined AI shocks to any scientific field to date.
In How Artificial Intelligence Shapes Science: Evidence from AlphaFold (NBER Working Paper 35143), Ryan R. Hill and Carolyn Stein study how AlphaFold2’s release affected both the narrow field of experimental structural biology and the downstream scientific and pharmaceutical research that depends on protein structures. They draw on data spanning the period 2017–2024 from several sources: the Protein Data Bank (PDB), which records experimental protein structures; the AlphaFold Protein Structure Database; UniProt’s manually curated Swiss-Prot subset linking roughly 570,000 proteins to a categorized scientific literature; and ChEMBL, which records early-stage pharmaceutical bioactivity experiments.
The release of Google DeepMind’s AlphaFold2 expanded scientific research on previously unstudied proteins.
Despite the algorithm’s ability to predict protein structures at near-experimental accuracy and negligible cost, the rate of experimental structure deposits in the PDB did not decline after July 2021. Publication rates in top general-interest journals similarly remained stable, and there is little evidence that experimentalists shifted toward proteins that AlphaFold predicts less accurately—a pattern that would be expected if the AI were substituting for experimental work in easier cases.
Instead, the researchers document a complementarity effect: AlphaFold predictions are being used as starting-point templates in the dominant computational method for converting X-ray crystallography data into three-dimensional structures. Among proteins with no experimentally solved structural neighbor, the probability of using starting points rose by 21 percentage points post-AlphaFold. Over 50 percent of such structures were using AlphaFold-derived templates by the end of the sample period.
The researchers next consider the impact of AlphaFold on related areas of protein science research, where new structure predictions might unlock insights where experimental structure determination had previously been a bottleneck. Because only about 8 percent of Swiss-Prot proteins had an experimentally solved structure before AlphaFold, the researchers compare research activity on previously unsolved proteins to that on previously solved proteins. They find that the volume of non-structure publications about previously unsolved proteins increased by between 16 and 25 percent relative to those about solved proteins after AlphaFold was released, and by more than 35 percent by 2024. By 2024, papers about previously unsolved proteins cite core AlphaFold references at an approximately 50 percent higher rate than papers about solved proteins.
Despite the evident changes in research activity, the researchers do not find any shift in early-stage pharmaceutical R&D. ChEMBL bioactivity data—which record small-molecule binding experiments against protein targets—show no statistically significant increase in activity targeting previously unsolved proteins. This is consistent with a bottleneck model in which basic biological validation of a protein’s disease relevance must accumulate before drug discovery investment occurs. The researchers suggest that pharmaceutical applications of AlphaFold’s structural information may materialize in the future as the foundation of basic research matures.