AI Tools and Scientific Innovation

08/01/2026
Summary of working paper 35143
Featured in print Digest

This figure is a bar chart titled "Public Insurance Expansion and Insurance Coverage by Income Group," showing the estimated effects of the 2014 ACA public insurance expansion on insurance coverage across different income groups. The y-axis is unlabeled with a title but shows values ranging from 0.0 percentage points (pp) to −1.0 pp. The x-axis is labeled "Income as a percentage of the federal poverty level" and includes four income groups: 0–150%, 151–250%, 251–400%, and 401–500%. The legend identifies two bar colors: blue for the change in share of the adult population uninsured per 1 percentage point increase in the public insurance rate, and gray for the change in share of the adult population with employer-sponsored insurance per 1 percentage point increase in the public insurance rate. The chart shows that for the lowest income group (0–150% of the federal poverty level), the uninsured rate declines by close to 1.0 pp while employer-sponsored coverage declines only slightly; as income rises across the groups, the decline in the uninsured rate becomes smaller, while the decline in employer-sponsored insurance becomes larger, with the 401–500% group showing the largest drop in employer-sponsored coverage (nearly −0.7 pp) and a smaller drop in the uninsured rate (around −0.4 pp), suggesting that public insurance expansion increasingly crowds out employer coverage at higher income levels. A note on the figure reads: Estimates relate to the 2014 ACA expansion of public insurance through Medicaid and subsidized Marketplace plans. The source line reads: Researchers' calculations using data from the American Community Survey.

Determining a protein’s structure—the three-dimensional shape that determines how it functions and how drugs might interact with it to influence disease pathways—has historically required months or even years of costly laboratory experiments. In July 2021, virtually overnight, Google DeepMind’s AlphaFold2 algorithm generated predicted structures for hundreds of thousands of proteins and made them freely available in a public database. This represented one of the most discrete and well-defined AI shocks to any scientific field to date.

In How Artificial Intelligence Shapes Science: Evidence from AlphaFold (NBER Working Paper 35143), Ryan R. Hill and Carolyn Stein study how AlphaFold2’s release affected both the narrow field of experimental structural biology and the downstream scientific and pharmaceutical research that depends on protein structures. They draw on data spanning the period 2017–2024 from several sources: the Protein Data Bank (PDB), which records experimental protein structures; the AlphaFold Protein Structure Database; UniProt’s manually curated Swiss-Prot subset linking roughly 570,000 proteins to a categorized scientific literature; and ChEMBL, which records early-stage pharmaceutical bioactivity experiments.

The release of Google DeepMind’s AlphaFold2 expanded scientific research on previously unstudied proteins.

Despite the algorithm’s ability to predict protein structures at near-experimental accuracy and negligible cost, the rate of experimental structure deposits in the PDB did not decline after July 2021. Publication rates in top general-interest journals similarly remained stable, and there is little evidence that experimentalists shifted toward proteins that AlphaFold predicts less accurately—a pattern that would be expected if the AI were substituting for experimental work in easier cases.

Instead, the researchers document a complementarity effect: AlphaFold predictions are being used as starting-point templates in the dominant computational method for converting X-ray crystallography data into three-dimensional structures. Among proteins with no experimentally solved structural neighbor, the probability of using starting points rose by 21 percentage points post-AlphaFold. Over 50 percent of such structures were using AlphaFold-derived templates by the end of the sample period.

The researchers next consider the impact of AlphaFold on related areas of protein science research, where new structure predictions might unlock insights where experimental structure determination had previously been a bottleneck. Because only about 8 percent of Swiss-Prot proteins had an experimentally solved structure before AlphaFold, the researchers compare research activity on previously unsolved proteins to that on previously solved proteins. They find that the volume of non-structure publications about previously unsolved proteins increased by between 16 and 25 percent relative to those about solved proteins after AlphaFold was released, and by more than 35 percent by 2024. By 2024, papers about previously unsolved proteins cite core AlphaFold references at an approximately 50 percent higher rate than papers about solved proteins.

Despite the evident changes in research activity, the researchers do not find any shift in early-stage pharmaceutical R&D. ChEMBL bioactivity data—which record small-molecule binding experiments against protein targets—show no statistically significant increase in activity targeting previously unsolved proteins. This is consistent with a bottleneck model in which basic biological validation of a protein’s disease relevance must accumulate before drug discovery investment occurs. The researchers suggest that pharmaceutical applications of AlphaFold’s structural information may materialize in the future as the foundation of basic research matures.