Google DeepMind released AlphaGenome Atlas on September 8, a database that precomputes the molecular impact of roughly 9 billion possible single-letter changes across the human genome. The dataset runs to 1 petabyte and is free for academic research through a web portal that requires no coding.
Charting the 98 percent of DNA we cannot read
Human DNA consists of about 3 billion base pairs, but only around 2 percent of it codes for proteins. The remaining 98 percent, known as the non-coding genome, governs when genes switch on and off, and much of its behavior is still poorly understood. Most variants linked to disease and human traits sit in this region, which has made interpretation a persistent bottleneck for genomics.
Google DeepMind released AlphaGenome in June 2025, an AI model that predicts molecular behavior such as gene expression, splicing and chromatin accessibility directly from DNA sequence. Until now researchers mostly queried it one variant at a time, which gave no genome-wide perspective. Testing 9 billion variants in a laboratory is not realistic. So the company ran AlphaGenome ahead of time across the entire genome and packaged the output as something researchers can search.
Compressing variant importance into a single AVI score
At the center of Atlas sits the new AlphaGenome Variant Impact (AVI) score. It merges predictions from AlphaGenome, which covers non-coding regions, with AlphaMissense, the company's model for protein-altering variants, and reports them as one number. Because the same yardstick applies to coding and non-coding regions alike, researchers can narrow thousands of candidates down to the ones most likely to matter.
The score can also be taken apart. Each AVI value is decomposed into contributions from interpretable categories such as chromatin accessibility, splicing and sequence conservation, so it is possible to trace why a variant was flagged. Atlas also ships a catalog of more than 2,500 recurring sequence motifs and their genomic locations. Per-variant molecular predictions span hundreds of human and mouse cell types and tissues.
Thirty times the size of the AlphaFold Database
One petabyte is more than 30 times the size of the AlphaFold Database, the company's protein structure resource. When the AlphaFold Database expanded in 2022, it grew from roughly 190,000 experimentally determined structures to more than 200 million predicted ones, and its portal made large-scale structural analysis usable for researchers who write no code. AlphaGenome Atlas is a clear attempt to repeat that outcome on the genomic side.
Early results in rare disease and the UK Biobank
External collaborators presented their findings alongside the release. Laura Covill and colleagues at the Broad Institute in the United States applied the AVI score to unsolved rare disease research under the GREGoR Consortium, prioritizing candidates that earlier work had passed over. The search surfaced a variant in a gene called DNM1, predicted to create an incorrect splice site that abnormally extends the resulting protein. DNM1 is strongly associated with epileptic encephalopathy, and experimental screens confirmed the prediction.
Gareth Hawkes of the University of Exeter in the United Kingdom applied Atlas to whole-genome data from more than 54,000 UK Biobank participants. By grouping rare variants according to their predicted molecular effects, he detected 22 percent more non-coding associations that had previously been buried in statistical noise. The analysis also pinpointed regulatory variants governing levels of proteins such as PLA2G7, which is linked to aging, and EGLN1, a cellular oxygen sensor. Narrowing to the top 1 percent of non-coding variants by predicted impact, he identified 19 genomic regions associated with body mass index.
Julia Zeitlinger and colleagues at the Stowers Institute for Medical Research in the United States used the motif data to separate transcription factors that only open up DNA from those that also switch genes on and off.
Availability and caveats
AlphaGenome Atlas is accessible through the web portal, through the AlphaGenome API, and as a skill inside Google Antigravity, the company's agentic development environment. Web access covers non-commercial use, with commercial access on Google Cloud planned. The underlying AlphaGenome model is already published on GitHub for academic use, with commercial access via Model Garden on Google Cloud.
The company states plainly that information from Atlas is not a substitute for professional medical advice, diagnosis or treatment, and that AlphaGenome has been neither validated nor approved for clinical use. It is a tool for speeding up research prioritization, and a prediction is not a conclusion.
Summary
Google DeepMind has released AlphaGenome Atlas, a 1-petabyte database that precomputes the molecular effects of 9 billion possible single-nucleotide variants in the human genome. The AVI score puts coding and non-coding regions on the same scale, letting researchers rank enormous candidate lists quickly. Concrete results are already in hand from rare disease investigation and UK Biobank analysis, and the step of guessing where to look before running experiments should get considerably shorter.
Note: the thumbnail image is AI-generated and illustrative.
