Deepmind's AlphaGenome Atlas maps every possible DNA change in the human genome

| Source: THE DECODER

Tags: DeepMind, AlphaGenome, genomics, variant prediction, noncoding DNA, rare disease, AlphaFold

Google DeepMind has precomputed predicted effects for all ~9 billion possible single-letter DNA mutations in the human genome, releasing the 1-petabyte AlphaGenome Atlas — 30x larger than AlphaFold's protein database — alongside a single-number variant impact score that outperforms existing tools on clinically classified noncoding variants.

Details

Google DeepMind has released the AlphaGenome Atlas, a dataset covering predictions for every one of the roughly nine billion possible single-letter changes in the human genome. Built on the AlphaGenome model introduced in 2025, which reads one-million-letter DNA stretches and predicts gene expression, regulatory binding, and splicing across hundreds of cell and tissue types, the atlas precomputes results that previously required querying the model variant by variant. The dataset spans one petabyte — more than 30 times the size of the AlphaFold protein structure database. Each mutation comes with approximately 27,000 individual prediction values on average, covering molecular processes across dozens of tissue types. The atlas targets the noncoding 98% of the genome, where disease-linked variants are most concentrated but hardest to interpret. To make the data usable in practice, DeepMind developed the AlphaGenome Variant Impact Score (AVI), a composite score combining AlphaGenome predictions with the protein model AlphaMissense and evolutionary conservation metrics — using only 18 input features, compared to 150+ for established tools like CADD. In benchmarks on clinically classified variants, AVI outperformed existing tools particularly in noncoding regions. The team demonstrated the atlas's utility in a real epilepsy case, identifying a previously overlooked variant as the likely causal mutation. This release accelerates rare-variant diagnostics and genetic research by providing a precomputed lookup rather than per-query model inference, with particular promise for rare disease diagnosis where most variants sit in noncoding regions.