AlphaGenome Atlas predicts the effect of nine billion DNA variants

Google DeepMind ouvre AlphaGenome Atlas, une base de 1 pétaoctet qui prédit l’effet moléculaire des neuf milliards de substitutions possibles dans le génome humain.

Three billion positions, with three possible substitutions at each one. That calculation produces roughly nine billion possible single-letter changes in the human genome. Google DeepMind used AlphaGenome to estimate their molecular consequences, then gathered the results into a database that researchers can browse online.

Launched on September 8, 2026, AlphaGenome Atlas is not a new model trained from scratch. It is a precomputed resource built with AlphaGenome, which was introduced in 2025 and described in a Nature paper in January 2026. Google ran the model at a massive scale so researchers would no longer need to generate every prediction individually.

The word “atlas” requires some clarification. The platform does not catalog nine billion mutations already observed in humans. It examines every theoretically possible substitution against a reference sequence. At each position containing an A, C, G, or T, the other three letters represent three potential variants.

The catalog therefore includes common changes, rare variants, and an enormous number of substitutions that may never have been observed in any population. It offers a collection of molecular hypotheses rather than an inventory of real-world genetic diversity.

Its scope is also limited to single-nucleotide variants, commonly abbreviated as SNVs. Insertions, deletions, duplications, repeat expansions, and chromosomal rearrangements are not systematically covered by these nine billion results. AlphaGenome can analyze some other types of changes through targeted queries, but Atlas should not be described as a comprehensive prediction of every possible form of genomic variation.

The underlying model processes up to one million DNA letters at a time. From that sequence, it predicts several measurements related to gene function, including gene expression, transcription initiation, chromatin accessibility, certain histone modifications, transcription factor binding, interactions between chromosomal regions, and RNA splicing.

The AlphaGenome paper describes 5,930 human genomic tracks and 1,128 mouse tracks across eleven output types. Several results reach single-base-pair resolution, while others, such as chromosomal contact maps, operate at a broader scale.

The one-million-letter input length helps the model account for regulatory elements located far from the genes they influence. It does not amount to a complete understanding of each chromosome. Interactions beyond that window, temporary changes tied to a particular cell state, and effects involving several distant regions may still be missed.

This capability is particularly relevant to researchers studying non-coding DNA. Roughly 2% of the human genome directly codes for proteins. The remainder includes sequences that control where, when, and how strongly genes are expressed. A change located far from a gene can therefore affect its activity without altering the protein itself.

These regulatory variants remain difficult to interpret. Their position does not immediately reveal which tissue is affected, which mechanism is disrupted, or which gene may be involved. AlphaGenome compares its predictions for the reference sequence and the modified sequence to estimate the resulting difference across several biological processes.

Atlas combines these results with those from AlphaMissense, another Google DeepMind system focused on changes that alter proteins. This combination produces the AlphaGenome Variant Impact score, or AVI.

AVI summarizes the predicted functional importance of a variant in a single number, whether it lies within a coding or non-coding region. Researchers can use it to rank a long list of candidates before examining the underlying results in greater detail.

That number is not a diagnosis. A high AVI score indicates that a change could strongly disrupt one or more molecular processes. It does not prove that the variant causes disease, directly measure a person’s likelihood of developing symptoms, or replace an analysis of its population frequency or inheritance within a family.

To avoid reducing interpretation to a single value, Atlas also provides detailed feature attributions. These show which components contribute to the score, such as a predicted change in splicing, gene expression, chromatin accessibility, or protein structure.

The platform also includes more than 2,500 recurring DNA motifs. These short sequences function like elements of a biological vocabulary. Some serve as binding sites for proteins that activate or suppress gene expression. Connecting a variant to a disrupted motif can help researchers develop a mechanistic explanation that can later be tested experimentally.

The full resource occupies about one petabyte, more than 30 times the reported size of the AlphaFold Database. This comparison primarily describes storage volume. It does not mean that Atlas is 30 times more accurate or automatically carries greater scientific value.

The website is designed for researchers who do not want to write code for every query. They can search for a variant, review its associated scores, and explore the surrounding genomic context. An API remains available for automated analyses, while integration with Google Antigravity is intended to support its use in agent-driven research workflows.

The “zero-code” label does not remove the need for expertise in genetics. Selecting the correct reference genome version, verifying a variant’s orientation, choosing the relevant tissue, and interpreting a molecular change still require specialist knowledge. An accessible interface lowers the technical barrier, not the biological complexity.

Google has presented several collaborations intended to demonstrate the resource’s practical value. Working with the GREGoR Consortium and researchers at the Broad Institute, scientists used the AVI score to reexamine candidate variants in rare disease cases that had remained unsolved.

The analysis highlighted a change affecting the DNM1 gene, which has already