Anthropic releases BioMysteryBench, a bioinformatics benchmark for Claude

Anthropic released BioMysteryBench, a bioinformatics benchmark where Claude Opus 4.6 achieved a 77.4% success rate, rivaling Genentech’s CompBioBench.

Anthropic has published the results of BioMysteryBench, a benchmark consisting of 99 bioinformatics questions designed to evaluate its models' ability to conduct scientific analyses on real and noisy datasets. The questions, drafted by domain experts, primarily focus on raw DNA or RNA sequencing data, supplemented by a few proteomics and metabolomics tasks. Each answer relies on an objective ground truth, validated experimentally or by metadata, rather than a researcher's interpretation, which distinguishes this protocol from other benchmarks like BixBench or SciGym. Of the 76 problems solved at least once by a human panel, Claude Opus 4.6 achieved a 77.4% success rate. For the 23 problems that experts could not definitively resolve, Claude Mythos Preview reached 30%, and several models already surpass the performance of a panel of five specialists. Anthropic identifies two recurring strategies employed by Claude: the direct mobilization of knowledge from training (ontologies, biological mechanisms, meta-analyses) and the convergence of multiple methods when the model hesitates on the answer to provide. However, the reliability analysis conducted by Mythos Preview highlights a notable fragility on difficult tasks: nearly half of the correct answers obtained on this subset come from a single success out of five attempts, indicating that the model reaches the solution without being able to reproduce it stably. In parallel, Genentech and Roche have published CompBioBench, a similar benchmark whose conclusions align, with a score of 81% for Opus 4.6. BioMysteryBench is publicly accessible, and Anthropic invites the community to submit other verifiable research tasks to scienceblog@anthropic.com.