LifeSciBench, OpenAI's benchmark on life sciences research
OpenAI releases LifeSciBench, a benchmark of 750 tasks where GPT-Rosalind scores 36% and GPT-5.5 scores 25.7% in evaluating AI for life sciences research.
OpenAI releases LifeSciBench, a benchmark designed to measure the real utility of models in life sciences research, and not just their ability to answer biology questions. The idea is to stick to real-world work: interpreting incomplete evidence, reconciling contradictory results, designing experiments, assessing risk, and deciding next steps under uncertainty.
The setup is extensive: 750 tasks authored by 173 PhD-holding scientists with pharmaceutical industry experience, distributed across seven workflows and seven biological domains, then scored using over 19,000 rubric criteria. More than half require the use of attached documents, figures, tables, or sequence files, and an independent review by 453 experts validated its relevance.
The main takeaway lies in the difficulty: absolute scores remain modest; the benchmark is far from saturated. Regarding its own models, OpenAI reports that GPT-Rosalind achieves a 36% success rate compared to 25.7% for GPT-5.5, with progress in scientific communication and preclinical translation, but persistent weaknesses in experimental design, analysis, and anything requiring an exact output.