OpenAI has released LifeSciBench, a new benchmark for measuring how useful AI can be in real life sciences research[1]. It consists of 750 tasks written by Ph.D.-level scientists with hands-on drug discovery experience, and rather than testing simple factual recall, it evaluates the harder work of research itself, such as interpreting evidence and designing experiments. OpenAI's life sciences model GPT-Rosalind posted gains over the previous generation, while the results also made clear where AI still falls short.
What LifeSciBench Measures
Many earlier life sciences evaluations leaned toward narrow knowledge questions or clean prediction problems with tidy reference answers[1]. Real research, however, is a layered process: interpreting incomplete evidence, reconciling conflicting results, designing difficult experiments, and deciding what to do next under uncertainty. LifeSciBench focuses on whether AI can handle this realistic complexity.
The benchmark contains 750 tasks spanning seven research workflows and seven biological domains[1]. The workflows are evidence handling, analysis, design and optimization, scientific reasoning, validation and operations, translation (bench-to-bedside), and scientific communication. Each task is framed like a request to a knowledgeable collaborator, providing a scientific prompt and relevant materials and asking for a free-response answer.
The task design reflects the complexity of real research[1]. Overall, 79 percent of tasks require multiple reasoning or decision-making steps, averaging four steps each. There are 1,062 attached artifacts, including figures, PDFs, sequence files, and chemical structure files, and 53 percent of tasks require interpreting at least one of them to answer.
Built by 173 Scientists and Validated by 453 Reviewers
A defining feature of LifeSciBench is the depth of expert involvement in its creation[1]. The tasks were written by 173 scientists with Ph.D.-level training and biotechnology or pharmaceutical industry experience. Accepted tasks underwent an average of six automated review cycles and at least two rounds of expert review, and only those with more than 90 percent agreement among reviewers in the relevant domain were kept.
Grading uses a detailed, task-specific rubric[1]. Each breaks the expected answer down into specific scientific claims, calculations, judgments, and justifications, totaling 19,020 criteria, or an average of 25 per task. The design assesses not only whether the final answer is correct, but whether the model reaches it through a scientifically valid and practically useful path.
The benchmark's validity was confirmed through an independent review by 453 reviewers who did not write the tasks[1]. Of these, 97 percent held a Ph.D. or equivalent, with an average of 12 years of experience and 14 peer-reviewed publications. They assessed whether tasks reflected real research and tested scientific reasoning appropriately, and agreement exceeded 96 percent in every category.
How the Latest Model "GPT-Rosalind" Scored
OpenAI compared its life sciences model GPT-Rosalind[2] with the previous-generation GPT-5.5[1]. Two metrics are used: pass rate, the share of tasks meeting the per-task success threshold of 70 percent, and score, the average rubric reward that grants partial credit. On overall exact pass rate, GPT-Rosalind improved to 36.1 percent, up from 25.7 percent for GPT-5.5.
The largest gains came in scientific communication and translation[1]. The scientific communication pass rate rose from 56.3 percent to 71.1 percent, though OpenAI notes this category should be read cautiously because it has only 9 tasks. Translation, the bench-to-bedside process, improved from 36.8 percent to 57.7 percent.
At the rubric level, GPT-Rosalind scored 44.7 percent on tasks requiring expert-useful, actionable outputs (versus 29.1 percent for GPT-5.5) and 44.8 percent on tasks involving uncertainty and caveat handling (versus 29.3 percent)[1]. The pattern suggests AI performs better when a task has a clear evidence boundary and calls for structured scientific judgment.
Where AI Still Struggles
On the other hand, AI weakness stands out on artifact-heavy and design-oriented tasks[1]. The design, optimization, and prediction workflow reached only a 30.7 percent pass rate for GPT-Rosalind, and analysis just 30.3 percent. When figures or sequence files are involved, performance drops from 45.1 percent on text-only tasks to 28.1 percent on tasks with artifacts.
Answer format also matters[1]. GPT-Rosalind reached only 14.8 percent on tasks requiring exact numeric answers and 24.0 percent on sequence or structure outputs. Because work such as CRISPR donor design requires outputs precise enough to use directly, this weakness carries real scientific significance.
OpenAI is careful to bound its claims[1]. LifeSciBench cannot fully reproduce real research, which iterates and revises hypotheses over time, and strong scores should be read as task-level capability rather than downstream research impact. The next step, OpenAI says, is to study whether AI actually accelerates discovery in live research settings. It is also worth noting that these benchmarks are designed by OpenAI itself, and even though external experts grade LifeSciBench, the results have not been independently reproduced[3].
Summary
LifeSciBench is a new benchmark that uses 750 expert-written tasks to measure how useful AI can be in life sciences research[1]. GPT-Rosalind lifted the overall pass rate from 25.7 percent to 36.1 percent, with particularly strong progress in scientific communication and translation[2]. At the same time, it still struggles with artifact interpretation and precise design, and absolute pass rates remain modest. The benchmark reflects both how AI is becoming a research partner and the limits of that partnership.
Source: https://openai.com/index/introducing-life-sci-bench
