On June 30, 2026, OpenAI released GeneBench-Pro, a research-level benchmark that measures how well AI can exercise judgment on hard problems in computational biology[1]. It spans 129 problems across genomics, quantitative biology, and translational medicine, and rather than testing rote knowledge, it asks the kind of judgment a scientist makes when deciding which analysis to run on ambiguous data. Even OpenAI's strongest model, GPT-5.6 Sol, reached a pass rate of only 28.7% (31.5% with Pro mode enabled)[1].
A Benchmark That Tests Judgment, Not Just Answers
Scientific data rarely arrive with instructions. Researchers must repeatedly judge whether a pattern reflects biology or noise, whether the data can support the question being asked, and what to do next[1]. AI has become capable of executing complex analyses, but real research depends less on following a predefined workflow and more on these higher-order judgments.
GeneBench-Pro is designed to measure that capability head-on[1]. OpenAI calls it "research taste" and defines it as the chain of judgment calls that shapes an analysis: which questions the data can support, and how early diagnostics should change the modeling approach. Each problem gives the model a realistic, messy dataset, brief experimental context, and a target value tied to a downstream decision. The model must explore the data itself, choose an appropriate method, iterate through experimentation, and produce a final answer.
A Test Design Covering 129 Problems Across 10 Domains
GeneBench-Pro contains 129 problems across 10 domains and 21 sub-domains[1]. The mix includes 21 problems in population genetics, 26 in clinical, pharmacogenomics, and diagnostics, and 17 each in statistical genetics, quantitative genetics, and regulatory omics, reaching as far as cancer genomics and forensic genetics.
A key feature is that each problem is built from synthetic data[1]. Because OpenAI knows the full causal structure and simulates the data-generating process, it can tune difficulty and grade answers deterministically against known targets. Plausible but incorrect analyses are verified to fail, and the problems are audited to ensure they cannot be solved through shortcuts or by matching an author's preferences. In addition, 82 of the 129 questions were sent to external domain experts—graduate students, postdoctoral researchers, industry scientists, and professors—to check each problem's realism and the validity of its answer. At launch, 10 representative questions are open-sourced on Hugging Face, and a 50-question subset will be provided to the third-party evaluator Artificial Analysis for independent benchmarking[1][2].
Fewer Than a Third Solved at Best, Yet Rapid Progress
In evaluation, OpenAI's top model GPT-5.6 Sol scored 28.7% at the highest reasoning level, and 31.5% with Pro mode enabled[1]. That is a large jump considering that when the original GeneBench was being built, the then-leading GPT-5 scored below 5%. OpenAI notes that at this pace the benchmark could be saturated by the end of the year.
Performance also rose as more test-time compute was applied[1]. At the lowest reasoning level GPT-5.6 Sol lands in the single digits, but at the highest level it solves nearly six times as many problems as GPT-5.2 while using about two-thirds as many tokens. OpenAI also reports that the gap with leading open-source models such as GLM 5.2 is larger than coding benchmarks would suggest, indicating that open-source models are more specialized for coding than for broad reasoning. Competitor models, it says, at best matched the corresponding GPT model of the same period, and many fell well short.
20 to 40 Hours for an Expert, a Few Dollars for AI
The weight of each problem is also quantified[1]. Reviewers estimated that a typical GeneBench-Pro problem would take a human expert around 20 to 40 hours. Even at a conservative 200 USD (about 33,000 yen) per hour, the human labor cost per problem reaches thousands of dollars (on the order of hundreds of thousands of yen). By contrast, the AI inference cost is only a few dollars (a few hundred yen) per problem. Current AI agents are still too unreliable to replace human experts, but this large cost gap shows that even partial automation could generate meaningful economic and scientific value.
※1 USD = 163 JPY (as of June 30, 2026)
Behind this lies a shift in biology's bottleneck[1]. The cost of generating data, such as genome sequencing, has fallen dramatically, and the limiting factor is moving from sample collection to the downstream computation and analysis that turn vast data into actionable insight. Because drug targets with genetic support are more likely to lead to approved treatments, models that can reliably perform analyses once handled by teams of human experts could speed up hypothesis triage and target follow-up, reshaping the research cycle.
Summary
GeneBench-Pro is a new benchmark that measures whether AI can choose the correct analytical path in computational biology—a test of scientific judgment[1]. Even the top GPT-5.6 Sol solves fewer than a third of the problems, exposing a weakness: models can make partial progress but struggle to close the inferential loop. At the same time, the jump from GPT-5's sub-5% shows that AI's scientific reasoning is steadily advancing. It is drawing attention as an example of AI entering scientific research not to replace expert judgment, but to accelerate the work.
Source: https://openai.com/index/introducing-genebench-pro/
Source: https://openai.com/index/genebench-pro/case-studies/
