OpenAI has released a field report covering eight projects in which researchers used coding agents to modernize scientific computing software[1]. The projects span genomics and other data-rich fields, with five relying on Codex alone and three combining Codex with Claude Code. The report is a hands-on look at how much of the maintenance burden on research software, long neglected for lack of engineering hands, agents can actually take on.

Code Written for a Paper Ends Up as Research Infrastructure

Scientific computing underpins both academia and industry, but the report starts from a simple observation: software has not kept pace with the rate at which data is generated[1]. Many widely used research tools began life as code accompanying a paper, written by small academic teams without software engineering backgrounds and with little time for packaging, testing, optimization, or long-term support.

What remains is research infrastructure that is slow, fragile, and in constant need of repair. Citing prior studies of published research code and omics tools, OpenAI notes that such software often fails to install in a fresh computing environment or to run as documented[1]. Researchers lose substantial time to configuration and debugging, and the pace of discovery slows accordingly.

Eight Projects: Five with Codex Alone, Three with Claude Code

The field report gathers eight projects, primarily in the life sciences, as case studies written by the teams behind each one[1]. The work ranges from routine maintenance and targeted optimization to large-scale language migrations and GPU-native redesigns.

One example is the modernization of cyvcf2, a Python library for reading and writing genomic variant data. GPT-5.5 replaced the library's legacy build and packaging system with a modern, unified process covering installation, testing, and release[1]. cyvcf2 is a fast VCF parser that wraps htslib in Cython, developed by Brent Pedersen and Aaron Quinlan at the University of Utah and published under the MIT license[2][3]. It grew out of the need for speed that existing libraries could not deliver on studies yielding tens of millions of variants[2].

Pedersen, the author of cyvcf2, is quoted in the report: "With coding agents, it's quite easy to go fast; for now, to go far in science, there's still a need for expert guidance, understanding, taste, and care."[1]

The Bottleneck Has Moved from Implementation to Verification

A theme running through the contributions is that the researcher's role shifted from implementation to verification and orchestration[1]. People specify what to build, define how correctness will be measured, and decide when a project is ready to ship. Scientific direction and the quality bar stay in human hands, while agents supply the velocity.

At the same time, agents handled well-scoped requests effectively but could not reliably judge whether their output was scientifically valid[1]. They frequently expressed confidence even when their work contained clear errors. Human reviewers therefore needed external references or measurable acceptance targets: exact output agreement, parity with an existing tool, appropriate statistical behavior, or answers established in advance using simulated data.

The process also tended to run in stages rather than one shot. Teams broke broad goals into smaller changes and refined the work against intermediate benchmarks and test systems. Initial implementations arrived quickly, but resolving edge cases and subtle numerical differences took far longer, making the last mile the heaviest part of the job[1].

Who Takes Over Maintenance

When implementation costs fall, near-duplicate rewrites become easy to produce. That fragments users and spreads thin the expert attention needed to keep any one tool reliable, which the report flags as a real risk[1]. Mature software carries undocumented conventions, compatibility requirements, and accumulated user trust that porting source code alone cannot reproduce.

The eight projects took different paths. Changes to MHCflurry and cyvcf2 were folded back into their original upstream projects, while rustar-aligner moved under new community stewardship because the original project had been abandoned[1]. Where existing maintainers can be involved, coordination should start as early as possible; where a separate implementation is necessary, it needs a clear owner and a credible maintenance plan. Without that, today's modern rewrite simply becomes tomorrow's abandoned code.

Summary

OpenAI's field report uses eight case studies to show that coding agents such as Codex are already lowering the cost of maintaining, migrating, and optimizing scientific software. Judging whether the results are scientifically valid remains a human task, and the bottleneck has moved from implementation to verification. The question of who will keep owning these tools only grows heavier as agents get faster.

Source[1]: https://openai.com/index/scientific-computing-agentic-ai

Source[2]: https://academic.oup.com/bioinformatics/article/33/12/1867/2971439

Source[3]: https://github.com/brentp/cyvcf2