OpenAI has unveiled "Patch the Planet," a new effort to support the maintainers of open-source software (OSS)[1]. It pairs AI-assisted vulnerability research with expert human review, and its defining feature is helping not just to find flaws but to fix them. Working with the security firm Trail of Bits, the program targets critical projects the world depends on, such as cURL, Python, and Go, and has already uncovered hundreds of issues.

A program that supports the full path "from discovery to patch"

Patch the Planet was launched together with Trail of Bits as part of Daybreak, OpenAI's cybersecurity-focused program[1][2]. AI is accelerating the discovery of vulnerabilities, but discovery alone does not protect users. Many maintainers are being asked to handle more reports faster with the same limited time and staff. Patch the Planet is designed to reduce that burden rather than add to it: before a report reaches a maintainer, security engineers review its contents and work alongside the project to develop patches and tests[1].

For the initial surge, Trail of Bits committed its entire security research organization to the effort[1]. The program also partners with HackerOne and Calif on triage, coordinated disclosure, and additional vulnerability research. Each collaboration begins with a consultation with the maintainer, aligning on where added effort is most useful, such as vulnerability validation, patch development, and improvements to build-to-release automation.

The first participating projects include cURL, NATS Server, pyca/cryptography, Sigstore, aiohttp, the Go project, freenginx, Python, and python.org[1]. These span networking, cryptography, the software supply chain, and language infrastructure, the areas that underpin a broad range of products and services. Participating projects gain access to "Codex Security," which finds and fixes vulnerabilities directly in code, along with ChatGPT Pro and API credits for development and release work[1].

A fuzzing setup built in under a day

Trail of Bits assigned dedicated security engineers to work continuously across 19 OSS projects using the AI agent "Codex" and the cyber-specialized model "GPT-5.5-Cyber"[1]. As a result, they identified hundreds of security issues, have already merged dozens of patches, and have many more in the coordinated-disclosure process.

This first sprint produced not only individual findings but reusable security infrastructure[1]: fuzzing harnesses (a technique that floods software with inputs to surface anomalies), pipelines for analyzing past CVEs (Common Vulnerabilities and Exposures), differential-testing systems, threat models, and expanded test suites. A few examples stand out.

One was a fuzzing setup built in under a day[1]. Engineers used GPT-5.5-Cyber to run Codex repeatedly toward a goal, assembling a complete setup covering dozens of entry points, variant builds, and multiple environments. Building the same thing by hand would normally take at least several weeks.

Another was a pipeline for finding "variants" of known vulnerabilities[1]. It ingests past CVEs, extracts vulnerability patterns, searches target code for similar flaws, and routes the results through dedicated judging agents that remove duplicates and filter out false positives. This turns years of accumulated vulnerability history into a search strategy reusable across projects.

Differential testing also delivered results[1]. Multiple implementations of the same specification should respond identically to the same input; when their behavior diverges, one of them may contain a bug. This used to require writing custom glue code for each implementation, but with Codex generating and refining that code, work that normally takes weeks or months was compressed into days.

Vulnerabilities found from Linux to Chrome and Safari

The findings spanned every layer of software[1]. In the Linux kernel, GPT-5.5-Cyber identified security-relevant areas across more than 30 million lines of code and generated 8 proof-of-concept (PoC) exploits leaking kernel pointer information and 24 PoCs for local privilege escalation. These are a subset that were auto-generated; hundreds of issues were identified in total.

In OpenBSD, the team found a 23-year-old use-after-free flaw in the kernel's System V semaphore implementation[1]. It could let an ordinary user escalate to administrator (root) privileges, and researchers reproduced and confirmed it. In FreeBSD, they confirmed 34 vulnerabilities and produced 7 privilege-escalation PoCs.

In the lightweight DNS software dnsmasq, Codex Security independently uncovered vulnerable patterns matching 4 of the 6 CVEs later fixed in the patched release "2.92rel2" (CVE-2026-4890, 4891, 4892, and 5172)[1]. A denial-of-service (DoS) technique that Calif identified with Codex, dubbed "HTTP/2 Bomb," affects major HTTP/2 implementations such as NGINX, Apache, IIS, and Pingora; analysis suggested more than 880,000 internet-facing sites were running affected software.

Major browsers also yielded results[1]. In Google's Chrome, the team reported 5 exploitable vulnerabilities in the JavaScript engine "V8," 3 of which were identified and fixed within days of being introduced. In Apple's Safari, about a week of work on WebKit surfaced more than 10 exploitable vulnerabilities. In Mozilla's Firefox, a WebAssembly vulnerability (CVE-2026-8390) found with GPT-5.5 during safety evaluations was patched by Mozilla two days before the hacking contest "Pwn2Own Berlin." As a result, 5 of the 6 registered Firefox entries withdrew, and no attack against Firefox succeeded at the event.

A design that assumes human review

A central principle of the effort is that a human always reviews findings before they reach a maintainer[1]. Frontier models are highly capable at finding and fixing vulnerabilities, but they also produce large volumes of false positives that can add to maintainers' already swelling workloads. In Patch the Planet, Trail of Bits researchers reproduce the evidence, cross-check it against project documentation and threat models, remove duplicates, reassess severity, and prioritize confirmed vulnerabilities for remediation.

Patch development and submission also follow each maintainer's preferences, and maintainers retain control over which patches are applied and how disclosure is handled[1]. The framing is that AI's broad search supports, rather than replaces, human filtering and final judgment. OpenAI says it is sharing early results for now while withholding attack mechanics and details still under disclosure, and plans to publish deeper technical reports on individual findings and methods once fixes are in place. Maintainers can apply to take part in the effort.

Summary

"Patch the Planet" combines AI-assisted vulnerability research with expert review to support OSS maintainers from discovery through patching[1]. Together with Trail of Bits, it uncovered hundreds of issues across 19 projects, with concrete results ranging from Linux-kernel privilege escalation and a 23-year-old OpenBSD flaw to vulnerabilities in Chrome and Safari. By filtering out false positives and leaving the final judgment to humans, the program aims to turn AI's search power into support for maintainers rather than an added burden.

出典:https://openai.com/index/patch-the-planet

出典:https://openai.com/index/daybreak-securing-the-world