OpenAI released the Agents API in public beta to all developers on September 10, making the harness and infrastructure behind Codex and ChatGPT available directly. Specifying four things, the task, the model, the tools, and the environment, is enough to create a production-ready agent. There is no separate fee for the API itself; you pay only for the tokens and tools your agents consume.

What developers no longer have to build is the plumbing

Running an agent in production tends to stall on everything around the model call rather than the call itself: repacking context when a session hits its limit, managing tool definitions that pile up, keeping a workspace for intermediate artifacts, and holding a session alive for hours without falling over. Every team has been writing that plumbing itself.

The Agents API bundles that layer into something OpenAI operates and maintains. The contrast with the existing options is clear. The Agents SDK runs the harness inside your own application, and the Responses API leaves orchestration to you. The Agents API hands that layer over so you can concentrate on the tools, knowledge, and workflows specific to your application.

Context compaction and tool search

OpenAI describes three harness improvements.

The first is automatic context compaction. As a session approaches its context limit, earlier material is compacted automatically while the information the agent still needs is preserved. Workflows that span several context windows can be built without writing compaction logic.

The second is tool search. Instead of loading every tool definition up front, only the relevant definitions are loaded when they are needed. The stated point is that this cuts token usage and cost while leaving the model cache intact.

The third is programmatic tool calling. Agents can run calls in parallel, chain related operations, and filter or combine results in code before bringing anything back into context, so large volumes of data can be processed while only the relevant portion returns. The API supports MCP, custom functions, and built-in tools such as web search.

Subagents run in parallel with separate contexts

Complex jobs can be broken into independent pieces and delegated to subagents that run concurrently. Each subagent keeps its own context, which helps it stay on its assignment, while the main agent coordinates the work and combines the results. The number of concurrent subagents is configurable, and tasks that divide cleanly, such as research, analysis, and coding, finish sooner.

One point deserves care here. Parallel execution shortens elapsed time, not total inference. Running three branches at once brings the wall clock close to one third, but the token spend remains roughly three branches worth.

You choose where the agent runs

Developers pick the place where the agent executes code and works with files. There are three options: an OpenAI-hosted sandbox, your own infrastructure, or an ecosystem provider. The named partners are Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop, and Vercel, covering deployments inside a VPC, specific file and secret storage mechanisms, different CPU, GPU, and memory configurations, and varying cold-start and cost profiles.

The hosted sandbox runs on the same infrastructure as Codex and ChatGPT and can be configured with your files, packages, skills, and plugins. Separating the harness from the execution environment is the substance of the design.

Numbers from early customers

Among the deployments OpenAI cited, Ciridae raised its evaluation score from 0.71 to 0.85 and cut latency by a factor of four using out-of-the-box subagent support. SafetyKit moved case review workloads over and reported a 60 percent reduction in cost per case. Hypha saw failed responses drop by 86 percent after separating the harness from the sandbox.

The foundation is the open-source Codex harness, so the core logic that coordinates model calls, tools, and context can be read from the public codebase. OpenAI runs and maintains it while developers can still inspect how it works.

Constraints worth checking first

The current limitations are stated plainly. Data residency is US-only, and Zero Data Retention is not supported, which holds even when the sandbox runs on your own infrastructure. Anyone evaluating this for regulated workloads should settle those two points first.

Cost deserves a note as well. There is no additional charge for the Agents API itself; you pay for tokens, tools, and container time when using the hosted sandbox. Still, this is built for work that runs for hours or days, so token consumption accumulates. The engineering time saved on plumbing has to be weighed against the running cost.

This is a public beta, and OpenAI says it will iterate quickly on developer feedback on the way to general availability. Approaching it as a moving target is the sensible stance.

Summary

On September 10, OpenAI released the harness and infrastructure behind Codex and ChatGPT as the Agents API in public beta. Automatic context compaction, on-demand tool search, subagents running in parallel, and a choice of execution environments spanning nine partners all sit behind a single API call. There is no extra fee, but US-only data residency and the lack of Zero Data Retention remain, so the nature of the data you handle is the practical place to start.

※The thumbnail image is AI-generated and for illustrative purposes.