OpenAI published a technical guide for teams building AI agents on August 13, titled The builder's guide to GPT-5.6[1]. It collects lessons from startups running the models in production, and argues that simply moving from a design that hands every step to a top-tier model, to one that mixes model selection with new API primitives, cuts costs sharply.

Handing Everything to the Flagship Is No Longer the Default

For long-horizon work, the standard approach used to be running the flagship model at the highest available reasoning effort. Frontier models were far better than cost-optimized ones at handling long contexts and tool calling[1].

OpenAI argues that the GPT-5.6 family breaks that assumption. Given more test-time compute, the smaller Luna and Terra models can often match GPT-5.4 and GPT-5.5 while costing significantly less[1].

One example the company cites is Agents' Last Exam, where GPT-5.6 Sol at low reasoning effort outscored GPT-5.5 at high reasoning effort with the harness held constant[1]. Startups running production tests report meaningful cost improvements across a range of workflows simply by dialing reasoning effort below the previous defaults.

Roughly One Twenty-Fifth the Cost at the Same Accuracy: the BrowseComp Numbers

The clearest illustration is BrowseComp, a benchmark that measures how well a model can search out obscure facts[1].

Three months ago, GPT-5.5 at Extra High reasoning scored 84.36 percent at a total cost of 33.27 USD (about 5,300 yen). GPT-5.6 Luna at the same Extra High setting delivers essentially the same result, scoring 84.04 percent for 1.33 USD (about 210 yen). Accuracy holds, and the bill drops to roughly one twenty-fifth. OpenAI adds that prices have come down further since then[1].

※1 USD = 159 JPY (based on the New York close on August 14, 2026)

The smaller models are positioned for high-volume workloads, latency-sensitive interactions, and steps that repeat inside an agentic workflow[1]. The concrete example given is a legal-tech company that parses handwritten memos before running agentic analysis. Moving just the extraction step off the frontier model and onto Terra or Luna brings the cost down.

Three New Primitives in the Responses API

GPT-5.6 was trained end-to-end around three architectural interventions, each exposed as a feature of the Responses API[1].

The first is not throwing away work already done. Reasoning can be persisted across model turns, and long-running conversations are compressed with native compaction. The model keeps its bearings over longer task horizons instead of getting confused or rebuilding prior context from scratch.

The second is decomposing work that can run in parallel. Native multi-agent orchestration lets several agents work concurrently so complex tasks finish faster.

The third is pushing deterministic work into code. With programmatic tool calling, the model writes JavaScript to orchestrate tools and processes the filtering and aggregation of their output outside the context window. Model tokens are reserved for judgment, which reduces cost, latency, and the context rot that sets in as conversations grow[1].

Retrieving 100 filings, filtering them by date, and identifying the relevant transactions is a good example: there is no reason for the model to reason over every intermediate result. The cleaner the split, the larger the gain.

Roughly Triple the Score Without Changing the Model: the ARC-AGI-3 Case

Combined, the effect is substantial. On ARC-AGI-3, GPT-5.6 Sol scored 13.3 percent with the standard harness[1]. With retained reasoning and compaction enabled, the score jumped to 38.3 percent, while using roughly six times fewer output tokens.

Nothing about the model changed. Only the way it was called did, and the score roughly tripled. The flip side is that running with defaults may leave you misjudging what a model can actually do. It is worth checking what harness settings sit behind any benchmark figure.

The guide also covers multi-agent behavior. A primary agent orchestrates subagents and delegates tasks, the subagents work in parallel, and the primary agent synthesizes the final output[1]. OpenAI notes that the ultra capability setting in ChatGPT works this way internally. GPT-5.6 already has a good sense of how many subagents to spawn and when, but the behavior is highly steerable, so instructing the model to spawn them only where the extra tokens will pay off is effective.

Prompt Cache TTL Extended to at Least 30 Minutes

A quieter change that matters in production is prompt caching. Across the entire GPT-5.6 family, the cache TTL has been extended to a minimum of 30 minutes[1]. Cache breakpoints can also now be set deterministically within the context window, which OpenAI says has significantly improved cache hit rates for some startups.

Continuing to use an appropriate prompt_cache_key also raises the odds that a request lands on the same inference engine that previously served the same prefix, which reduces latency[1].

The guide was written by Samarth Madduru, Prashant Mital, Dave Leo, and Julien Reiman, based on their work with startups building on GPT-5.6 from early testing through production[1].

Summary

OpenAI's builder's guide to GPT-5.6 comes down to one point: the economics of building agents have changed. On BrowseComp, essentially the same accuracy now costs 1.33 USD instead of 33.27 USD, roughly one twenty-fifth. On ARC-AGI-3, enabling retained reasoning and compaction alone lifted the score from 13.3 percent to 38.3 percent while cutting output tokens by about six times. With the flagship at maximum reasoning no longer the automatic answer, choosing the right model and rethinking how it is called now translates directly into lower costs.

Source: https://openai.com/index/builders-guide-to-gpt-5-6