OpenAI has published a technical post explaining how it raised the efficiency of GPT-5.6[1]. The notable part is who did the optimization work: GPT-5.6 Sol, the flagship model itself. Working through Codex, it rewrote production GPU kernels and cut end-to-end serving costs by 20 percent. The post covers gains stacked across both the inference stack and the agentic harness, not just the model.

Efficiency Is Not Decided by the Model Alone

GPT-5.6 comes in three tiers, Sol, Terra and Luna, designed to balance capability against cost. According to OpenAI, Sol with max reasoning outperforms Claude Fable 5 on the Artificial Analysis Coding Agent Index at less than half the cost. Terra matches GPT-5.5 on intelligence benchmarks at half the price, and Luna is the fastest and most affordable option, priced 80 percent below Sol[1].

Training is not what makes those numbers possible on its own. Over the past four years the company has scaled to 1 billion active users and more than 2 million businesses, which makes serving more tokens from the same hardware a top design priority. A model can be efficient in isolation and still be expensive to run if requests are distributed poorly, hardware sits idle, or data movement stalls computation[1].

So OpenAI worked across layers: routing (where requests are sent), scheduling (when they are sent), the kernels that run on GPUs, caching that reuses earlier work, and model implementation down to the ordering of GPU code. Globally, requests are routed by geography, available capacity and accelerator type; within a cluster, work is distributed by load, context length and cache availability[1].

The AI Rewrote the Kernels Itself

This is where GPT-5.6 Sol takes center stage. Sol analyzed production traffic, identified previously overlooked sources of imbalance, tested new routing strategies and continuously tuned the heuristics. OpenAI says these load balancing improvements alone dramatically reduced the cost of serving its models[1].

It went further with the forward pass, the computation that turns inputs into next-token predictions. Running in Codex, Sol found work that could be precomputed, avoided or parallelized, and autonomously rewrote production kernels, the core code that executes the mathematical operations making up the model. That was possible because GPT-5.6 had been trained to write and improve code in Triton and Gluon, two open-source GPU programming languages maintained by OpenAI. Together, these kernel improvements reduced end-to-end serving costs by 20 percent[1].

Code written by an AI cannot simply be pushed to production, so verification received investment at the same time. OpenAI released FpSan, a floating-point sanitizer, as an open-source tool used to validate the correctness of the kernels Sol produced[1]. Technical press coverage has framed the whole effort as work delegated to Sol inside a human-led process, stressing that this is not fully autonomous operation[2].

Tuning Speculative Decoding and Cache Configuration

Speculative decoding, a lever for gaining both speed and efficiency, also got attention. It runs a smaller draft model alongside the primary model, proposing several tokens that the primary model verifies in parallel. When proposals are accepted, a single pass of the primary model can produce multiple output tokens, cutting expensive sequential computation[1].

Sol improved that draft model as well. It designed and ran hundreds of experiments varying size, structure and features, then launched and monitored the speculator training process. When hardware failures or training instability appeared, it intervened autonomously. Token-generation efficiency rose by more than 15 percent as a result[1].

Caching followed the same pattern. When processing uncached input tokens, the model builds the key-value (KV) cache in one compute-intensive pass, then repeatedly reads from and extends it during generation. The optimal setup for batching, sharding and KV management depends on the workload, but the configuration space was too large to tune systematically, leaving engineers to rely on broad heuristics. With Sol in Codex analyzing production workloads and generating and evaluating candidate configurations, per-scenario tuning became practical[1].

The Harness Fights Context Bloat

The other pillar is the agentic harness shared by Codex and ChatGPT Work. It is implemented as a Rust orchestration layer connecting the models, the tools and the user's environment[1].

Agent work stacks up model requests within a single turn: inspecting source code, searching deployment history, reading incident reports, editing a file, running tests. If a task needs 30 requests, one wasted second per request becomes 30 seconds overall. That is why cutting repeated work across the system matters as much as making the model faster[1].

One concrete measure is deferred discovery, which surfaces integrations, custom MCP tools, skills and plugins only when they are needed. Tool output is also capped at 10,000 tokens by default unless the model requests a different limit, preventing any single tool from unexpectedly consuming the context window[1].

Prompt caching gets equally careful treatment. All model-visible history is treated as append-only, so new messages and tool results are added at the end rather than inserted into earlier context. Tools are presented in a deterministic order, and runtime settings such as approval policies are applied during execution instead of being embedded in tool definitions. OpenAI credits this design for the high prompt-cache hit rates of Codex and ChatGPT Work[1].

Summary

OpenAI describes the efficiency of GPT-5.6 as the product of gains stacked across three fronts: the model, inference and the agentic harness. The standout result is that GPT-5.6 Sol rewrote production kernels through Codex to cut serving costs by 20 percent, and improved speculative decoding to raise token-generation efficiency by more than 15 percent. With compute still constrained, it is a concrete sign that the work of optimization is shifting toward the AI side.

Source[1]: https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency

Source[2]: https://thenewstack.io/gpt-5-6-serving-efficiency/