OpenAI revised its GPT-5.6 API pricing on July 30. The company also published a piece explaining how it cycles efficiency gains back into investment, and mentioned there that Luna had been cut by 80 percent and Terra by 20 percent[1]. Luna input now costs 0.20 USD (about 32 yen) per million tokens, while the flagship Sol keeps its price and gains a new speed option instead[2][3].

Luna at a Fifth, Terra 20 Percent Cheaper

The new prices took effect on July 30. Here is how they compare with the previous rates[2][3].

Model Input (per million tokens) Output (per million tokens)
GPT-5.6 Luna (before) 1 USD (about 160 yen) 6 USD (about 950 yen)
GPT-5.6 Luna (after) 0.20 USD (about 32 yen) 1.20 USD (about 190 yen)
GPT-5.6 Terra (before) 2.50 USD (about 400 yen) 15 USD (about 2,400 yen)
GPT-5.6 Terra (after) 2 USD (about 320 yen) 12 USD (about 1,900 yen)

※1 USD = 159 JPY (as of July 31, 2026)

The change is not limited to the API. Both models now consume fewer usage credits in ChatGPT Work and Codex, so subscribing companies can run more work without paying more[3].

The GPT-5.6 family launched on July 9, which puts this revision roughly three weeks later. The lineup runs Sol, Terra, and Luna from the top down, with Luna positioned as the fastest and cheapest tier.

Sol Holds Its Price and Gains Fast Mode

Sol, the flagship, sees no price change. Instead, the API replaces the earlier Priority Processing with a new Fast mode[3]. It delivers up to 2.5 times the speed of standard processing at twice the price, with no change in intelligence[1]. Existing requests tagged as priority continue to work[3].

OpenAI argues that the framing itself needs to change. Rather than asking which model belongs to which task, the questions are how much intelligence the outcome demands, how quickly it is needed, and what it should cost. That balance may shift several times inside a single workflow[1].

Serving Efficiency Paid for the Cut

The company attributes the lower prices to efficiency work across its training and inference stack[3]. The clearest example is GPT-5.6 Sol helping optimize the production GPU kernels, which brought end-to-end serving costs down by 20 percent. Improvements to speculative decoding raised token-generation efficiency by more than 15 percent[1].

Efficiency is not determined by the model alone. In a recent benchmark analysis, better retained reasoning and context management raised Sol's score on the public ARC-AGI-3 task set from 13.3 percent to 38.3 percent while using six times fewer output tokens. The model did not change; only the system around it did[1].

The Real Metric Is the Cost of a Successful Outcome

OpenAI also notes that customers do not buy tokens for their own sake. They want the support issue resolved, the software shipped, or the contract reviewed. The right measure is the cost of a successful outcome, including the time, retries, oversight, and errors it takes to get there. A stronger model that finishes the job once can end up cheaper than a low-cost model that needs repeated attempts[1].

The scale behind that argument is substantial. The models now reach more than one billion active users and more than two million businesses. Six months after signing up, people send roughly 50 percent more messages per day and use ChatGPT for about twice as many kinds of work. Internally, agentic work through Codex accounts for 99.8 percent of weekly output tokens[1].

Enterprises Will Spend Differently, Not Less

Analysts do not expect the cut to shrink enterprise AI budgets. Pareekh Jain, principal analyst at Pareekh Consulting, says the biggest impact for CIOs is wider deployment rather than cost reduction. Lower prices make it easier to move pilots into production and make agentic workflows that require multiple model calls economically viable. Most enterprises will reinvest the savings into higher usage, he says, pointing to the Jevons paradox, where efficiency gains tend to raise total consumption[3].

Chandrika Dutt, research director at Avasant, similarly expects teams to spend the new headroom on sophisticated agentic workflows that were previously hard to justify. Jain adds that new chips, better software, and heavy competition make it likely that inference costs will keep falling over the next 24 months, so CIOs should design for AI FinOps and flexible architectures that let them swap models as price-performance shifts[3].

Summary

This revision effectively puts Luna at a fifth of its former rate and Terra 20 percent lower, while the flagship Sol keeps its price and adds a speed option through Fast mode. Behind the cut is a reduction in serving costs that the model itself helped deliver, which OpenAI frames as a cycle linking efficiency to investment. What to do with the cheaper tokens is now a design decision on the customer's side.

Source[1] https://openai.com/index/building-abundant-intelligence/

Source[2] https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/

Source[3] https://www.infoworld.com/article/4203865/openai-drops-gpt-5-6-luna-and-terra-api-prices-by-up-to-80.html