On August 13, OpenAI gave an early look at Ultrafast, a new service tier that runs its most capable model, GPT-5.6 Sol, up to 14 times faster than Standard processing[1]. It launches first in the OpenAI API, with inference powered by Cerebras. Output reaches up to 750 tokens per second, and the model retains the same intelligence as the Standard tier[1][2].
The speed-versus-intelligence tradeoff no longer holds
Until now, getting real-time responses usually meant switching to a smaller or more specialized model[1]. Frontier models were smart, but you waited for them.
Ultrafast points in a different direction. OpenAI frames it as getting more useful work done per second[1]. Once speed no longer costs you intelligence, AI can move into the most time-sensitive parts of a business, and new kinds of work become possible[1].
Cerebras makes the same point. Andrew Feldman, CEO and co-founder, said GPT-5.6 Sol on Ultrafast is proof that speed and intelligence are no longer mutually exclusive[2]. Sachin Katti, VP Compute Strategy and GPT-Infra at OpenAI, said the company is exploring what becomes possible when customers can get the intelligence of its most capable models at significantly lower latency[2].
The Wafer-Scale Engine behind 750 tokens per second
The speed comes from the Cerebras Wafer-Scale Engine architecture[2]. An entire wafer serves as a single chip, carrying 44 GB of SRAM on-chip so model weights stay resident during inference[2].
GPU-based inference has to shuttle weights between on-chip memory and off-chip storage, and that memory-bandwidth bottleneck caps frontier-model inference speed[2]. Keeping the weights on-chip removes the constraint. The result is up to 750 output tokens per second with the full intelligence of GPT-5.6 Sol intact[1][2].
Cerebras also offers comparisons. Based on output speeds for Anthropic models reported by Artificial Analysis, Ultrafast is 5 times faster than Claude Opus 4.8 in Fast mode and 11 times faster than Claude Fable 5[2].
Three days of compute compressed into 11 hours
Benchmarks built around real work put numbers on the difference.
On Humanity's Last Exam, a 2,500-question set spanning graduate-level chemistry, economics, and literature, GPT-5.6 Sol Ultrafast finished the full set in just over 11 hours[2]. Claude Fable 5 needed more than three days of continuous compute for the same questions, meaning Ultrafast reached comparable accuracy nearly 7 times faster[2].
On GDP-Val, a benchmark of economically valuable knowledge work such as legal briefs, financial models, and engineering reports, Ultrafast delivered a 5.6x end-to-end speedup with no loss in quality[2]. The focus is not on benchmark scores themselves but on how long it takes to reach the same result.
Built for work that cannot wait
OpenAI listed several scenarios where Ultrafast looks promising[1].
In incident response and reliability, when a critical system fails, the model analyzes application logs, recent code changes, and engineer reports to identify a likely cause and help prepare a fix while the outage is still unfolding[1]. In financial research and security, it analyzes market signals, assesses transactions, and flags suspicious activity while conditions are still changing[1].
In customer support and voice, the aim is resolving complex issues in real time without interrupting the conversation, even when the answer requires multiple steps or systems[1]. In commerce, it answers product questions, checks inventory, personalizes recommendations, and clears checkout problems while the shopper is still deciding, before hesitation turns into an abandoned cart[1].
For live research and experimentation, work that previously took an overnight run becomes an interactive session[1]. Teams test an idea, examine the results, adjust their approach, and run another experiment without breaking their flow[1].
OpenAI is using it internally for incidents and research
Inside OpenAI, a group of developers has been testing GPT-5.6 Sol on Ultrafast to understand which workflows benefit from frontier intelligence that answers in real time[1].
Incident response is one example. When an alert fires, the system and the evidence are both still changing. Teams read logs, analyze traces, synthesize conversations, identify the next checks, and help prepare or validate a fix, all in a fraction of the time and with the intelligence of Sol[1]. The delay between observing a signal, testing a hypothesis, and choosing the next action shrinks, while engineers remain responsible for judgment and deployment[1].
For research, teams use it to search knowledge sources, query data, and gather, organize, and summarize information across connected tools[1]. The familiar pattern of launching a batch of experiments overnight and reviewing results in the morning is tightening into a loop that runs multiple times during the workday[1].
Limited preview, with pricing and general availability unannounced
GPT-5.6 Sol on Ultrafast mode is currently a limited preview available to a select group of customers[1][2]. OpenAI is working with companies across coding, commerce, financial research, and support to study these conditions in real production environments before widening deployment[1].
Access will expand as capacity grows, but the announcement does not include pricing or a general availability date[1]. Since this is an added speed tier, how the price gap against Standard is set will shape whether teams can actually build it into their workflows.
Summary
OpenAI has previewed Ultrafast, which runs GPT-5.6 Sol up to 14 times faster than Standard processing. Inference runs on the Cerebras Wafer-Scale Engine, which keeps model weights in 44 GB of on-chip SRAM to remove the memory-bandwidth bottleneck and reach 750 output tokens per second. On Humanity's Last Exam, a workload that took more than three days was compressed into just over 11 hours, and GDP-Val showed a 5.6x speedup with no loss in quality. Availability remains a limited preview for select customers, with pricing and a general release date still unannounced.
