Inception released Mercury 2.5 on September 8, a large language model built on diffusion rather than the usual left-to-right generation. It runs at 1,107 tokens per second and claims a 40 percent gain in intelligence over Mercury 2. A series that sold itself on speed is now lining up against cost-optimized frontier models on quality as well.
Placing tokens in parallel instead of one at a time
Most large language models write text one token after another, front to back. The Mercury series drops that assumption. It borrows diffusion, the technique behind modern image generation, and applies it to language: the model starts from a rough draft and rewrites tokens in parallel until the output settles. Because tokens are not queued up single file, the GPU spends less time idle and responses come back faster on the same hardware.
Inception describes Mercury 2.5 as the most capable diffusion LLM on the market and, to its knowledge, the largest diffusion language model ever trained. The speed figures are measured on widely available NVIDIA GPUs rather than on some specialized accelerator, which matters for anyone planning a deployment.
Same serving profile, more capability
The numbers tell the story. Generation runs at 1,107 tokens per second, and the context window holds 260,000 tokens. Intelligence is up 40 percent from Mercury 2, putting the model on par with cost-optimized frontier models such as GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5.
Pricing is 0.20 USD per million input tokens (about 31 yen) and 0.75 USD per million output tokens (about 115 yen). At launch, an 80 percent discount brings that down to 0.04 USD (about 6 yen) and 0.15 USD (about 23 yen). New capabilities include tunable reasoning, parallel tool calls, and schema-aligned JSON output.
※1 USD = 153 JPY
Inception says the gains came out of production use rather than benchmark chasing. Since Mercury 2 shipped, usage has grown by more than an order of magnitude across latency-sensitive workloads in search, voice, and coding. Customer feedback and real failure cases were folded into the evaluation set, and training was focused around them. NVIDIA also commented that the release shows how quickly a new architecture can mature into a production-ready system.
Where speed compounds: calling the model many times
Speed pays off less in a single visible answer than in the dozens of model calls hidden behind it. A single search request can trigger a chain of work: plan the search, rewrite queries, rerank results, structure the facts, summarize sources, and verify the answer. If those steps are slow, the whole thing falls outside the window a user is willing to wait.
In voice, latency is simply the silence a caller hears. OpenCall, which builds AI phone agents, brought median model response latency to roughly 170 milliseconds after switching to Mercury. The company reports that its slowest responses dropped from several minutes to about one second, and the median fell from 0.4 seconds to under 0.2 seconds.
Coding assistants work the same way. Augment Code uses Mercury for context compaction, model routing, and MCP tool search. Moving compaction over cut the time from roughly 150 seconds to 27 seconds, an 82 percent reduction, and lowered cost by 90 percent. Tool-search summaries return in under a second.
Mercury Voice and Mercury Router in preview
Two related models arrived alongside Mercury 2.5. Mercury Voice is tuned for voice agents and delivers time to first token under 170 milliseconds. Mercury Router reads incoming prompts with a diffusion model and dispatches each one to whichever model offers the best mix of quality, speed, and cost, drawing on both open and closed options.
The Mercury family is available through Inception's own API as well as Baseten and OpenRouter. The API is OpenAI compatible, so it drops into existing code with little change. Enterprise deployments add dedicated capacity, autoscaling, compliance controls, and configurable data retention. Inception has already begun training its next model, which it says will be its largest yet.
Summary
Mercury 2.5 raises intelligence by 40 percent while holding the speed and price profile of its predecessor. At 1,107 tokens per second, a 260,000-token context, and 0.20 USD per million input tokens, it now sits alongside the cost-optimized tier of frontier models. In search, voice, and the supporting calls inside coding assistants, where a single interaction fans out into dozens of model invocations, that speed and price accumulate into a visibly different experience. The release suggests diffusion is no longer a trade of quality for throughput.
