NVIDIA has announced that Groq 3 LPX, an accelerator built specifically for interactive inference, is now in full production. It extends the Vera Rubin platform with a single goal: raising the rate at which tokens come out. The announcement landed alongside the Hot Chips conference, with AI cloud provider Nebius named as the first adopter.

Agent responsiveness comes down to generation speed

An AI agent takes hundreds or thousands of inference steps to finish a single request. It reads files, writes code, runs tests, calls external tools, checks the results and tries again. Every one of those steps is built on token generation, so slow generation turns the sheer number of steps into waiting time.

What NVIDIA targeted with Groq 3 LPX is the measure of how many tokens per second a single user gets back. The company calls this interactivity and describes it as the factor that determines how long an agent takes to complete each step of its work. Faster generation means an agent can fit more inspection and iteration into the same window of user patience.

Agentic workloads carry two different kinds of load at once: processing enormous amounts of context, and emitting tokens with very low latency. Groq 3 LPX is positioned to handle the second, extending the inference performance of Vera Rubin NVL72 from the generation side.

3,400 tokens per second on Gemma 4 31B

The headline figure comes from benchmarking by Artificial Analysis. Running Gemma 4 31B, an open source agentic model, under a 100,000-token context, the system recorded 3,400 output tokens per second. NVIDIA says this is the fastest performance ever recorded for that model.

Translated into practical terms, the company claims agentic tasks such as coding finish in minutes rather than hours, and that latency-sensitive workloads see 4x better responsiveness than the nearest alternative platform.

That said, NVIDIA does not identify which product the alternative platform refers to. The benchmark itself comes from a third party, but the substance behind the 4x comparison is hard to reproduce from what has been published.

Nebius goes first

Nebius is opening the commercial rollout. The company plans to bring Groq 3 LPX to Nebius Token Factory, its production inference platform. CTO Danila Shtan emphasized that generation is the phase that determines how responsive an AI system actually feels, and that developers get the speed through the same API they already use, with no migration to a new stack.

Groq, an AI cloud focused specifically on inference, is expected to be among the earliest adopters after Nebius. The Groq and LPU names in the product branding are used by NVIDIA under license from Groq, Inc.

Seven chips and five racks, designed together

Groq 3 LPX is not a standalone product but part of the Vera Rubin generation of rack systems. NVIDIA describes codesign spanning seven chips and five purpose-built racks, with Vera Rubin NVL72 and Groq 3 LPX splitting the work between them. The former serves as the versatile base for training and inference, while the latter takes on generation speed.

The racks carry BlueField-4 DPUs and operate in combination with Vera CPU racks, Vera BlueField-4 STX storage and Spectrum-6 SPX Ethernet. The stated design goal is to optimize multi-agent systems for the highest throughput per watt and the lowest-latency inference.

Founder and CEO Jensen Huang described inference as the growth engine of AI, and said Vera Rubin brings workload-optimized AI factory configurations designed for the era of agentic AI. The phrasing reflects how attention that used to sit on giant training clusters has been shifting toward the side that keeps models running.

Summary

Groq 3 LPX goes after per-user generation speed rather than raw aggregate compute. The 3,400 tokens per second figure on Gemma 4 31B is a direct answer to how tightly agent usability is currently bound to latency. Nebius comes first, followed by Groq, and how well those numbers hold up in real deployments is still ahead of us.