On June 10, 2026 (local time), Google DeepMind released DiffusionGemma, an experimental open model built for exceptionally fast text generation[1]. Instead of producing text one word at a time, it adopts a diffusion approach that refines a whole block of text at once, much like image generation models. NVIDIA announced it has optimized the model for its GPUs, including GeForce RTX and DGX Spark[1]. For single-user workloads, it delivers up to 4x faster generation than a comparable conventional model[1].
A New Approach: Generating Text in Blocks
Almost every large language model (LLM) in wide use today is autoregressive, generating one token at a time with each new word depending on the one before it. That sequential process is why interactive AI feels like it is typing out its answer[1].
DiffusionGemma takes an entirely different path. Like diffusion models for images, it starts from noise and progressively denoises an entire block of text, generating up to 256 tokens in parallel per step[1]. It is a model that thinks in blocks rather than sequentially.
The model is built on Gemma 4, a 26-billion-parameter mixture-of-experts model that activates just 3.8 billion parameters per step[1]. It pairs a diffusion head with the Gemma 4 architecture[1]. DiffusionGemma is released as open weights under the permissive Apache 2.0 license and runs entirely on RTX and DGX Spark — no cloud, no per-token cost[1].
Why It Gets Faster on GPUs
Generating one token at a time is fundamentally a memory-bandwidth-bound process. At batch size 1 (a single user), a traditional LLM spends most of its time waiting on memory rather than computing[1]. The GPU's compute power sits largely idle.
Diffusion flips this equation. Pulling a full 256-token block through the transformer in parallel is a compute-bound workload — exactly what GPUs excel at[1]. NVIDIA Tensor Cores accelerate the dense parallel math, and the CUDA software stack lets the model run efficiently from day one without bespoke tuning[1].
Measured numbers have been published as well. DiffusionGemma delivers 1,000 tokens per second at batch size 1 on a single H100 GPU, 150 tokens per second on the deskside DGX Spark, and up to 2,000 tokens per second on DGX Station — roughly 4x faster than an equivalent autoregressive model[1]. That translates into noticeably shorter waits for chat responses and agentic reasoning loops.
Broad Support from Local PCs to Workstations
The model runs across NVIDIA's full lineup[1].
- DGX Spark: a deskside personal AI supercomputer powered by the GB10 Grace Blackwell Superchip with 128GB of unified memory. With the preinstalled AI software stack, it is ready for prototyping and fully local agent workflows[1]
- RTX PRO 6000 workstations: give developers and researchers the headroom to run low-latency generation and agentic loops as part of a professional workflow[1]
- DGX Station: offers best-in-class inference at up to 2,000 tokens per second with 748GB of coherent memory[1]
- GeForce RTX: llama.cpp support is coming soon[1]
Getting started is straightforward. Hugging Face Transformers runs DiffusionGemma on a GeForce RTX 5090 or DGX Spark out of the box, and vLLM provides day-zero serving support for higher-throughput inference[1]. For adapting the model to a specific task, fine-tuning is available through Unsloth and the NVIDIA NeMo framework, with ready-made playbooks for DGX Spark[1]. You can try it on Hugging Face or test it for free via NVIDIA-hosted APIs at build.nvidia.com[1].
Summary
DiffusionGemma is an experimental open model that uses parallel diffusion-based generation to boost single-user text generation — long a weak point of local AI — by up to 4x[1]. Released under Apache 2.0 and running entirely on RTX and DGX Spark, it is an appealing option for developers who want fast local agents without cloud costs. If llama.cpp support brings it to the broader GeForce RTX user base, this could become a new standard approach for local LLMs.
Source: [1] https://blogs.nvidia.com/blog/rtx-ai-garage-local-gemma-diffusion/
