On May 5, 2026, Google released Multi-Token Prediction (MTP) drafters for the Gemma 4 family of open-weight models. According to the company, the new drafters use a specialized speculative-decoding architecture to deliver up to a 3x speedup without any degradation in output quality or reasoning logic. They ship under the Apache 2.0 license and are available immediately on Hugging Face and Kaggle [1]. The release also closes a long-standing community gripe that the original Gemma 4 weights on Hugging Face had their MTP heads stripped out [2].

What MTP Drafters Are Solving

Standard large-language-model inference is memory-bandwidth bound — Google notes that the processor spends most of its time moving billions of parameters from VRAM to the compute units just to produce a single token, and the latency hit is most visible on consumer-grade hardware [1]. Speculative decoding mitigates that inefficiency by pairing a heavy target model (for example Gemma 4 31B Dense) with a lightweight drafter (the MTP model). In the time it takes the target model to process one token, the drafter "predicts" several future tokens, and the target model then verifies them in parallel [1].

The technique itself is not new. Google's Yaniv Leviathan, Matan Kalman, and Yossi Matias introduced it in their 2022 paper "Fast Inference from Transformers via Speculative Decoding," published at ICML 2023, where the same approach delivered a 2x to 3x acceleration on T5-XXL with identical outputs [3]. The new MTP drafters can be read as that line of work tuned specifically for Gemma 4.

The Source of the 3x Number — and the Caveats

Google says it benchmarked the drafters across LiteRT-LM, MLX, Hugging Face Transformers, and vLLM, and the announcement page features a side-by-side video of Gemma 4 26B running on an NVIDIA RTX PRO 6000, where standard inference is paired against MTP-drafted inference at "the same output quality, half the wait time" [1].

That 3x figure, however, is an upper bound. For the 26B Mixture-of-Experts model on Apple Silicon, batch size 1 hits MoE-specific routing challenges, and Google itself notes that batch sizes of 4 to 8 are needed to unlock up to roughly a 2.2x local speedup. Similar gains were observed on NVIDIA A100 when batch sizes were increased [1].

External validation widens the range. A community-trained EAGLE3 drafter head for Gemma 4 31B has reported a 1.72x speedup with no change to outputs, with conversation-heavy workloads such as MT-Bench getting more lift than code-heavy SWEBench tasks where token sequences are harder to predict [4].

Distribution and the Architectural Tweaks

The MTP drafters carry the same Apache 2.0 license as Gemma 4. Weights are distributed via Hugging Face and Kaggle, and the supported runtimes include Hugging Face Transformers, MLX, vLLM, SGLang, and Ollama. On mobile, the drafters can be tried directly in the Google AI Edge Gallery for Android and iOS [1].

On the architecture side, the drafter models reuse the target model's activations and share its KV cache, which avoids re-computing context the larger model has already worked out. For the edge-class E2B and E4B models, where the final logit calculation is the dominant cost, Google added a clustering-based optimization in the embedder to keep generation fast [1].

It is worth recalling that when Gemma 4 first shipped in April, the MTP heads used during training were stripped out of the Hugging Face release, and the heads were initially available only through Google's own LiteRT-exported artifacts. The community filled the gap with EAGLE3-based drafters of its own. Today's release effectively pulls that capability back into the official open-weight package [2][4].

Summary

The new MTP drafters are open-source, Apache 2.0 helper models for speculative decoding on Gemma 4, advertised as up to 3x faster inference. The realized speedup varies sharply with hardware and batch size — about 2.2x on Apple Silicon MoE and around 1.72x for the community EAGLE3 drafter, for instance — so the practical move is to try them on Hugging Face, Kaggle, or the AI Edge Gallery against your own workload before betting production traffic on the full 3x figure.

出典:[1] Google, "Accelerating Gemma 4: faster inference with multi-token prediction drafters" (https://blog.google/innovation-and-ai/technology/developers-tools/multi-token-prediction-gemma-4/)

出典:[2] FlowHunt, "Gemma 4 Was Released Without MTP Data — Here's Why That Matters" (https://www.flowhunt.io/blog/gemma-4-released-without-mtp-multi-token-prediction/)

出典:[3] Yaniv Leviathan, Matan Kalman, Yossi Matias, "Fast Inference from Transformers via Speculative Decoding," arXiv:2211.17192 (https://arxiv.org/abs/2211.17192)

出典:[4] Hugging Face Blog, "Google Released Gemma-4 Four Days Ago. We Already Made It 1.72× Faster." (https://huggingface.co/blog/lujangusface/tw-eagle3-gemma4)