Ollama, the tool that lets you run large language models locally, has significantly increased the generation speed of Google's open model Gemma 4 in its latest release, version 0.31[1]. On Apple Silicon Macs, token generation is reported to be about 90 percent faster on average across a benchmark that simulates coding work[1]. The speedup is enabled by default, and the model's actual output does not change. The technique behind it is called multi-token prediction (MTP).

Gemma 4 Gets Much Faster in Ollama 0.31

The target of this improvement is Gemma 4, the lightweight open model published by Google. In Ollama 0.31 and later, inference for this model has been optimized for Apple Silicon, and the company measured an average speed increase of roughly 90 percent in token generation[1].

What matters is that the output itself does not change even as things get faster. This is an improvement aimed purely at "returning the same answer more quickly," without sacrificing accuracy or response quality. No configuration is required either; simply installing a supported version delivers the speed benefit automatically[1].

Multi-Token Prediction (MTP), the Engine Behind the Speedup

At the heart of the speed gain is multi-token prediction (MTP). Gemma 4 ships with a small, fast "draft model" that runs alongside the main model and proposes several of the next tokens at once[1]. The main model then verifies that proposal in a single pass and keeps only the tokens it judges to be correct. Because the draft model is far smaller than the main model, its predictions are cheap, and when they are right, several tokens can be committed for the cost of processing one[1].

This approach is said to work especially well for generating code. Program code contains many closing brackets, repeated identifiers, and boilerplate, which makes the next content easy to predict[1]. Coding agents, which read files and run tools as they work, call the model continuously. When each individual generation is faster, the responsiveness of these agents improves noticeably.

The technique of predicting and verifying multiple tokens for a speedup is known as speculative decoding, and Google itself has also introduced a similar acceleration for Gemma 4 using draft models[2].

Three Technical Refinements

Ollama explains that it combined three elements to draw out this speed[1].

The first is automatic tuning of how many tokens to draft. The optimal number of tokens to predict at once changes depending on the model, the quantization method, the hardware, and how predictable the text is during generation. Drafting too few leaves performance on the table, while drafting too many spends time verifying rejected proposals and can end up slower. Ollama tracks the acceptance rate of predictions and the time each verification takes as it runs, then selects the length that yields the most tokens per second at any moment. When proposals stop being accepted, it automatically returns to plain one-at-a-time generation[1].

The second is a refinement in how the process is carried out. The whole sequence, from prediction by the draft model to batch verification by the main model, runs together on the GPU. Rejected tokens have already been written into the cache mid-process, but recovery only requires rewinding to a "rollback point" recorded in advance, keeping the cost of undoing them small[1].

The third is greater efficiency in the GPU computation itself. The heaviest workload is not drafting but verification. Verification handles a batch of roughly 2 to 8 tokens, an awkward size, whereas conventional computation is optimized for either a single token or a large batch and struggles with the middle ground. Ollama contributed a routine for this case to Apple's machine learning framework, MLX. It reads weight data once and reuses it across the batch; on an M5 Max with nvfp4, this makes Gemma 4's largest matrix multiplications 2 to 2.5 times faster[1].

How to Get Started

You need Ollama 0.31 or later for macOS. After installing, you can launch a coding agent powered by Gemma 4 with the following command.

ollama launch claude --model gemma4:12b-mlx

If you obtained Gemma 4 earlier, re-pull the model to get the MTP-enabled version.

ollama pull gemma4:12b-mlx

This launch method also works with other coding assistants such as Codex, Droid, OpenCode, and Copilot[1]. According to Ollama, Gemma 4 is the first model to receive this speedup, with more to follow[1].

Summary

Ollama 0.31 combines multi-token prediction with computation optimized for MLX to raise Gemma 4's generation speed on Apple Silicon by about 90 percent on average. It gets faster without changing the output, requires no configuration, and is especially effective for continuous calls such as those made by coding agents. The usability of running models locally has taken another step toward practical, everyday use.

出典:https://ollama.com/blog/faster-gemma-4-mlx-mtp

出典:https://blog.google/innovation-and-ai/technology/developers-tools/multi-token-prediction-gemma-4/