Ollama, the local AI runtime, announced a major update to its MLX engine for Apple Silicon on June 11, 2026. The update is built around three pillars: support for NVIDIA's model-optimized NVFP4 format, output speeds up to 20 percent faster thanks to Metal kernel optimizations, and a new snapshot system that streamlines agent workloads. By leaning more heavily on Apple's unified memory and the Metal-backed MLX framework, Ollama says local AI on the Mac has reached its highest performance yet[1].
NVFP4 Support Roughly Halves the Quality Loss of 4-bit Quantization
One of the highlights of this update is support for NVIDIA's model-optimized NVFP4 format. NVFP4 is a 4-bit quantization format designed to track the local dynamic range of model weights more closely, which keeps quantization-induced quality loss to a minimum[1].
When Ollama compared perplexity (a metric for measuring a language model's predictive accuracy) on the Gemma 4 12B model across q4_K_M (a common 4-bit quantization format), NVFP4, and unquantized bf16 weights, model-optimized NVFP4 roughly halved the quality loss while maintaining performance[1].
Another point worth noting is that models optimized for datacenter deployment can now be imported and run directly on Ollama's MLX engine. Since NVFP4 has primarily been used for datacenter inference, the same model can now be carried between the datacenter and the desktop without maintaining separate builds[1].
Output Speeds Up to 20 Percent Faster via Fused Metal Kernels
Output performance has been strengthened as well. Using MLX's just-in-time compiler features, several operations are now fused into single Metal kernels, and Ollama's GPU-backed sampling has been reworked to run more efficiently, making the MLX engine up to 20 percent faster than before[1]. Ollama only released its MLX support for Apple Silicon as a preview in March 2026[2], so it is steadily stacking up performance gains within just a few months.
A New Snapshot System Built for the Agent Era
The most interesting part of this update is the snapshot system designed with agent use cases in mind. In workloads like coding agents, every tool call resends the entire conversation, including the system prompt, tool definitions, and every file read so far, so the same context ends up being processed dozens of times. Prefix caching (a mechanism that resumes from where the previous request left off) has traditionally avoided this duplicated work, but it only functions when each request picks up exactly where the last one ended[1].
Real agent sessions do not stay that simple for long. The new snapshot system therefore saves model state at the points conversations are likely to return to, using the same approach that serves agent workloads in Ollama's cloud. It pays off in scenarios like the following[1].
When an agent hands off to a subagent or multiple sessions run in parallel, each one resumes from its own saved state, and anything they share, such as system prompts and tool definitions that can span tens of thousands of tokens, is only processed once. With reasoning models, thinking tokens are dropped from the conversation history, which would normally force the entire conversation to be reprocessed every turn, but a snapshot taken right before the response starts gives the next turn somewhere to resume from. When a conversation branches through a regenerated response or a different follow-up question, only the new direction needs to be processed because snapshots exist where conversations split[1].
Sliding-window attention and recurrent layers, which are common in recent models, carry state that cannot be rewound once the model moves past a point in the conversation, so this kind of state saving is harder than it sounds. Ollama says it keeps snapshots selective and incremental, leaving more memory for the model itself[1].
Getting Started
To use the MLX engine, simply download the latest version of Ollama and run a model[1].
ollama run gemma4:12b-mlx
For use in a coding agent, use the ollama launch command[1].
ollama launch pi --model gemma4:12b-mlx
Summary
With quality gains from NVFP4 support, speeds up to 20 percent faster, and a snapshot system that underpins agent workloads, this update is packed with improvements that raise the practical value of local AI. The portability of bringing datacenter-optimized models straight to the Mac should also widen the options for working with open models. Personally, I am eager to see how noticeable the improved responsiveness feels when running agents locally.
Source [1]: https://ollama.com/blog/mlx-performance
Source [2]: https://ollama.com/blog/mlx
