Muse Glimmer, the new open model from Meta Superintelligence Labs, is now available through Ollama[1]. It is a 30-billion-parameter multimodal model built specifically for agent workloads, and the first wave of support arrives via Ollama's MLX engine on Apple Silicon. In practical terms, the brains behind coding agents such as Claude Code and Codex can now sit on your own Mac instead of in the cloud.
Apple Silicon and the MLX engine come first
According to Ollama's August 10 announcement, support for Muse Glimmer is currently limited to initial coverage through the company's MLX engine on Apple Silicon[1]. Further Apple Silicon support and optimizations, along with NVIDIA, AMD, and other platforms, are due to arrive in the coming days.
Getting started is almost anticlimactically short. Install the latest release of Ollama and run a single command.
ollama run muse-glimmer:30b-mlx
Muse Glimmer also supports controllable reasoning strength, with four levels available: low, medium, high, and xhigh[1]. Ollama suggests high or xhigh for complex coding and agentic work, and lower strengths when speed matters more. On local hardware, where compute is finite, having that dial makes a real difference to day-to-day usability.
Swapping the brains of Claude Code and Codex for a local model
The most striking part of this release is how little you have to change. Ollama documents running Claude Code on Muse Glimmer with the following command[1].
ollama launch claude --model muse-glimmer:30b-mlx
The same pattern applies to Pi, a lighter-weight coding agent, and to long-running personal assistant frameworks such as OpenClaw and Hermes.
ollama launch pi --model muse-glimmer:30b-mlx
ollama launch openclaw --model muse-glimmer:30b-mlx
ollama launch hermes --model muse-glimmer:30b-mlx
The launch command reportedly works with Codex, OpenCode, GitHub Copilot, and more as well[1]. The agent shell stays exactly as it is while only the underlying model moves local, which lowers the barrier for anyone who has already settled into a workflow.
DFlash delivers 1.5x to 1.8x, and a 1.8B perception encoder reads images
Ollama's MLX engine now supports DFlash, building on its earlier multi-token prediction (MTP) support. With it, Muse Glimmer runs 1.5 to 1.8 times faster on Apple Silicon[1].
DFlash is a form of speculative decoding. Meta explains that Muse Glimmer ships with a lightweight drafter model that proposes entire blocks of tokens at once, after which the main model verifies those proposals in parallel, accepting the correct tokens and correcting the wrong ones[2]. The result is faster generation than conventional token-by-token decoding with identical output quality. Meta's own measurements show speedups of 3.1 times on an RTX 5090, 1.8 times on an M5 Max, and 1.5 times on an M4 Max[2].
Image input is the other headline capability. Muse Glimmer carries a dedicated 1.8-billion-parameter perception encoder that gives it native image understanding[1]. Now that the MLX engine handles image input, coding agents can make low-latency, back-to-back tool calls on tasks that involve images. Ollama lists building websites or applications from a drawing or mockup, computer use applications driven by screenshots, and reading documents, receipts, and charts[1].
Why 30 billion parameters fit on consumer hardware
Meta describes Muse Glimmer as an open-weights model released under the Apache 2.0 license and optimized for always-on local agent workflows[2]. Pre-training used logit distillation on outputs from the larger teacher model Muse Spark, followed by mid-training on longer-context, agent-heavy data and post-training that combined supervised fine-tuning with on-policy distillation and reinforcement learning.
The size problem was solved with quantization. A 30-billion-parameter model at full precision would need more than 55 GB of memory, well beyond any consumer GPU[2]. Compressing the weights to roughly 4-bit precision shrinks the language model to under 20 GB, leaving enough headroom to run the KV cache, the perception encoder for image understanding, and the speculative decoding drafter simultaneously inside a 24 GB or 32 GB envelope. Meta states that this compression introduces minimal to no degradation on agentic tasks.
On performance, Meta reports that Muse Glimmer holds its own against same-class models Gemma4-31B and Qwen3.6-27B across several widely used benchmarks[2]. The weights are distributed on Hugging Face[3], and beyond Ollama, Meta points to LM Studio, Unsloth, llama.cpp, ExecuTorch, vLLM, and SGLang as routes to running the model[2].
Summary
Ollama's Muse Glimmer support is, for now, an initial rollout confined to the MLX engine on Apple Silicon. Even so, the combination on offer is significant: a 30-billion-parameter multimodal model squeezed under 20 GB through 4-bit quantization, accelerated by DFlash, capable of taking image input, running on a Mac, and able to drop straight into Claude Code or Codex in place of a cloud model. That pushes the practical floor for local agents noticeably higher. With NVIDIA and AMD support expected within days, the range of agent setups that never need a network connection looks set to widen further.
Source [1]: https://ollama.com/blog/muse-glimmer
Source [2]: https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model
Source [3]: https://huggingface.co/meta-models/Muse-Glimmer-30B
