Ollama, the tool for running large language models (LLMs) locally, has released its latest version, 0.30[1]. Through the open-source inference engine llama.cpp, it now supports models in the GGUF format, letting a wider range of models and hardware run than before[1]. On NVIDIA hardware, throughput is now up to 20% faster, and with Vulkan—the technology used to tap the GPU—enabled by default, AMD and Intel devices can use the GPU without any extra setup[1].

Up to 20% Faster on NVIDIA, with Vulkan Widening GPU Support

In version 0.30, performance on NVIDIA hardware is now up to 20% faster[1]. This gain comes from optimizations contributed by the NVIDIA and llama.cpp teams[1]. The figure was measured by running the Gemma 4 26B model, quantized to the Q4_K_M format (a quantization scheme that makes models lighter), on an NVIDIA RTX 5090[1].

Vulkan, a cross-platform technology for GPUs, is now enabled by default[1]. This extends GPU acceleration to AMD and Intel devices as well[1]. More users can now run models on the GPU right out of the box, without having to install vendor-specific libraries[1]. It is a step that meaningfully widens the range of hardware on which you can try local AI.

A Wider Choice of Models Thanks to GGUF Support

At the heart of 0.30 is support for the GGUF format[1]. GGUF is the model file format that llama.cpp handles natively, and supporting it broadens the set of models from the GGUF ecosystem that run as-is[1]. Specifically, that now includes model families such as LFM and Prism, as well as fine-tuned models published by Unsloth[1]. This support does not replace the MLX engine prepared for Apple silicon; rather, it augments it to broaden compatibility[1].

The blog also shows how to run GGUF models distributed on Hugging Face and elsewhere[1]. First, download the GGUF file, or a directory containing GGUF files[1]. Next, prepare a Modelfile (a configuration file for the model) with a FROM command that points to that path[1].

FROM ./my-model.Q4_K_M.gguf

Then create and run the model with the following commands[1].

ollama create -f Modelfile my-model
ollama run my-model

Working with Coding Agents

If a model supports tool calling (the ability to invoke external tools to carry out tasks), that capability carries over on Ollama as well[1]. This lets you combine the model with your favorite coding agents and personal assistants in a single command[1].

For example, to combine it with Claude Code, you specify the following[1].

ollama launch claude --model my-model

For Hermes Agent, the command is as follows[1].

ollama launch hermes --model my-model

For OpenClaw, run the following[1].

ollama launch openclaw --model my-model

You can check whether a given GGUF file supports tool calling with the ollama show command[1]. If the displayed information includes a tools entry, that is the sign it supports tool calling[1].

ollama show my-model

In releasing this update, Ollama offered thanks to Georgi Gerganov and the maintainer teams behind llama.cpp, as well as to hardware partners such as NVIDIA, AMD, Qualcomm, and Intel, who worked to optimize performance within the GGML ecosystem[1].

Summary

Ollama 0.30 is an update that greatly broadens the range of supported models and hardware through GGUF support via llama.cpp. On NVIDIA hardware, throughput is up to 20% faster, and enabling Vulkan by default lets AMD and Intel devices use the GPU without extra setup. GGUF models can be brought in via a Modelfile, and models that support tool calling can be paired with coding agents such as Claude Code and OpenClaw in a single command. It looks set to further widen the options for running local AI on your own hardware.

Source: https://ollama.com/blog/improved-performance-and-model-support-with-gguf