Google has released Gemma 4 QAT, a memory-efficient version of its open Gemma 4 AI models. By building quantization into the training process itself, it sharply cuts memory use while limiting the drop in response quality. The lightest configuration runs in just 0.84GB of memory, making it possible to run AI locally on an ordinary smartphone or laptop. The models are free to download under the Apache License 2.0.

The Hurdle of Running High-End AI Locally

When you run an AI model locally on a PC or phone, it is first loaded into fast VRAM. Whatever does not fit spills over into slower RAM, and if that is not enough, into a swap file on the SSD, slowing things down. To run smoothly, you need a model that fits within VRAM.

The problem is that high-performance models often need tens to hundreds of gigabytes of memory, far beyond consumer hardware. The common workaround has been quantization, a technique that lowers numerical precision to reduce memory use.

How QAT Limits the Quality Drop

Quantization saves memory, but because it lowers precision, response quality tends to suffer. Most widely available quantized models are produced by quantizing a finished model after the fact, which makes them especially prone to this effect.

Gemma 4 QAT instead uses Quantization-Aware Training (QAT). Because it simulates the quantized state during training, the final quantized model holds up far better in quality. Achieving memory savings and quality retention at the same time is the heart of these models.

Supported Models and the Memory Savings

Gemma 4 QAT covers every Gemma 4 variant: E2B, E4B, 12B, 26B A4B, and 31B. On top of that, E2B and E4B come in versions optimized for mobile use.

The memory savings are substantial. The original Gemma 4 E2B needs 11.4GB, but the QAT version (Q4_0, 4-bit) shrinks that to 2.9GB. The mobile version comes in at 1.1GB, and a text-only model that drops image and audio recognition runs in as little as 0.84GB. Google says the other model sizes also achieve large memory reductions while keeping the quality drop in check.

How to Get It and Where It Runs

The Gemma 4 QAT models are distributed in a collection on Hugging Face and can be downloaded by anyone for free. They are released under the Apache License 2.0, which permits a wide range of uses including commercial ones. For runtimes, Google states that the models work as-is with the popular local-AI tools llama.cpp, Ollama, and LM Studio. Being easy to slot into the tools people already use, and ready to try right away, is a real advantage.

Summary

Gemma 4 QAT is a set of models that achieves both low memory use and quality retention through QAT, which bakes quantization into training. With configurations that cut memory from the original 11.4GB down to as little as 0.84GB, running high-end AI locally on a smartphone or laptop has become far more realistic. Free to use under the Apache License 2.0 and ready to run in llama.cpp, Ollama, and LM Studio, it looks like an option well worth trying for anyone who wants to run AI on their own device.