Z.ai released GLM-5.3-Flash on August 26, 2026, the first natively multimodal model in the GLM-5 family. It carries 320 billion total parameters, activates 18 billion per token, and handles a context window of one million tokens. The weights ship under the MIT License, so commercial use comes with almost no strings attached. The design is aimed squarely at running coding and agent workloads cheaply.

Only 18 billion of 320 billion parameters move

GLM-5.3-Flash is a Mixture-of-Experts model. Total parameters reach 320 billion, but only about 18 billion are actually computed for each token. GLM-4.5 activated 32 billion, so the working set has been cut nearly in half. The priority here is clearly cheap, fast inference.

Training used a multimodal corpus of roughly 30 trillion tokens. The attention stack combines two mechanisms: linear attention for local dependencies, and sparse attention to pull in the distant context that matters. As the context approaches one million tokens, the indexer key vectors pile up and latency suffers, so a component called IndexPool compresses groups of them. Z.ai reports three times less attention compute and a 4.4-fold reduction in KV cache size compared with GLM-5.3.

Benchmarks move sharply from GLM-5.2

On the coding side, DeepSWE v1.1 comes in at 63.4, up from 46.2 for GLM-5.2. AutomationBench, which measures automation ability, scores 48.8 against GLM-5.2's 26.2, close to double. On the Artificial Analysis Intelligence Index v4.1.1 the model scored 57 at a discounted cost of 0.045 USD (about 7 yen) per task.

Z.ai says the model beats GLM-5.2 across the reported coding and agentic tests at one-tenth the price, and claims it lands within half a point of Claude Opus 4.8 on its internal coding benchmark.

That said, numbers like these shift with the evaluation harness, the context limit, and the generation settings. Whether two scores were measured on the same footing is not obvious without checking each setup, so the table is worth reading as a set of conditions rather than a ranking.

※1 USD = 160 JPY (as of August 31, 2026)

Trained to look at its own screen and fix it

The headline capability in this release is visual verification. Z.ai trained the model to inspect rendered interfaces, gameplay footage and 3D output, then judge and revise its own work from what it sees. Writing code is not the end of the loop; the model can check what actually appeared on screen and go back in.

The same approach extends to documents, spreadsheets, presentations, dashboards and meeting materials. The goal is reasoning that spans text, images and structure at once, which should show up most clearly in front-end work and document production.

What "ox-alpha" turned out to be, and the domestic silicon behind it

Before launch, the model sat anonymously on OpenCode and OpenRouter under the name ox-alpha. It drew attention without an identity attached, and Z.ai says it became the most used model of the week on those services. This release formally confirms that ox-alpha was GLM-5.3-Flash.

One more detail deserves attention: inference runs entirely on Chinese AI chips. Z.ai built an SGLang-based stack that separates encoding, prefill and decoding, and reports a threefold gain in end-to-end serving performance across tens of thousands of domestic accelerators. With access to foreign GPUs hard to plan around, demonstrating that inference at this scale can run on domestic silicon carries weight independent of the model's benchmark scores.

Where to run it and what it costs

GLM-5.3-Flash is available to every GLM Coding Plan user, with three times the usable quota of GLM-5.3. The multimodal features are exposed in ZCode through Browser Use and Computer Use.

The weights are distributed on Hugging Face, and local deployment is supported by SGLang, vLLM and TokenSpeed. Because the license is MIT, both internal use and embedding the model in a product are straightforward.

Summary

GLM-5.3-Flash is a multimodal Mixture-of-Experts model with 320 billion total and 18 billion active parameters, offering a one-million-token context under the MIT License. Coding and automation benchmarks climb sharply over GLM-5.2 while the stated price drops to one-tenth. The release also revealed that the anonymous ox-alpha was this model, and that its inference runs entirely on Chinese chips, which leaves plenty to discuss beyond the scores. Since those scores depend on evaluation conditions, the reliable way to judge it is to run it yourself.

※The thumbnail image is AI-generated.