Chinese AI company Z.ai officially announced its large language model "GLM-5.2" on June 17, 2026 (Japan time). With roughly 753 billion total parameters, it outperforms GPT-5.5 on several software-development benchmarks and ranks 2nd in the world on a coding leaderboard. Notably, the model itself is distributed as an open model under the MIT License, so anyone can download and use it.
Launched as a Freely Available Open Model
GLM-5.2 was first mentioned on June 13, 2026, when Anthropic shut down the "Claude Fable 5" and "Claude Mythos 5" services. At that point only its direction had been revealed — strong coding ability, 1-million-token input support, and continuous execution of long-horizon tasks — but this official announcement filled in the concrete performance details.
The biggest feature is that the model weights themselves are released as an open model under the MIT License. They can be obtained for free from "zai-org/GLM-5.2" on Hugging Face, allowing companies and individuals to download and evaluate the model or fine-tune it for their own needs. In contrast to many cutting-edge models that can only be reached through the cloud, GLM-5.2 keeps open the option of running it in your own environment.
A MoE Design That Stays Light Despite 753 Billion Parameters
GLM-5.2 adopts a Mixture-of-Experts (MoE) architecture that activates only the experts it needs. Although the total parameter count is about 753 billion, only around 40 billion are actually active in a single inference, a design that balances sheer scale with light processing.
It also refines efficiency when handling long contexts. A mechanism called "IndexShare," which reuses computation results across multiple sparse attention layers, is said to substantially cut the per-token computation even when processing the vast 1-million-token context. The goal is to ease, through the architecture itself, the long-standing weakness in which computational load spikes as the context grows.
Strength Shown in Coding and Long-Horizon Tasks
What Z.ai emphasizes most is performance on coding and long-horizon tasks. On "SWE-bench Pro," which measures real bug-fixing ability, it scored 62.1, beating GPT-5.5's 58.6. On the more difficult "FrontierSWE," it reached 74.4 percent, surpassing GPT-5.5's 72.6 percent and closing in on Claude Opus 4.8's 75.1 percent.
It was also highly rated in blind evaluations where humans judge quality without knowing the model's name. On "Code Arena," which compares coding performance, it ranked 2nd in the world, edging out both Claude Opus 4.7 and the latest Claude Opus 4.8. In some comparisons it was even rated above "Claude Fable 5," whose service had just ended. That said, it is worth keeping in mind that many of these figures were published by Z.ai or by the various leaderboards themselves.
How to Use It, the Cost, and Points to Watch
GLM-5.2 can be used through Z.ai's chat AI service and API, and it also works with more than 20 coding environments such as "ZCode," "Claude Code," and "OpenCode." Being easy to drop into the development tools you already use is a practical plus. On cost, the per-token price is said to be about one-sixth that of GPT-5.5, which makes it easier to keep expenses down for workloads that run long-running agent processes.
At the same time, some caution is warranted. While running it locally as an open model is unlikely to be a problem, using Z.ai's hosted API for business has prompted concerns that China's data regulations could come into play over how input data is handled. Companies dealing with highly sensitive information will want to carefully weigh on-premise operation against the use of an external API.
Summary
GLM-5.2 releases a large 753-billion-parameter model as an open model under the MIT License, demonstrating performance that surpasses GPT-5.5 in coding and long-horizon tasks. It ranks 2nd in the world on Code Arena, at times approaching the latest Claude Opus. As a high-performance model you can run on your own hardware, it becomes a strong option for developers — though it is worth keeping an eye on how data is handled when using the hosted API for business.
