Alibaba's model team Qwen released the open-weight AI model Qwen3.8-Flash-Next on August 26, 2026 (JST). It carries 125 billion total parameters but activates only 6 billion per token, and it ships four new mechanisms that preview the upcoming Qwen4 architecture. The weights are already downloadable from Hugging Face and ModelScope.

125 Billion Total Parameters, 6 Billion Active

Qwen3.8-Flash-Next is a Mixture-of-Experts (MoE) model that handles images and video in addition to text. An MoE model splits its interior into many smaller networks with different roles and fires only the ones each input needs, which keeps the compute cost down. The 125B-A6B label means 125 billion total parameters with 6 billion active.

Separately from the main body, the model carries 51 billion parameters worth of N-gram embeddings. Context length is 262,144 tokens by default and can be extended to 1 million tokens.

Compared With Opus 4.6 and the 27B Model

In benchmark results published by Qwen, the model was measured against Anthropic's Claude Opus 4.6 (Max). It scored 62.5 percent on SWE-bench Pro, which measures software development work, against 53.4 percent for Opus 4.6. On CoWorkBench, which covers long-running office work, it scored 73.9 percent against 68.2 percent. On JobBench, which covers occupational tasks, it scored 55.7 percent against 36.6 percent.

It also came out ahead of Qwen's own Qwen3.7-Plus, a 397-billion-parameter model, on many items, and Qwen says the training cost was roughly one-ninth as much.

The comparison against Qwen3.8-27B, a staple for local AI users, is the more interesting one. The 27B model is a dense design in which all 27 billion parameters run, while Qwen3.8-Flash-Next runs only 6 billion. Even so, Qwen reports parity or better on 22 language and vision benchmarks, with wins on 21 of them. The gap on SWE-bench Pro is narrow at 62.5 percent against 61.7 percent, but JobBench opens up to 55.7 percent against 33.4 percent.

These figures all come from Qwen itself and have not been independently verified by a third party, so they should be read with that caveat in mind.

Four New Mechanisms That Preview Qwen4

Qwen credits four design changes for the performance it gets out of so few active parameters.

The first is the attention layout, attention being the mechanism that decides which parts of the context to focus on. Every four layers combine three layers of Gated DeltaNet, which compresses and stores context, with one layer of Qwen Sparse Attention, which selects and re-reads only the important blocks. At a 90 percent cache hit rate with 1 million tokens of input, prefill throughput reaches 8.6 times that of Qwen3.7-Plus.

The second is Gated Residual, which increases the number of pass-through paths carrying each layer's input into the deeper layers from one to four, with gates controlling how much each path reads and writes. Early information is less likely to wash out by the time it reaches deeper layers, and training is more stable.

The third is the N-gram embedding described above. A lookup table worth 51 billion parameters, held separately from the main body, maps recent token sequences to vectors. Because it is only a table lookup, per-token compute barely increases.

The fourth is Muon, a new method that makes parameter updates more efficient during training. Fewer steps are needed, which is where the roughly one-ninth training cost comes from.

Why Holding 51B in RAM Matters

For anyone running models locally, the decisive point is that the 51 billion parameters of N-gram embeddings can sit in RAM rather than VRAM. Normally, parameters that do not fit in VRAM have to be transferred from RAM to the GPU on demand, which slows things down badly. FreeToken, a runtime that got attention for running 35B-class models on an 8GB GPU, cleared that same wall, but it did so through runtime-side engineering that assumes NVIDIA GPUs.

Qwen3.8-Flash-Next was designed from the model side on the assumption that the table lives in RAM and is prefetched, so it is not tied to a particular GPU environment. For server deployments, Qwen lists the inference frameworks SGLang, vLLM, and TokenSpeed. For local use, llama.cpp support is also mentioned, though it is not yet clear whether the prefetching is implemented there.

Licensing and Commercial Pricing

The license is the Qwen Community License 1.0, which is free in principle, including for commercial use. A separate contract is required for MaaS businesses and for products whose primary purpose is AI coding or office assistance. Products with more than 100 million monthly active users, or more than 20 million USD (about 3.2 billion yen) in monthly revenue, must display the model name in their UI.

The production version, Qwen3.8-Flash, is due shortly through the QwenCloud API. Pricing is listed at 0.16 USD (about 25 yen) per million input tokens and 0.47 USD (about 75 yen) per million output tokens.

※1 USD = 159 JPY

Summary

Qwen3.8-Flash-Next amounts to an early release of the Qwen4 design. The key points are that it posts Opus 4.6-class scores while activating just 6 billion parameters, and that it offloads the heavy part to RAM in a way that does not depend on a specific GPU. If prefetch support in llama.cpp comes together, the range of models people can run on their own machines will widen further.