Alibaba's Qwen team made its flagship large language model, Qwen3.8-Max, generally available on August 3. The headline number is 2.4 trillion parameters, but the more interesting figure is a stretch of 16 days during which the model wrote code without a single human intervention. Weights for this model and a smaller sibling are due next week, adding more momentum to the open-weight race.

Only 95 billion of the 2.4 trillion actually fire

Qwen3.8-Max uses a Mixture-of-Experts design that activates just the expert networks it needs for each token. Total parameters come to 2.4 trillion, but only 95 billion are active at inference time. The point is to keep the knowledge capacity of a very large model while holding down the compute and latency cost of every response.

The context window reaches roughly 1 million tokens. Maximum input is about 991,000 tokens in standard mode and about 983,000 tokens with thinking enabled, with a maximum output of 131,000 tokens. The model reads images and video natively as well as text, so workflows such as building a website from a screenshot are in scope.

Sixteen days from an empty folder, with nobody watching

The capability Qwen is pushing hardest is sustained software development with no human in the loop. Starting from an empty folder, a roughly 16-day fully autonomous run left the repository with 265 commits, 127 pull requests and 151 issues. That is not one-shot code generation but more than two weeks of finding problems, fixing them and extending features.

For comparison, the previous generation Qwen3.7-Max was being discussed in terms of 35-hour autonomous runs. This is an order of magnitude beyond that.

Reproducing a paper, then beating its method

Qwen also published a research task. The subject was a paper on selecting training examples that sit near the decision boundaries a language model finds hardest. Qwen3.8-Max was handed the paper and a GPU, nothing else, so it had to build the data processing, training and evaluation code and the experimental environment from scratch.

The run took about 125 hours, produced roughly 7,600 lines of code, involved more than 1,100 operations and 33 GPU training runs. The first 37 hours or so reproduced the paper's 6 main results; the remaining 88 hours cycled through hypothesis, implementation, experiment, analysis and retry for 4 rounds covering 18 candidate improvements. It arrived at its own selection method and lifted the AIME24 math benchmark from the reproduced 49.58 percent to 52.29 percent, a gain of 2.71 points.

There was also a head-to-head against humans. Entered into an Alibaba Cloud multimodal dialogue intent recognition contest, the model built a classification system on its own within 24 hours and raised accuracy from 0.600 to 0.853 across 45 submissions. It finished ahead of 458 of the 526 participating teams, inside the top 13 percent[1].

A 365-day e-commerce simulation returned 4.16x

To test whether the model learns from experience, Qwen ran a 365-day e-commerce operations simulation built on anonymized data from Alibaba's Taobao and Tmall. Starting capital of 100,000 yuan (about 2.3 million yen) grew to 416,252 yuan (about 9.6 million yen), a 4.16x return. That beat second-placed GLM 5.2 by 38 percent and the previous-generation Qwen3.7-Max by 152 percent.

The reason for the gain is the interesting part. Over more than 2,000 interactions, the model gradually changed how it negotiated with the same suppliers and drove purchase prices down over time. Adjusting strategy based on past exchanges appears to be where long-horizon tasks are won or lost.

Benchmarks show a clear split between strengths and gaps

The published scores draw a sharp line. Terminal-Bench 2.1 came in at 86.6, ahead of the 84.6 posted by Claude Fable 5 and Claude Opus 4.8, but short of GPT-5.6 Sol at 88.8. PaperBench, which measures paper reproduction, reached 93.0 and instruction-following on IFBench hit 82.8, both above Claude Fable 5.

On the other side, SWE-bench Pro at 67.7 and FrontierSWE at 73.5 fell below Claude Fable 5, while Humanity's Last Exam stalled at 43.6 and Toolathlon Verified at 72.5. Editing existing code and handling complex tool use remain the weaker areas. Against the previous generation, DeepSWE 1.1 moved from 21.6 to 56.6 and JobBench from 31.3 to 53.4. On visual and document work, RealWorldQA scored 88.0, MMMU-Pro 82.3 and OmniDocBench 1.5 92.1.

Qwen says that jointly scaling its reinforcement learning environments and compute produced steady, consistent gains across dozens of in-house and public benchmarks[1].

Pricing, and the weights arriving next week

Qwen3.8-Max is available through Qwen Chat and Alibaba Cloud's Model Studio, among other routes. API pricing is 2 USD per million input tokens (about 310 yen), 6 USD per million output tokens (about 940 yen) and 0.25 USD per million cached input tokens (about 40 yen). The model supports function calling, structured output, batch processing, prefix completion and fine-tuning, and the Responses API adds built-in tools for code execution, web search, web page extraction and image search.

Next week, weights for Qwen3.8-Max itself and for the 27-billion-parameter Qwen3.8-27B are set to be released. Running a 2.4-trillion-parameter model in-house calls for data center scale hardware, but 27 billion is well within reach of ordinary GPUs. With Moonshot AI having just published the 2.8-trillion-parameter Kimi K3, the contest to release frontier-class models as open weights looks set to continue.

※1 USD = 157 JPY, 1 CNY = 23 JPY (as of August 5, 2026)

Summary

Qwen3.8-Max is a multimodal Mixture-of-Experts model that activates 95 billion of its 2.4 trillion parameters and handles up to 1 million tokens of context. The 16-day unattended development run, the 125-hour paper reproduction and improvement, and the 4.16x return over a simulated year of e-commerce are all framed as measures of endurance rather than single-response quality. Gaps remain in editing existing code and complex tool use, but the real question is whether next week's weight release puts a model at this level within reach of everyone else.

References