Tencent's Hunyuan team released Hy4 preview, a next-generation large language model, on August 28, 2026, and opened the weights at the same time. The model carries 770 billion total parameters, activates 49 billion of them per token, and handles a context window of more than 1 million tokens. It ships under the Apache 2.0 license, which places few restrictions on commercial use.
Only 49B of the 770B Runs at Any One Time
Hy4 preview uses a Mixture-of-Experts (MoE) design. Many small expert networks sit inside the model, and only a subset fires depending on the input. So while the total parameter count is 770 billion, just 49 billion are actually computed for each token.
The backbone runs 78 layers. The first layer uses a conventional dense FFN, and the remaining 77 replace it with MoE. Each MoE layer holds 256 routed experts plus 1 shared expert, and every token activates the top 8 routed experts alongside the shared one. Hidden size is 6,144 and the vocabulary comes to 120,832 entries.
For attention, the team adopted Gated DSA (Gated DeepSeek Sparse Attention), drawing on work from DeepSeek and GLM, combined with IndexCache for reusing sparse indices across layers. Both are there to keep a 1 million token context computationally practical. A native MTP layer for speculative decoding (10 billion total parameters, 0.7 billion activated) is built in as well, and an FP8-quantized build called Hy4 preview-FP8 shipped alongside the main release.
A Blind Evaluation Scored by 163 In-House Experts
Tencent positions the model as built for real work rather than for benchmarks. Training data was created together with in-house specialists including software engineers, game developers, financial analysts, and security staff.
To measure the payoff, 163 internal experts ran a blind side-by-side evaluation across 203 engineering tasks. Hy4 preview averaged 2.99 out of 4.00, edging past GLM 5.3 at 2.92 and Kimi K3 at 2.94. Against GLM 5.3 the split was 46.8 percent wins, 12.8 percent ties, and 40.4 percent losses; against Kimi K3, 51.2 percent wins, 7.9 percent ties, and 40.9 percent losses. Read plainly, these are narrow wins rather than a rout.
Published benchmark figures include 92.3 on GPQA Diamond, 65.7 on SWE-bench Pro, 82.9 percent resolved on SWE-bench Multilingual, 64.3 on Deep SWE, and 62.9 on SkillsBench v1.1.
The Model Joined In on Optimizing Its Own Training and Inference
The most striking part of the announcement is that Hy4 preview took part in its own development. For the first time, the model was involved in automated optimization of training methods, data strategies, evaluation frameworks, and low-level operators. It proposed approaches, ran experiments, and iterated on the results, with the resulting code, logs, and feedback flowing into the next round. Tencent describes this as an early-stage recursive self-improvement loop.
The same thing happened on the serving side. The model analyzed bottlenecks in the inference system it runs on and applied several rounds of optimization to areas such as operator fusion and communication. End-to-end throughput rose 31.8 percent against the baseline, and the gains held consistently across different context lengths and concurrency levels.
Tencent is candid that this is unfinished work. Known issues at release include spending longer than necessary reasoning through complex tasks and a tendency to over-verify its own output. The team plans to keep the same approach it used with the previous generation, Hy3: ship early and listen for what breaks.
Where to Run It, and What It Costs
Weights are distributed on Hugging Face, ModelScope, GitCode, and CNB. For self-hosting, official vLLM and SGLang images are available, each with its own recipe. The recommended vLLM setup spreads the FP8 build across 8 GPUs.
docker run --gpus all -p 8000:8000 --ipc=host \
vllm/vllm-openai:hy4-preview tencent/Hy4-preview-FP8 \
--tensor-parallel-size 8 \
--attention-backend FLASHMLA_SPARSE \
--served-model-name hy4-preview
If you would rather not line up GPUs yourself, the model can be tried directly through Tencent products including WorkBuddy, CodeBuddy, Yuanbao, and ima. WorkBuddy and CodeBuddy offer it free for two weeks from launch. Free access to the previous generation, Hy3, on both platforms has also been extended through September 30.
API access runs through Tencent Cloud TokenHub and OpenRouter. Pricing per million tokens is 0.834 USD (about 130 yen) for input, 2.501 USD (about 400 yen) for output, and 0.042 USD (about 7 yen) for cache hits. For a model in the 770 billion parameter class, that is a restrained price tag.
※1 USD = 159 JPY
Summary
Hy4 preview is a MoE model with 770 billion total parameters and 49 billion activated, opening a 1 million token context under Apache 2.0. It edged past GLM 5.3 and Kimi K3 in an internal blind evaluation, and it stands out for having taken part in optimizing its own training and inference, lifting throughput by 31.8 percent. Tencent treats this strictly as a preview, and says the next models in the Hy4 series are due soon.
※The thumbnail image is an AI-generated illustration.
