Moonshot AI has released the weights of its flagship model, Kimi K3, to the public. At 2.8 trillion total parameters, it is the largest freely downloadable open-weight model to date, and the company calls it the world's first open 3T-class model. Kimi K3 itself launched on July 16, with the complete weights following shortly afterward on Hugging Face. Running it locally, however, demands compute far beyond what an individual can reasonably assemble.

Only 16 of 896 experts fire at a time

Kimi K3 uses a Mixture-of-Experts design, in which specialized modules are switched in and out rather than run all at once. Of the 2.8 trillion total parameters, only 104 billion are active per token — just under 4 percent of the model. There are 896 experts, of which 16 are selected for each token, plus 2 shared experts that always run.

The model has 93 layers, and its attention stack is a hybrid: 69 layers of Moonshot AI's own Kimi Delta Attention (KDA) and 24 layers of Gated MLA. Combined with a framework the company calls Stable LatentMoE, Moonshot AI says the design delivers roughly 2.5 times the intelligence per unit of compute compared with the previous Kimi K2. The company goes out of its way to stress that this is not simply a matter of adding parameters, which hints at a desire to stand apart from the raw scaling race.

Other specifications: a context window of 1,048,576 tokens, a 160K vocabulary, and MoonViT-V2 as the vision encoder with 401 million parameters. Text and images are understood inside the same model, in what is usually described as a natively multimodal architecture.

The download runs about 1.56TB, and the bar for self-hosting is high

The weights are stored in MXFP4, a 4-bit format. This is not post-hoc compression: quantization-aware training is applied from the supervised fine-tuning stage onward, with MXFP8 used for activations. Even so, the distribution comes to roughly 1.56 terabytes across 96 shards.

Moonshot AI itself recommends deployment on supernode configurations with 64 or more accelerators for inference efficiency. Research labs, cloud providers, and well-funded startups can run it in-house, but this is not a model an individual developer tries out on a desktop GPU. The recommended inference engines are vLLM, SGLang, and TokenSpeed, and because KDA does not sit well with conventional prefix caching, the company contributed an implementation to vLLM.

The license is not a standard open-source license but a custom Kimi K3 License. The weights can still be obtained, modified, and redistributed, but the terms are worth checking individually.

Benchmarks put it behind the top two

The evaluation table published by Moonshot AI places Kimi K3 behind Claude Fable 5 and GPT-5.6 Sol. The company states plainly that overall performance still trails the most capable proprietary models.

Even so, it leads in a fair number of individual categories. On SWE-Marathon, which covers long-running software engineering, it posts the top score of 42.0. It reaches 91.2 on BrowseComp, which measures navigating the web, and 94.5 on MCPMark-Verified. On GPQA Diamond it scores 93.5, ahead of Claude Fable 5's 92.6. On the other hand, FrontierSWE comes in at 81.2 against Claude Fable 5's 86.6, and its AA-Briefcase Elo of 1548 for knowledge work falls short of 1583.

One notable demonstration: Moonshot AI had Kimi K3 design a chip. In a single 48-hour autonomous run using open-source EDA tools, it produced a circuit that closes timing at 100 MHz within 4 square millimeters and sustains more than 8,700 tokens per second of decode throughput in simulation. The example is clearly meant to show how far the model can carry a task on its own judgment over an extended period.

API pricing keeps the per-token rate low

Self-hosting is not the only option. Selecting the model name kimi-k3 on the Kimi API Platform gives access at 0.30 USD (about 47 yen) per million tokens for cache-hit input, 3.00 USD (about 470 yen) for cache-miss input, and 15.00 USD (about 2,350 yen) for output. Moonshot AI says the official API achieves a cache hit rate above 90 percent in coding workloads.

※1 USD = 157 JPY (based on the August 1, 2026 close)

The model is also available through the Kimi mobile app, the Kimi Work desktop app, and Kimi Code in the terminal.

The company discloses its caveats too

Moonshot AI lists the limitations as well. One concerns thinking history: Kimi K3 was trained on the assumption that all prior reasoning content is passed back, so output becomes unstable on agent harnesses that do not support this, or when a session is switched over from another model mid-way. The other is excessive proactiveness — because training emphasized long, difficult tasks, the model may make decisions on the user's behalf when instructions are ambiguous. The company also acknowledges a remaining gap in user experience against Claude Fable 5 and GPT-5.6 Sol.

The weights were not the only thing released. High-performance attention kernels, the MoE communication library, and infrastructure for running agent environments at scale were opened up alongside them. US labs keep their frontier models closed, while Chinese firms have built their presence by opening weights. With a 3-trillion-parameter-class model now available at no cost, that dividing line looks set to shift further.

Summary

Kimi K3 is the first open 3T-class model, combining a 2.8-trillion-parameter MoE with a 1M-token context window and native image understanding. An extremely sparse design that activates 16 of 896 experts, paired with a custom attention mechanism, is said to deliver roughly 2.5 times the efficiency per unit of compute over the previous generation. Benchmarks place it behind the leading proprietary models, though it comes out ahead on several coding and agentic measures. At roughly 1.56 terabytes, the barrier to running it yourself remains high.

Source: https://huggingface.co/moonshotai/Kimi-K3