On April 24, 2026, China's DeepSeek released its new large language model, DeepSeek V4, as a preview. It carries 1.6 trillion total parameters and uses a Mixture-of-Experts (MoE) design that activates only the parts it needs. The weights are distributed under the commercially friendly MIT License, while its API pricing lands at roughly one-sixth that of the leading closed models. The default context length reaches 1 million tokens, and the model is already available through Hugging Face and DeepSeek's API. Arriving 484 days after V3, it is being described as a "second DeepSeek moment."

Competing on efficiency: frontier-class capability drops into a cheaper tier

The most immediate impact of this release is pricing. The higher-end DeepSeek V4-Pro is priced through the API at 0.435 USD (about 71 yen) per 1 million input tokens and 0.87 USD (about 143 yen) per 1 million output tokens. With cached input, that input price falls to 0.003625 USD (about 0.6 yen).

The lighter V4-Flash is cheaper still, at 0.14 USD (about 23 yen) per 1 million input tokens and 0.28 USD (about 46 yen) per 1 million output tokens. Against published list prices, V4-Pro undercuts a top-tier closed model such as Gemini 3.1 Pro (2.00 USD per 1 million input tokens and 12.00 USD per 1 million output tokens, for prompts up to 200,000 tokens) by roughly five times on input and fourteen times on output. The gap shifts with a workload's input-to-output mix, but overall it lands in the neighborhood of one-sixth. This does not make intelligence free, but it clearly widens the range of tasks that were previously "too expensive to automate" and can now pay off.

Close to the leaders, but not a clean sweep

On the benchmarks, the numbers stand out for an open-weight model. On BrowseComp, which tests demanding information-seeking, V4-Pro scored 83.4 percent, closing in on the top closed models. It reached 90.1 percent on GPQA Diamond, which probes graduate-level science knowledge, 55.4 percent on SWE-Bench Pro for software engineering, and 73.6 percent on MCP Atlas for tool integration.

That said, on many of the directly comparable tests, the newest top-tier closed models from U.S. players still lead. The value of DeepSeek V4 is less about "taking every first place" and more about "getting close while being dramatically cheaper and openly available." The evaluation firm Vals AI ranked V4 as "the number one open-weight model, and it's not close" on its own coding benchmark.

The architecture behind 1 million tokens

V4's technical centerpiece is making a 1-million-token context usable at a realistic cost. DeepSeek adopted a hybrid attention design that combines Compressed Sparse Attention (CSA), which compresses the dimensionality of the input, with Heavily Compressed Attention (HCA), which aggressively compresses the memory needed for long-range dependencies. As a result, even when running at 1 million tokens, V4-Pro reportedly uses only 10 percent of the KV cache and 27 percent of the per-token inference compute compared with the previous V3.2.

To train 1.6 trillion parameters stably, DeepSeek replaced conventional residual connections with a scheme it calls Manifold-Constrained Hyper-Connections (mHC), aimed at widening the flow of signals across layers while keeping training from destabilizing. It used the Muon optimizer and pre-trained on more than 32 trillion tokens of high-quality data. Thanks to the MoE design, only 49 billion of the 1.6 trillion parameters are active on any given token, which holds down compute. Inference offers three switchable modes: "Non-think" for fast replies, "Think High" for careful logical analysis, and "Think Max" for pushing reasoning to its limit.

An open distribution model, with an eye on reducing NVIDIA dependence

Drawing as much attention as the model itself was the surrounding software. DeepSeek revealed that it validated its Expert Parallelism scheme, which distributes experts across hardware, on Huawei's Ascend NPUs rather than NVIDIA parts. It reported a 1.50x to 1.73x speedup on non-NVIDIA platforms, sketching a path to deployment that is less exposed to Western GPU supply chains and export controls. It did note, however, that officially licensed NVIDIA GPUs were also used to train V4 itself.

DeepSeek also open-sourced its MegaMoE kernel, which it says speeds up latency-sensitive workloads by up to 1.96x. The license is the industry's most permissive, MIT, allowing modification, redistribution, and commercial use free of charge. The model integrates out of the box with developer agents such as Claude Code, OpenClaw, and OpenCode, making it easy to slot into existing tooling.

Legacy models retire today, and 1 million tokens becomes the standard

DeepSeek is also moving quickly to cut off older generations. The legacy deepseek-chat and deepseek-reasoner endpoints shut down completely at 15:59 UTC on July 24, 2026, with all traffic rerouted to V4-Flash. It is close to a declaration that million-token support is now the baseline across all services.

The overall picture is a claim, pressed on both price and openness, that architectural ingenuity can approach frontier performance without brute-force scale. For companies that have charged high fees for closed models, it becomes harder to justify the premium on capability alone.

*1 USD = 164 JPY (as of July 24, 2026)

Sources: DeepSeek API Docs: Models & Pricing / DeepSeek-V4 Preview Release / Gemini Developer API pricing

Summary

DeepSeek V4 is one of the largest openly licensed models yet, pairing a 1.6-trillion-parameter MoE with a 1-million-token context under the MIT License. It approaches the top closed models — scoring 83.4 percent on BrowseComp, for instance — while holding its API pricing to about one-sixth of theirs. Between validating on Huawei's Ascend NPUs and open-sourcing its kernel, it reads as a move to pull the initiative toward the "cheap and open" side of AI development. The legacy endpoints closed today, and 1 million tokens is becoming the new standard.