DeepSeek has published DeepSeek-V4-Flash-0731, promoting its lightweight large language model from preview to official release. The shape of the model is untouched: 284B total parameters, 13B active at inference, and a 1M-token context window. Only the post-training changed. Even so, the agentic benchmarks now land above the preview build of the larger V4-Pro. The weights ship under the MIT license.

The design stayed put, the post-training did not

The first thing to register about this update is that the skeleton of the model was left alone. DeepSeek states plainly in the model card that the architecture and size match the preview build, and that the improvements come from redoing post-training rather than from a new design.

Underneath, V4-Flash is a 284B-parameter mixture of experts (MoE) that activates only 13B per token. Each MoE layer holds 1 shared expert and 256 routed experts with an intermediate dimension of 2048, and 6 routed experts fire per token. The first 3 MoE layers use hash routing instead.

Attention is hybrid, pairing Compressed Sparse Attention (CSA) with Heavily Compressed Attention (HCA). Manifold-Constrained Hyper-Connections (mHC) stand in for conventional residual connections, with an expansion factor of 4 and 20 Sinkhorn-Knopp iterations. Pre-training ran on more than 32 trillion tokens using the Muon optimizer.

One detail worth clearing up: the Hugging Face repository lists 304B parameters. That figure includes the bundled DSpark speculative decoding draft module on top of the 284B base, so the model itself has not grown.

Agentic scores now sit above the larger model's preview

The clearest gains show up in coding and agent work. On Terminal Bench 2.1, which measures terminal operation, the score climbed from 61.8 in the preview to 82.7. The preview build of the larger V4-Pro sits at 72.1, so the lightweight model has passed its bigger sibling here.

DeepSWE, which covers software engineering tasks, moved even further: from 7.3 to 54.4, better than a 7x improvement. DeepSeek also reports 68.7 on DSBench-FullStack and 59.6 on DSBench-Hard, both internal test sets.

Outside evaluation points the same way. Artificial Analysis, which aggregates benchmark results, placed V4-Flash-0731 at 50 on its Intelligence Index, 10 points above the earlier V4-Flash.

Taking the published figures at face value would still be unwise. DeepSeek notes that code agent tasks ran under the minimal mode of its own DeepSeek Harness, and that harness has not been released. Agent scores move with the harness, so plan on measuring again in your own environment.

Output pricing lands at roughly a third of the larger model

API pricing was left unchanged, and it remains aggressive. Input for deepseek-v4-flash costs 0.14 USD (about 22 yen) per 1M tokens on a cache miss and 0.0028 USD (about 0.4 yen) on a cache hit. Output runs 0.28 USD (about 44 yen) per 1M tokens. The concurrency ceiling is set at 2,500.

For comparison, output on the larger deepseek-v4-pro costs 0.87 USD (about 137 yen) per 1M tokens. V4-Flash comes in at roughly a third of that. Agent loops burn output tokens quickly, so the gap translates directly into running costs. This is a price band where solo developers and internal platform teams without a GPU budget can afford to keep agents running.

On the interface side, deepseek-v4-flash now supports the Responses API format natively and has been adapted for Codex. The V4-Pro API and the app and web models were not part of this update.

Self-hosting demands more than 100 GB of memory

The weights are MIT-licensed and ungated, which clears the way for commercial on-premise deployment. The hardware bar, though, is a separate matter entirely.

Only 13B activate per token, but all 256 experts stay resident in memory. The vLLM recipe serves the model on a single node with four GB300 GPUs. If you push into quantization, Unsloth's dynamic GGUF builds put the lossless 8-bit version at 162 GB and the 3-bit version at 103 GB, requiring roughly 110 GB of combined RAM and VRAM.

Practically, self-hosting suits mid-size and larger companies with a serving cluster, or a single well-specced workstation running aggressive quantization. This is not a model you casually spin up on a laptop.

Implementation details worth knowing in advance

A few things tend to trip up integrations. There is no Jinja chat template. DeepSeek ships an encoding/ folder instead, providing encode_messages and parse_message_from_completion_text. Existing code that assumes a template will need rewriting.

Reasoning depth is controlled through reasoning_effort, which accepts low, high, and max. At high and max, output can extend to 384K tokens. For sampling, DeepSeek recommends temperature 1.0 with top_p 0.95 for agentic use, and 1.0 for both otherwise.

Enabling the bundled DSpark takes a single vLLM flag.

--speculative-config '{"method":"dspark","num_speculative_tokens":7,"draft_sample_method":"greedy"}'

The DSpark paper reports 60 to 85 percent faster per-user generation compared with an MTP-1 baseline at matched aggregate throughput. Since that lands squarely on perceived speed, it is worth turning on from the start when self-hosting.

Summary

DeepSeek-V4-Flash-0731 leaves the design alone and rebuilds the post-training stage, and the payoff shows in agentic and coding results. Output pricing sits at about a third of the larger model, and the weights are open under MIT. Set against that, the published scores are vendor-measured on an unreleased harness, and self-hosting needs north of 100 GB of memory. Anyone weighing adoption should start by running evaluations on their own tasks.

References