DeepSeek opened a limited-time beta of an intermediate build of DeepSeek V4.1 Flash to developers on September 8. It is the largest structural change to the V4 line since the series arrived in April, and the centerpiece is that image and audio handling now sits inside the model rather than in a bolt-on extension. This is not framed as a formal release, and the window closes on September 10.
From an Extension Pack to the Model Itself
This is not DeepSeek's first move into multimodal work. On August 21 the company added V4-Flash-Vision-Exp to its API, a text model with a vision extension pack layered on top. That build was aimed at describing images, reading text out of screenshots, and analyzing charts.
What V4.1 Flash changes is the assembly itself. Instead of attaching vision after the fact, multimodal handling is woven into the model structure, so text, image, and audio inputs are processed on the same footing. DeepSeek describes the intermediate build as more capable, faster to generate, and cheaper to run.
The venue for the announcement says something too. It went out through the company's official community group rather than its social accounts, which frames this as groundwork for the next production model rather than a product launch.
Calling It Is Just a Model Name Change
The barrier to trying it is deliberately low. Developers already on the API keep their existing base_url and simply swap the model name.
model = "deepseek-v4.1-flash-expires-on-0910"
No beta application is required, and billing stays on the same terms as the existing deepseek-v4-flash. The constraint is a cap of 20 concurrent requests per account, which is not enough to push production traffic through.
The date at the end of the model name is the deadline: the build goes offline automatically on September 10. Technical specifications, benchmark figures, and standalone pricing have not been published, so the only available evidence is what developers get from calling it directly.
Third-Party Measurements Arrived First
With no official numbers on the table, the developer community filled the gap. Reported figures include time to first token of roughly 178 milliseconds and sustained generation around 355 tokens per second, with peaks above 365 tokens per second. One test produced a 1,500-token block of code in 4.40 seconds, with the built-in thinking phase clearing in about 1.1 seconds.
The same measurements put the higher-tier V4 Pro at 63 tokens per second with 766 milliseconds to first token, which is a wide gap. Other testers reported a range of 300 to 600 tokens per second, though some suspect those figures reflect a separately provisioned beta environment. None of this is confirmed by DeepSeek, so whether the production release lands in the same territory is still unknown.
The Real Question Is Whether Flash Can Replace Pro
The most revealing part of the beta is the survey DeepSeek circulated alongside it, which asks developers whether the intermediate V4.1 Flash build could fully replace V4 Pro in production.
Flash has been the tier that prioritizes speed and price, while Pro is the tier that prioritizes capability. DeepSeek is now checking whether the lower tier can carry the upper tier's workload. If it can, developer decision-making shifts from "Flash for speed, Pro for accuracy" to "try Flash first and see whether it holds." The difference matters most where call volume drives cost directly, such as long-running agent workflows and batch code generation.
Seen from the other side, the 20-request cap and the roughly two-day window are settings designed to collect impressions, not to serve traffic. What DeepSeek appears to be measuring is not a speed record but the line at which Flash-tier pricing stops being enough for real work.
Summary
V4.1 Flash is an intermediate build that pulls previously bolt-on multimodal support into the model itself, released to developers only through September 10. Getting to it takes nothing more than a model name change, and billing is unchanged, but the 20-request cap and the short window make clear that the goal is feedback rather than deployment. With no official benchmarks published, the speed numbers should be treated as reference points for now. The part worth watching is DeepSeek's attempt to claim Pro territory from the Flash tier, and that answer will have to wait for the production release.
