DeepSeek made DeepSeek-V4-Flash-Vision-Exp, its first image-capable model, available on the DeepSeek API Platform on August 21. The model is explicitly experimental, yet its text-side behavior matches the existing DeepSeek-V4-Flash and there is no surcharge for image input. Each image is billed at up to 384 tokens, and the Files API launched free of charge on the same day.
Images Are Capped at 384 Tokens
The billing rule for images is arguably the most practical detail here. Images are converted into tokens and billed alongside text as input tokens, but every image is resized before that conversion happens.
Images whose total pixel count falls below roughly 384×384 are scaled up while preserving their aspect ratio. Larger images are scaled down, also preserving aspect ratio, until the total pixel count lands near that of an 800×800 image. The result is a hard ceiling of 384 tokens per image. A 2000×2000 image and a 5000×5000 image therefore cost exactly the same after resizing. When a request carries multiple images, each one is counted independently under the same rule.
When fine visual detail does not matter, setting detail to low downscales the image to 512×512. Setting original or auto keeps the source resolution.
response = client.chat.completions.create(
model="deepseek-v4-flash-vision-exp",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image_url",
"image_url": {"url": "https://example.com/image.jpg", "detail": "low"}},
],
}],
)
Text Performance Holds Steady, Agent Work Is Where It Moves
According to DeepSeek, the experimental model matches DeepSeek-V4-Flash on the text side, covering agent behavior, reasoning, and world knowledge. Teams already running V4-Flash for text workloads should see little change in how the written output behaves after switching.
What did improve is multimodal agent performance. DeepSeek says the model makes a major leap over V4-Flash on multimodal agent benchmarks, landing close to Opus-4.8. Reading a dashboard screenshot, pulling figures out of a chart, or deciding the next action from a rendered screen may now take fewer steps than piping in a text summary instead.
The execution side moved on the same day: DeepSeek Harness 0.1.1 shipped with out-of-the-box support for the new model, so existing setups can try it without a significant rewrite.
Three Ways to Send an Image, and a Free Files API
There are three input paths. Embed the image directly in the request as base64, pass a public URL and let the model download it, or upload it once through the Files API and reference the returned file_id. Supported formats are JPEG, PNG, GIF, and WebP, and the format is detected from the file contents rather than from the file name or declared MIME type.
The Files API, which launched alongside the model, is free and saves request bandwidth when the same image is reused across calls. Upload once, then reference the file_id. Inline base64 counts against the 48 MiB request body limit, while an image referenced through the Files API can reach 64 MiB.
The limits are spelled out clearly: up to 600 images per request, external URLs of at most 8192 characters, and a maximum of 8192 px per side, which drops to 4096 px once a request contains 15 or more images. Images are accepted only in user messages; placing one in a system or assistant message returns a 400 error. The model is reachable through the OpenAI-compatible Chat Completions endpoint, the Anthropic-compatible Messages endpoint, and the Responses API.
Pricing Is Unchanged, and Weekends Go Fully Off-Peak
Rates are identical to DeepSeek-V4-Flash. Per 1M tokens, cache-hit input runs 0.007 USD (about 1 yen) off-peak and 0.014 USD (about 2 yen) at peak. Cache-miss input runs 0.22 USD (about 35 yen) off-peak and 0.44 USD (about 70 yen) at peak. Output runs 0.66 USD (about 105 yen) off-peak and 1.32 USD (about 210 yen) at peak. Sending images adds no separate line item; the converted tokens simply land on the input side.
※1 USD = 159 JPY (as of August 22, 2026)
Peak hours run 01:00–04:00 and 06:00–10:00 UTC, with everything else counted as off-peak. From 00:00 Beijing time on August 23, the split changes again so that Saturdays and Sundays are off-peak all day. For batch work that can be shifted in time, that single change halves the effective rate. Context length is 1M tokens, maximum output is 384K tokens, and the concurrency limit is 2500.
Summary
DeepSeek-V4-Flash-Vision-Exp adds image understanding while keeping text performance level with V4-Flash. The 384-token ceiling per image and the absence of any image surcharge make it far easier to forecast the cost of workloads built around screenshots and charts. A free Files API, three input paths, and the weekend off-peak change all arrived at the same time. It is still an experimental model, so measuring accuracy against your own images before production remains unavoidable, but the barrier to trying it is unusually low.
