
DeepSeek-V4-Flash-Vision-Exp goes live with image input for V4-Flash agents
DeepSeek’s experimental V4-Flash-Vision-Exp is live on its API with image inputs, multimodal agent benchmarks near Opus-4.8, and clear implementation limits.
DeepSeek made DeepSeek-V4-Flash-Vision-Exp live on its API Platform on August 21, adding image input to the V4-Flash line. The experimental model keeps DeepSeek's claimed V4-Flash text capabilities for agents, reasoning, and world knowledge. DeepSeek says its multimodal agent results move close to Opus-4.8; that comparison comes from DeepSeek's own evaluation table. 12
What launched
| Signal | Confirmed detail | Why it matters |
|---|---|---|
| Release | The API model ID is deepseek-v4-flash-vision-exp; DeepSeek Harness 0.1.1 adds out-of-the-box support. 1 | Developers can start with an existing agent harness instead of waiting for a separate integration. |
| Inputs | The model accepts mixed text and images through Chat Completions, Messages, and Responses. Images can arrive as base64 data, public URLs, or Files API references. 3 | Screenshot, chart, document, and visual tool workflows can share one endpoint. |
| Cost boundary | DeepSeek says images are billed at V4-Flash pricing and capped at 384 tokens per image after resizing. The Files API supports reusable uploads up to 64 MiB. 14 | Image count and file size still affect request design and spend. |
| Action window | The release materials document hosted API access and a Files API. A local-weight deployment path is absent from this launch. 1 | API users can test now; teams planning self-hosting need more information. |
What the benchmark table says

top_p=0.95, and temperature=1.0. 1The multimodal scores are the launch's main claim. Vision-Exp scores 36.5 on ApexBench versus 26.2 for V4-Flash 0731 and 27.3 on Agents' Last Exam versus 25.2. The model also reaches 64.3 on Chartography and 35.0 on ZeroBench, close to Opus-4.8's 65.0 and 34.0. The same table shows a mixed text-agent result: Vision-Exp leads the older Flash model on Terminal Bench 2.1 and DeepSWE, while trailing it on Cybergym. 1
Why this matters
DeepSeek has added a hosted multimodal option to a model line that already targets reasoning and tool use. OpenRouter's independent listing reports a 1,048,576-token context window and one provider on its platform, adding a second hosted access route for developers. 5
A Hacker News thread linking to DeepSeek's Vision guide had 388 points and 126 comments when retrieved, an early attention signal rather than a quality verdict. 6 The practical next step is a small test loop with screenshots, charts, or visual tool outputs. Treat DeepSeek's benchmark table as a set of hypotheses to check against the workflows your team actually runs.
References
- 1DeepSeek-V4-Flash-Vision-Exp release
api-docs.deepseek.com
- 2
- 3DeepSeek Vision guide
api-docs.deepseek.com
- 4DeepSeek Files API guide
api-docs.deepseek.com
- 5OpenRouter model listing
openrouter.ai
- 6Hacker News discussion
news.ycombinator.com
AI Model & Product Launch Alerts
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.
Related content
- Sign in to comment.