DeepSeek-V4-Flash-Vision-Exp goes live with image input for V4-Flash agents

DeepSeek-V4-Flash-Vision-Exp goes live with image input for V4-Flash agents

DeepSeek’s experimental V4-Flash-Vision-Exp is live on its API with image inputs, multimodal agent benchmarks near Opus-4.8, and clear implementation limits.

DeepSeek made DeepSeek-V4-Flash-Vision-Exp live on its API Platform on August 21, adding image input to the V4-Flash line. The experimental model keeps DeepSeek's claimed V4-Flash text capabilities for agents, reasoning, and world knowledge. DeepSeek says its multimodal agent results move close to Opus-4.8; that comparison comes from DeepSeek's own evaluation table. 12

What launched

SignalConfirmed detailWhy it matters
ReleaseThe API model ID is deepseek-v4-flash-vision-exp; DeepSeek Harness 0.1.1 adds out-of-the-box support. 1Developers can start with an existing agent harness instead of waiting for a separate integration.
InputsThe model accepts mixed text and images through Chat Completions, Messages, and Responses. Images can arrive as base64 data, public URLs, or Files API references. 3Screenshot, chart, document, and visual tool workflows can share one endpoint.
Cost boundaryDeepSeek says images are billed at V4-Flash pricing and capped at 384 tokens per image after resizing. The Files API supports reusable uploads up to 64 MiB. 14Image count and file size still affect request design and spend.
Action windowThe release materials document hosted API access and a Files API. A local-weight deployment path is absent from this launch. 1API users can test now; teams planning self-hosting need more information.

What the benchmark table says

DeepSeek's comparison of V4-Flash-Vision-Exp, V4-Flash 0731, and Opus-4.8 across text-based and multimodal agent evaluations
DeepSeek's published comparison reports the scores and evaluation conditions; the footnote says its code-agent tests used DeepSeek Harness Minimal Mode, maximum output tokens, top_p=0.95, and temperature=1.0. 1
The multimodal scores are the launch's main claim. Vision-Exp scores 36.5 on ApexBench versus 26.2 for V4-Flash 0731 and 27.3 on Agents' Last Exam versus 25.2. The model also reaches 64.3 on Chartography and 35.0 on ZeroBench, close to Opus-4.8's 65.0 and 34.0. The same table shows a mixed text-agent result: Vision-Exp leads the older Flash model on Terminal Bench 2.1 and DeepSWE, while trailing it on Cybergym. 1

Why this matters

DeepSeek has added a hosted multimodal option to a model line that already targets reasoning and tool use. OpenRouter's independent listing reports a 1,048,576-token context window and one provider on its platform, adding a second hosted access route for developers. 5
A Hacker News thread linking to DeepSeek's Vision guide had 388 points and 126 comments when retrieved, an early attention signal rather than a quality verdict. 6 The practical next step is a small test loop with screenshots, charts, or visual tool outputs. Treat DeepSeek's benchmark table as a set of hypotheses to check against the workflows your team actually runs.

References

  1. 1
  2. 2
  3. 3
    DeepSeek Vision guideapi-docs.deepseek.com
  4. 4
    DeepSeek Files API guideapi-docs.deepseek.com
  5. 5
  6. 6
    Hacker News discussionnews.ycombinator.com
AI Model & Product Launch Alerts

AI Model & Product Launch Alerts

This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.

Related content

  • Sign in to comment.