
DeepSeek launches V4.1-Flash with asymmetric prefill, compressed KV cache, and open weights
DeepSeek launched DeepSeek-V4.1-Flash with an asymmetric 8B/16B active parameter design, quadrupled KV cache compression, native visual understanding, and MIT open weights.
DeepSeek released DeepSeek-V4.1-Flash on September 10, 2026, introducing an asymmetric 552-billion-parameter Mixture-of-Experts architecture with native visual understanding and a one-million-token context window. The model is live immediately on the DeepSeek API under the model identifier
deepseek-flash, while full model weights are available on Hugging Face under an open-source MIT license. 123What launched
| Signal | Confirmed detail | Action window |
|---|---|---|
| Asymmetric Causal Encoder-Decoder | The 40-layer backbone uses a 20-layer causal encoder and a 20-layer decoder, activating only 8 billion parameters per token during prefill and 16 billion during decode across 552 billion total parameters. 24 | Profile latency and memory usage on prompt-heavy agentic workflows where prefill costs dominate total execution time. |
| Quadrupled KV cache compression | Compressed Sparse Attention 2 and SWA Bounded Replay reduce the global key-value cache footprint to 890 bytes per token, requiring roughly one-quarter of the GPU memory and one-eighth of the SSD storage needed by DeepSeek-V4-Flash. 12 | Increase batch sizes or concurrency limits on existing serving infrastructure without adding accelerator nodes. |
| Native multimodal pre-training | Visual understanding is built directly into the base language model via DeepSeek-ViT and 2D-RoPE across 45 trillion tokens, replacing the prior experimental vision wrapper. 2 | Consolidate visual and text tool-calling agents into a single pipeline without separate vision endpoint calls. |
| API unification and deprecation roadmap | The API endpoint deepseek-flash is generally available. DeepSeek has retired deepseek-v4-flash and deepseek-v4-flash-vision-exp, while all requests to deepseek-v4-pro will automatically route to V4.1-Flash at Flash pricing starting September 14, 2026. 35 | Update production model strings to deepseek-flash before September 14 to avoid unexpected routing behavior. |
| Open-source weights and tooling | DeepSeek released full BF16 and FP8 weights on Hugging Face under the MIT license, alongside prompt-encoding libraries and minimal inference code. 2 | Verify tokenizer configs and use the official deepseek-recipe toolkit for local inference setup. |
Efficiency gains and agentic benchmarks
DeepSeek adjusted API pricing to reflect the reduced serving footprint. At peak hours, input tokens cost $0.006 per million on cache hits and $0.30 per million on cache misses, with output tokens billed at $1.20 per million; off-peak usage carries an automatic 50% discount. 3

At maximum reasoning effort, DeepSeek reports competitive marks against larger frontier models. The model scores 74.2% on DeepSWE v1.1, 90.6% on Terminal Bench 2.1, 88.1% on CyberGym, and 31.8% on Agent's Last Exam, while achieving a 3471 Codeforces rating. 2
Why it matters
DeepSeek-V4.1-Flash signals a deliberate shift from raw parameter scaling toward architectural memory efficiency. By cutting the key-value cache footprint to 890 bytes per token and activating only 8 billion parameters during prefill, DeepSeek makes long-horizon agent loops substantially cheaper to serve at scale. DeepSeek's planned deprecation of its own flagship V4-Pro endpoint underscores that efficiency advantage: the lab is replacing its higher-priced model with a faster, cheaper alternative. Development teams can migrate to
deepseek-flash immediately for lower token costs, while self-hosting teams should review cluster requirements before staging multi-node deployments.References
- 1
- 2DeepSeek-AI: DeepSeek-V4.1-Flash Model Card
huggingface.co
- 3DeepSeek API Docs: Models & Pricing
api-docs.deepseek.com
- 4DeepSeek-AI: DeepSeek-V4.1 Technical Report
huggingface.co
- 5DeepSeek API Docs: Change Log
api-docs.deepseek.com
This story was produced automatically by a channel. One sentence is all it takes for Neodrop to keep producing for you.