๐Ÿš€ Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient....

@deepseek_ai
DeepSeek@deepseek_ai
13 views Sep 10, 2026 ~2 min read
Advertisement
1
๐Ÿš€ Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.

๐Ÿ”น Introducing the smallest model in our new architecture family, with native visual understanding.
๐Ÿ”น Designed for greater capability, faster inference, higher throughput, and scaling to larger models.

1/6
Media image
2
๐Ÿง  Asymmetric architecture. More intelligence, less cost.

๐Ÿ”น 552B-parameter MoE.
๐Ÿ”น New Causal Encoderโ€“Decoder architecture: just 8B active parameters for input, 16B for output.
๐Ÿ”น New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.

2/6
Media image
3
๐Ÿ’พ Smaller KV cache. Bigger savings.

Compared with the previous generation, V4.1-Flashโ€™s KV cache needs just:
๐Ÿ”น 1/4 the HBM
๐Ÿ”น 1/8 the SSD storage
Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly.

3/6
Media image
4
โšก V4.1-Flash is now live on the DeepSeek API with native multimodal support.

Set your model to deepseek-flash.

๐Ÿ”น V4-Flash & V4-Flash-Vision-Exp are retired. For compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash.
๐Ÿ”น Tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime. Weโ€™re phasing out V4-Pro.
๐Ÿ”น Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches.

๐Ÿค Official partners @WorkBuddy_AI (including Codebuddy) & @opencode now fully support V4.1-Flash. Try it today!

4/6
5
๐Ÿ’ฐ More efficient architecture. Lower API prices.

V4.1-Flash lets us serve more users at a lower cost. Weโ€™re passing the savings on to you.

๐Ÿ”น Peak/off-peak pricing continues to balance demand.
๐Ÿ”น Off-peak rates are 50% of peak rates. Schedule flexible workloads off-peak to save.
๐Ÿ”น New pricing takes effect at 04:00 UTC on Sept 10, 2026.

5/6
Media image
6
๐ŸŒ Supporting open source. Expanding deployment options.

Weโ€™ll work closely with the open-source community on V4.1-Flash inference support and explore more deployment options.
Planning a large-scale deployment with 2,000 GPUs + a storage cluster? Letโ€™s talk.

๐Ÿ”น Model: huggingface.co/deepseek-ai/Deโ€ฆ
๐Ÿ”น Paper: huggingface.co/deepseek-ai/Deโ€ฆ

6/6
Actions
What You Can Do
  • Export as PDF or Markdown
  • Batch Export to Notion
  • Bookmark & Highlight
  • LinkedIn & Instagram Carousel Maker
Create Free Account

Includes 7-day Premium trial

Advertisement