๐ Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient....

๐น 552B-parameter MoE.
๐น New Causal EncoderโDecoder architecture: just 8B active parameters for input, 16B for output.
๐น New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.
2/6
Set your model to deepseek-flash.
๐น V4-Flash & V4-Flash-Vision-Exp are retired. For compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash.
๐น Tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime. Weโre phasing out V4-Pro.
๐น Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches.
๐ค Official partners @WorkBuddy_AI (including Codebuddy) & @opencode now fully support V4.1-Flash. Try it today!
4/6
V4.1-Flash lets us serve more users at a lower cost. Weโre passing the savings on to you.
๐น Peak/off-peak pricing continues to balance demand.
๐น Off-peak rates are 50% of peak rates. Schedule flexible workloads off-peak to save.
๐น New pricing takes effect at 04:00 UTC on Sept 10, 2026.
5/6
Weโll work closely with the open-source community on V4.1-Flash inference support and explore more deployment options.
Planning a large-scale deployment with 2,000 GPUs + a storage cluster? Letโs talk.
๐น Model: huggingface.co/deepseek-ai/Deโฆ
๐น Paper: huggingface.co/deepseek-ai/Deโฆ
6/6



