DeepSeek launches V4.1-Flash and begins retiring V4-Pro
The 552B-parameter model uses a new encoder-decoder design with 8B active input parameters, and will absorb all V4-Pro traffic from September 14.

DeepSeek has released DeepSeek-V4.1-Flash, which it describes as the smallest model in a new architecture family with native visual understanding. The model is a 552B-parameter mixture-of-experts system built on what the company calls a Causal Encoder-Decoder architecture, with 8B active parameters for input and 16B for output. DeepSeek says new pretraining methods and larger-scale reinforcement learning post-training deliver benchmark results ahead of its flagship models, including DeepSeek-V4-Pro. The company emphasizes cost. Compared with the previous generation, the V4.1-Flash KV cache needs a quarter of the HBM and an eighth of the SSD storage. DeepSeek notes that cache-hit charges often make up a large share of agent costs, so compressing the cache cuts those costs significantly, and it says the more efficient architecture allows lower API prices. V4.1-Flash is live on the DeepSeek API under the model name deepseek-flash with native multimodal support. V4-Flash and V4-Flash-Vision-Exp are retired, with their old model names temporarily routing to the new model. DeepSeek is also phasing out V4-Pro. Starting at 04:00 UTC on September 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates until a V4.1-Pro launches. Official partners WorkBuddy, including CodeBuddy, and OpenCode support the new model. Developers relying on V4-Pro have a short window to test before the automatic switch.