DeepSeek has officially launched V4.1 Flash, a new model positioned as the smallest member of its next generation architecture family. The release introduces native visual understanding, higher throughput, and improved efficiency for API users. It also marks the start of a transition away from the V4 Pro model and a new Flash series pricing structure.
Technical Architecture
V4.1 Flash is built on an asymmetric Causal Encoder Decoder architecture within a 552 billion parameter Mixture of Experts framework. The model activates only 8 billion parameters for input and 16 billion parameters for output generation. New pretraining methods and larger scale reinforcement learning post-training help the model deliver benchmark results ahead of flagship models including DeepSeek V4 Pro.
Memory and Cache Efficiency
The new model also reduces memory and storage demands compared with the previous generation. Its key value cache needs one quarter of the HBM capacity and one eighth of the SSD storage. DeepSeek says this compression lowers costs for agent workloads, where cache hit charges often account for a significant share of total spending.
Performance and API Availability
DeepSeek reports that V4.1 Flash has surpassed V4 Pro in performance, cost, speed, and total completion time during internal and external testing. The model is now live on the DeepSeek API with native multimodal support, and users can set the model name to deepseek-flash. It is also available on Hugging Face alongside the official technical report.
Compatibility and Partner Support
Several official partners and developer tools already support the new model. WorkBuddy, CodeBuddy, and OpenCode have fully integrated V4.1 Flash for users. DeepSeek is also retiring the older V4 Flash and V4 Flash Vision Exp identifiers, while those legacy model names will temporarily route to V4.1 Flash for compatibility, and the previous deepseek-v4-flash and deepseek-v4-flash-vision-exp endpoints continue to function.
New Flash Series Pricing
New Flash series pricing took effect on Sept. 10, 2026 with the model release. For off peak usage, cache hit input is priced at RMB 0.02 per million tokens, cache miss input at RMB 1, and output at RMB 4. Peak hour pricing is set at twice the off peak level, while off peak rates are 50 percent of peak rates to help balance demand.
Workload Scheduling and Cost Efficiency
The pricing structure continues to support workload scheduling for cost conscious users. DeepSeek encourages flexible workloads to be run during off peak periods to capture lower rates. The company states that V4.1 Flash allows it to serve more users at a lower cost and pass the savings on to customers, reflecting the efficiency gains from the new architecture.
Model Migration Timeline
Beginning at 04:00 UTC on Sept. 14, 2026, all deepseek-v4-pro requests will be routed to V4.1 Flash and billed at V4.1 Flash rates. This migration will remain in place until the launch of V4.1 Pro. DeepSeek is phasing out V4 Pro as part of its broader transition to the new model family.
Open Source and Enterprise Deployment
DeepSeek plans to work closely with the open source community to support V4.1 Flash inference and explore additional deployment options. The company is also inviting organizations planning large scale deployments with at least 2,000 GPUs and a storage cluster to discuss tailored infrastructure solutions. This move signals an effort to expand adoption beyond the standard API channel.
DeepSeek V4.1 Flash represents a strategic shift toward more efficient models with lower operating costs and stronger multimodal capabilities. The combination of reduced cache requirements, simplified pricing, and broad partner support positions the release as a practical option for developers and enterprises. With V4 Pro being phased out and further open source work planned, DeepSeek is clearly aligning its product line around the new Flash architecture.