DeepSeek is preparing to launch its V4.1 Flash model around September 10, 2026, according to a notice released by the company’s official open platform. The new model is expected to bring improvements across performance, response speed, operating costs and overall task completion time.
DeepSeek says that V4.1 Flash has gone through testing involving both internal teams and external evaluators. Based on those tests, the company claims the new Flash model has outperformed DeepSeek V4 Pro across multiple key metrics, making it a significant upgrade within the company’s model lineup.
V4 Pro Requests Will Move to V4.1 Flash
DeepSeek has also outlined how it plans to handle existing V4 Pro requests after the new model becomes available.
Once V4.1 Flash officially launches, but before the company releases V4.1 Pro, requests that would previously have been processed by V4 Pro will be automatically routed to V4.1 Flash. This pricing structure could make V4.1 Flash particularly attractive for developers and businesses running large volumes of AI workloads.
The move suggests that DeepSeek sees V4.1 Flash not simply as a faster alternative, but as a model capable of taking over a broader range of workloads previously handled by its Pro model.
DeepSeek Announces New Flash Pricing
Alongside the upcoming model launch, DeepSeek is also changing the pricing for its Flash series.
The new pricing is scheduled to take effect at 12:00 PM Beijing time on September 10, 2026. During off-peak periods, cached input will cost 0.02 yuan per million tokens, while uncached input will be priced at 1 yuan per million tokens. Output will cost 4 yuan per million tokens.
During peak hours, each of these rates will be twice the off-peak price.
This pricing structure could make V4.1 Flash particularly attractive for developers and businesses running large volumes of AI workloads, especially when response speed and API costs are both important considerations.
New Model Architecture and Native Multimodal Support
DeepSeek had already begun testing an intermediate version of V4.1 Flash before the official release.
According to a notice shared in one of DeepSeek’s official discussion groups, the test version uses a new model architecture and introduces native multimodal capabilities. The company also indicated that the model delivers stronger overall capabilities while responding faster and requiring lower operating costs.
Native multimodal support could be one of the more important changes in the V4.1 generation, potentially allowing the model to handle different types of information more naturally rather than focusing primarily on text-based tasks.
What to Expect From DeepSeek V4.1 Flash
The upcoming V4.1 Flash launch appears to be focused on improving the balance between AI performance, speed and cost.
Rather than relying solely on higher model capability, DeepSeek is positioning the new Flash model as a more efficient option for real-world workloads. Its planned replacement of V4 Pro requests also indicates that the company is confident in the model’s ability to handle demanding applications.
DeepSeek has not yet provided a complete breakdown of every benchmark or technical specification. More details should become available once V4.1 Flash officially goes live.
For now, the combination of a new architecture, native multimodal support, faster processing and lower pricing makes DeepSeek V4.1 Flash one of the company’s more notable model updates to watch in September 2026.









