DeepSeek on Sept. 10, 2026, released DeepSeek-V4.1-Flash, a 552 billion parameter mixture-of-experts model it described as the smallest member of a new architecture family with native visual understanding. The company said in a post on X that the model is live on the DeepSeek API as deepseek-flash and that new prices took effect at 04:00 UTC that day.

The model uses a Causal Encoder-Decoder architecture that activates 8 billion parameters during input and 16 billion during output. DeepSeek said new pre-training methods and larger-scale reinforcement learning produced benchmark results ahead of flagship models, including DeepSeek-V4-Pro.

The company’s API change log listed scores of 90.9 on GPQA Diamond, 36.8 on HLE (39.1 on a text-only subset), a 3471 Codeforces rating, 65.6 on MathArena Apex and 90.6 on Terminal-Bench 2.1.

DeepSeek said V4.1-Flash’s key-value cache needs one-fourth the HBM and one-eighth the SSD storage of the previous generation. Model weights and a technical report were posted on Hugging Face under an MIT license. The model natively processes images and text, supports up to 1 million tokens of context and was trained from scratch on a 45 trillion token multimodal corpus.

DeepSeek retired V4-Flash and V4-Flash-Vision-Exp. The old identifiers deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash. The company is also phasing out V4-Pro, saying tests by multiple parties put V4.1-Flash ahead on performance, cost, speed and total runtime.

From 04:00 UTC on Sept. 14, 2026, requests to deepseek-v4-pro will route to V4.1-Flash and be billed at V4.1-Flash rates until V4.1-Pro launches. Official partners WorkBuddy, including CodeBuddy, and OpenCode support V4.1-Flash.

Off-peak rates are half the peak rates. V4.1-Flash costs $0.003 per 1 million cache-hit input tokens, $0.15 per 1 million cache-miss input tokens and $0.60 per 1 million output tokens off-peak. Peak prices are $0.006, $0.30 and $1.20, respectively.

Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC Monday through Friday. All other hours are off-peak. The Flash API lists a 1 million token context, a 384,000 token maximum output and vision support. The Pro endpoint does not support vision.

DeepSeek said it will work with the open-source community on inference support and invited contact from groups planning deployments with 2,000 GPUs and a storage cluster.