In Focus
V4.1 Flash is the smallest model in DeepSeek’s new architecture family
The AI model is built on a 552 billion-parameter framework
DeepSeek’s V4.1 Flash delivers processing speeds that exceed 400 tokens per second
Chinese AI developer DeepSeek has released the smallest model in its latest architecture family, V4.1-Flash. According to the company, DeepSeek the V4.1 Flash outperforms the previous model while reducing inference and increasing speeds. DeepSeek introduced V4.1 Flash months after it released the V4 Flash AI model.
How DeepSeek Designed its Latest AI Model
DeepSeek’s V4.1 Flash model uses the new “Causal-Encorder-Decoder” architecture. It is built on a massive 552 billion-parameter framework and utilizes the Mixture-of-Experts (MoE) design.
Unlike conventional AI systems where each query runs through the entire system, MoE designs route tasks to subnetworks that are best suited for the job. For greater efficiency, DeepSeek V4.1 Flash uses only 8 billion parameters to process inputs and 16 billion to generate responses. This significantly reduces the computing power needed to process each request.
Despite being the smallest model in the company’s new architecture family, V4.1 Flash model understands text and images. It also features a massive one-million-token context window that enables it to efficiently handle heavy workflows.
According to DeepSeek, V4.1 Flash combines the MoE design with advanced sparse attention mechanisms to deliver high processing speeds that surpass 400 tokens per second. The highly optimized performance enables the model to handle real-time coding tasks, complex simulations, and data-heavy enterprise tasks using fewer active parameters.
How Did V4.1 Flash Perform on Benchmark Tests?
DeepSeek says V4.1 outperforms several AI models, including Moonshot AI’s Kimi K3, in a range of benchmark tests. Published tests showed the new model scored 90.9 on the scientific question answering benchmark, GPQA Diamond,. Kimi K3, Claude Opus 5, and GPT-5.6 Sol maintained higher scores on this benchmark at 92.9, 93.4, and 94.1 respectively.
However, V4.1 Flash performed better at end-to-end task execution. DeepSeek’s new model scored 90.6 on Terminal-Bench 2.1, compared to Kimi K3’s 88.3, Claude Opus 5’s 89.1, and GPT-5.6 Sol’s 88.8. DeepSeek’s V4.1 Flash performed better than Kimi K3 in Terminal Bench 3.0 and 4.0 tests. The new model also registered higher scores on the cybersecurity benchmark CyberGym at 88.1, ahead of Kimi K3’s 80.0 and GPT-5.6 Sol’s 84.5.
DeepSeek has also revised its API pricing, introducing lower rates for the Flash tier models. The company is also directing traffic from its older Pro models to the V4.1 Flash infrastructure temporarily. The new pricing comes as the company prepares for potential listing on Shanghai’s STAR Market.
Will V4.1 Flash Rollout Affect Other AI Models?
The launch of DeepSeek’s new model could push AI model developers into offering high performance at a lower cost. V4.1 Flash's ability to offer high throughput and low active-parameter usage will likely make efficiency an important factor for coding and enterprise workloads.
Additionally, DeepSeek’s lower API pricing could intensify competition and pressure rival AI developers to rethink the cost of high-volume inference. V4.1 Flash does not lead in every test benchmark. However, its performance suggests that smaller, highly optimized models can compete with large systems on specific tasks.


.webp&w=750&q=75)
.webp&w=750&q=75)
.webp&w=750&q=75)
