Chinese lab DeepSeek rolls out its V4.1-Flash model, pushing inference speeds past 400 tokens per second alongside advanced multimodal features
Chinese artificial intelligence lab DeepSeek has officially rolled out its V4.1-Flash model, marking a significant milestone in high-velocity inference technology.
As the smallest iteration within the company’s new architecture family, the model introduces native multimodal visual understanding and a massive one-million-token context window designed to handle heavy-duty workflows efficiently.
By implementing an asymmetric mixture-of-experts layout and advanced sparse attention mechanisms, the system achieves blistering processing throughput speeds exceeding 400 tokens per second.
This performance optimization heavily targets real-time coding, complex multi-modal simulations, and data-dense enterprise tasks while maintaining low active parameter consumption.
As reported by Reuters, the release came as the company is starting to prepare for an initial public offering on Shanghai's tech-focused STAR Market.
Alongside the technical rollout, DeepSeek has restructured its API tier pricing, offering aggressive flash-tier rates and temporarily routing incoming traffic from its older Pro models directly to the V4.1-Flash infrastructure.
DeepSeek said the new model is designed for greater capability, faster inference, higher throughput and scaling to larger models.
This aggressive market positioning coincides with the company's preparations for an upcoming public listing on Shanghai's STAR Market, intensifying competitive pressures across the global generative AI landscape.