DeepSeek has released its DeepSeek V4 Preview models, introducing two variants, DeepSeek-V4-Pro and DeepSeek-V4-Flash, alongside open-source weights and updated API access.
Both models support a 1 million-token context window and are designed for improved efficiency in long-context and agent-based workloads, the Chinese AI firm said.
DeepSeek-V4-Pro has 1.6 trillion total parameters with 49 billion active parameters, while V4-Flash has 284 billion total and 13 billion active parameters.
The Pro model is positioned for high-performance reasoning and coding tasks, while Flash is aimed at faster and lower-cost usage with comparable performance on simpler workloads.
The company said the release focuses on long-context efficiency, using compressed sparse attention and hybrid attention techniques to reduce compute and memory requirements.
It also highlighted agent-oriented optimizations and compatibility with external AI coding tools. API access is available immediately, with support for multiple interfaces and dual “thinking” and “non-thinking” modes.
Analysts at Jefferies see the launch as part of a broader wave of rapid AI model releases, noting more than 10 new model announcements in April alone across the industry.
They highlighted DeepSeek’s reported improvements in agentic capabilities, reasoning, and long-context performance across coding, tool use, and knowledge benchmarks, as well as integrations with agent frameworks such as coding tools and development platforms.
Jefferies also highlighted the company’s focus on long-context efficiency, pointing to compressed sparse attention and related architectural techniques aimed at reducing computational cost.
On pricing, the analysts described DeepSeek’s API as highly competitive versus both international and domestic peers, particularly for long-context workloads.