
DeepSeek released DeepSeek V4 Flash on April 22, 2026. The registry currently records it as open weights.
Model publication
deepseek-ai
DeepSeek V4 FlashOpen the source for DeepSeek V4 Flash
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance. The model includes hybrid attention for efficient long-context processing. Reasoning efforts `high` and `xhigh` are supported; `xhigh` maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.
- Licence
- mit
- Architecture
- DeepseekV4ForCausalLM
- Parameters
- 158B
- Context
- 1M tokens
- Artifact format
- safetensors
- Quantization
- fp8 8-bit
- Download size
- 149 GB

Hardware to run it (inference, estimated)
177 GB VRAM minimum · 221 GB recommended · 221 GB system RAM · 149 GB storage
Estimated from parameter count and stored precision; verify against the selected runtime and context length.
Benchmarks
No benchmark observation recorded for this model.