Skip to main content
AI Socratic
DeepSeek V4 Flash

DeepSeek released DeepSeek V4 Flash on April 22, 2026. The registry currently records it as open weights.

Model publication

deepseek-ai

DeepSeek V4 FlashOpen the source for DeepSeek V4 Flash

Open weights

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance. The model includes hybrid attention for efficient long-context processing. Reasoning efforts `high` and `xhigh` are supported; `xhigh` maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.

Licence
mit
Architecture
DeepseekV4ForCausalLM
Parameters
158B
Context
1M tokens
Artifact format
safetensors
Quantization
fp8 8-bit
Download size
149 GB
Figure 17: DeepSeek V4-Flash
DeepSeek V4-Flash (284B) · architecture · Sebastian Raschka · LLM Architecture Gallery

Hardware to run it (inference, estimated)

177 GB VRAM minimum · 221 GB recommended · 221 GB system RAM · 149 GB storage

Estimated from parameter count and stored precision; verify against the selected runtime and context length.

Benchmarks

No benchmark observation recorded for this model.

Weightshuggingface.co/deepseek-ai/DeepSeek-V4-FlashHugging Facehuggingface.coOpenRouteropenrouter.aiSebastian Raschka LLM Architecture Gallerysebastianraschka.com2 artifacts in the registry, including third-party rebuilds

About the Authors

A

AI Socratic