This is a base model, trained only for raw next-token prediction.
DeepSeek released DeepSeek V3.1 Base on August 19, 2025. The registry currently records it as open weights.
Unlike instruct/chat models, it has not been fine-tuned to follow user instructions. Prompts need to be written more like training text or examples rather than simple requests (e.g., “Translate the following sentence…” instead of just “Translate this”).
DeepSeek-V3.1 Base is a 671B parameter open Mixture-of-Experts (MoE) language model with 37B active parameters per forward pass and a context length of 128K tokens. Trained on 14.8T tokens using FP8 mixed precision, it achieves high training efficiency and stability, with strong performance across language, reasoning, math, and coding tasks.
The structured profile below stays connected to the canonical model record, including its architecture, parameters, papers, benchmarks, media, licence, and source links as those fields are enriched.
Model publication
deepseek-ai
DeepSeek V3.1 BaseOpen the source for DeepSeek V3.1 Base
This is a base model, trained only for raw next-token prediction. Unlike instruct/chat models, it has not been fine-tuned to follow user instructions. Prompts need to be written more like training text or examples rather than simple requests (e.g., “Translate the following sentence…” instead of just “Translate this”). DeepSeek-V3.1 Base is a 671B parameter open Mixture-of-Experts (MoE) language model with 37B active parameters per forward pass and a context length of 128K tokens. Trained on 14.8T tokens using FP8 mixed precision, it achieves high training efficiency and stability, with strong performance across language, reasoning, math, and coding tasks.
- Licence
- mit
- Architecture
- DeepseekV3ForCausalLM
- Parameters
- 685B
- Context
- 163.8K tokens
- Artifact format
- safetensors
- Quantization
- fp8 8-bit
- Download size
- 641 GB

Hardware to run it (inference, estimated)
766 GB VRAM minimum · 957 GB recommended · 957 GB system RAM · 642 GB storage
Estimated from parameter count and stored precision; verify against the selected runtime and context length.
Benchmarks
No benchmark observation recorded for this model.