
DeepSeek has put the DeepSeek-V4-Flash API into public beta, alongside a new 0731 checkpoint whose agent scores land far above both the earlier Flash preview and the larger V4-Pro-Preview. The official endpoint now speaks the Responses API format natively and is configured for Codex.
The model itself is not new — V4 Flash shipped in April as an open-weight 284B mixture-of-experts with 13B active parameters and a 1M-token context. What changed is the agent tuning.
| Benchmark | V4-Flash-0731 | V4-Flash-Preview | V4-Pro-Preview | GLM-5.2 | Opus 4.8 |
|---|---|---|---|---|---|
| Terminal Bench 2.1 | 82.7 | 61.8 | 72.1 | 81.0 | 85.0 |
| NL2Repo | 54.2 | 39.4 | 38.5 | 48.9 | 69.7 |
| Cybergym | 76.7 | 38.7 | 52.7 | — | 83.1 |
| DeepSWE | 54.4 | 7.3 | 12.8 | 46.2 | 58.0 |
| Toolathlon-Verified | 70.3 | 49.7 | 55.9 | 59.9 | 76.2 |
| Agents' Last Exam | 25.2 | 15.8 | 16.5 | 23.8 | 25.7 |
| AutomationBench (Public) | 25.1 | 10.8 | 12.8 | 12.9 | 27.2 |
| DSBench-FullStack | 68.7 | 37.0 | 41.8 | 61.8 | 71.6 |
| DSBench-Hard | 59.6 | 25.8 | 31.1 | 54.5 | 71.7 |
DeepSeek's own figures. Public code-agent tasks were run through its forthcoming DeepSeek Harness in minimal mode (max tier, top-p 0.95, temperature 1.0); DSBench-FullStack and DSBench-Hard are internal benchmark sets.
Two things stand out. V4-Flash-0731 beats GLM-5.2 on every row the two share, which moves the open-weight lead. And it is within about two points of Claude Opus 4.8 on Terminal Bench, Agents' Last Exam and AutomationBench — while still trailing it by fifteen on NL2Repo and twelve on DSBench-Hard. The gap has not closed, but on agentic terminal work it is now narrow.
Announced by @deepseek_ai.