Thread · 8 stories · Apr 22 – Aug 24
V4 Pro and V4 Flash: open weights and a 1M context, then a local engine, a public API, the Arena cost frontier and an 11x hosting price spread.
Jump to timeline ↓DeepSeek shipped V4 as a pair — the 1.6T-parameter Pro and the efficiency-tuned 284B Flash — both with hybrid attention, open weights and a practical million-token context. Antirez had a local inference engine for Flash within weeks, and DeepSeek followed with DSpark speculative decoding and a public Flash API whose agent scores jumped.
By August, Flash owned the cost-performance frontier on Agent Arena — and exposed how little of a model's price is the model, with an 11x spread between the cheapest and priciest host of identical weights.