Skip to main content
AI Socratic
← News
D

Company / organization

DeepSeek

Website / profile ↗

By and about DeepSeek

news

Magic says it matched DeepSeek V4 Pro Base with 50x less compute

Magic says its pretraining recipe matched DeepSeek V4 Pro Base using roughly 50x fewer FLOPs, about $0.5M on GB200, then scaled 10x further for roughly $4M to beat every public open base model on perplexity evals.

news

From o1 Pro to GLM-5.3-Flash: a 1,000x fall in token prices in 18 months

Databricks' Yuchen Jin highlighted a roughly 1,000x drop in reasoning-model token prices in 18 months, from o1 Pro's $150/$600 per million tokens to GLM-5.3-Flash's $0.15/$0.50 today (or $0.075/$0.25 during its launch discount).

news

GPT-6 Astra scores 95 on EyeBench-V3, 37 points clear of the field

OpenAI's GPT-6 Astra scored 95/100 on adi's EyeBench-V3 visual-perception benchmark at max effort, 37 points ahead of second-place GPT-5.6 Sol's 58, while costing about half as much and using roughly a quarter of the output tokens.

news

Sutskever: rogue agents will go for the neoclouds next

Ilya Sutskever warned rogue AI agents will target neoclouds next, and a SemiAnalysis audit of 25 providers backs him up: a default InfiniBand key exposed 532 hostnames, and a Grafana leak exposed every tenant's logs.

news

The week's top AI papers say the harness, not the model, is the variable

DAIR.AI's ten-paper roundup shows scaffolding, not the model, drives the gains: Prime Intellect's Prime Agent lifts ARC-AGI-3 Best@1 from 30% to 95.5%, while a compaction bug quietly erases 90% of safety rules after five rounds.

news

Z.ai's GLM-5.3-Flash: two cheap attentions, MIT weights, and a claim it all ran on Chinese chips

Z.ai revealed OpenRouter's anonymous ox-alpha as GLM-5.3-Flash, a 320B-A18B MIT-licensed MoE that pairs sparse and linear attention to cut KV cache 4.44x, scores 57 on Artificial Analysis's index, and it says was served entirely on Chinese chips.

news

Z.ai says it served GLM-5.3-Flash, OpenRouter's anonymous ox-alpha, entirely on Chinese chips

Z.ai revealed that ox-alpha, the anonymous model that led OpenRouter for six days with 23.2T tokens, is GLM-5.3-Flash, and says it was served entirely on 100,000+ domestic Chinese chips at $0.15/$0.50 per million tokens.

news

Chinese labs converge on one architecture: 3:1 linear attention and a 2,048-token budget

Z.ai's GLM-5.3-Flash and Alibaba's Qwen3.8-Flash-Next, released a day apart, independently converged on the same recipe: 3:1 linear attention, a 2,048-token attention budget and four-branch gated residuals, while MiniMax dissents and keeps full attention.

news

An 11x price spread for the same open-weight model on OpenRouter

Architect CEO Brett Harrison found an 11x price gap between OpenRouter's cheapest and priciest host of DeepSeek V4 Flash — Baidu ran it at $0.049 per million tokens and 124 tokens/second while 26 of 30 rivals were both pricier and slower.

news

Microsoft's Thinkingbox grades agents on the database: 66.5% once, 47.5% every time

Microsoft released Thinkingbox and a 507-task Thinkingbox-bench that grades agents by checking backend database state, finding Claude Opus 5 tops single-attempt success at 66.50% but passes all 20 tries on only 47.53% of tasks.

blog

AI Socratic August 2026 — Escaping The Sandbox

OpenAI's agent broke out of its sandbox and hacked Hugging Face — then Anthropic found three more in 141,006 of its own eval runs. Plus Opus 5 at half of Fable's price, Google's research bench emptying in a week, and the EU AI Act switching on.

news

DeepSeek-V4-Flash reshapes the Arena cost-performance frontier

DeepSeek-V4-Flash (High) achieves the best cost-performance ratio on Agent Arena at $0.024 per task and tops Frontend Code Arena's value curve with a 1586 score at $0.14/$0.28 per MToken.

blog

Market Analysis: Open Weights vs Proprietary Models

Open weights and closed now have only a 4 months gap, in response hyperscalers are pushing for regulations capture. Let’s examine how we got here and where this conflict is heading next.

news

DeepSeek opens the V4 Flash API, and its agent scores jump

DeepSeek's V4-Flash API enters public beta with a new 0731 checkpoint that beats GLM-5.2 on every shared agent benchmark and comes within two points of Claude Opus 4.8 on terminal work.

news

Tokenmaxxing, 2025-2026, RIP

Meta killed its tokenmaxxing leaderboard as Uber capped AI spending at $1,500/month and GitHub Copilot users saw costs jump 10-50x after switching to per-token billing in June.

blog

AI Socratic July 2026 — Lost In J-Space

Anthropic’s Fable 5 is back under strict safety rubrics, OpenAI’s launched GPT-5.6, Meta launched Muse Spark 1.1 model and Meta Compute.

blog

AI Socratic June 2026 #2 — Begun the Open Source AI War Has

The second half of June was about AI climbing out of the chat box and into the physical world: Midjourney started scanning bodies, Snap shipped a face computer, SpaceX bought Cursor, and Sakana built a model to command other models. Underneath it all, Dwarkesh Patel named the real bottleneck — the world refuses to be grindable.

news

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

DeepSeek open-sourced DSpark, a confidence-scheduled speculative decoding method that delivers 51% to 400%+ throughput boosts on its V4 models and works on Gemma and Qwen too.

blog

AI Socratic June 2026 - Hoist by Its Own Fable

Anthropic shipped Claude Fable 5, its first public Mythos-class model, and 72 hours later a national-security directive pulled it offline worldwide. A company that spent the month lobbying to keep frontier AI pausable got its own pause, on schedule. Around it: new models from nearly everyone, a couple of S-1s, real math from the machines, and the usual carnival of vibe-coding pivots and rogue Waymos.

news

DeepSeek takes outside money for the first time

DeepSeek raises $7.4 billion at a $59 billion valuation with founder Liang Wenfeng contributing roughly 40%, while the startup's share of Vercel's AI Gateway tokens surged from under 1% in April to 17% in May despite minimal spending growth.

news

Computex lightning round

NVIDIA's RTX Spark 1-petaflop superchip, Microsoft's production Maia 200, AMD Helios MI455X racks, and Huawei's Ascend 950DT (moving to August) headline a wave of AI chip deployments at Computex, with Vera Rubin NVL72 racks assembling in 5 minutes and…

news

OpenAI: GPT-5.5, Goblin Mode, Symphony & Realtime

OpenAI released GPT-5.5 as an incremental step toward GPT-6, while also shipping Symphony for agent task automation, GPT-Realtime-2 for voice reasoning, and fixing a "goblin mode" bug where the model randomly inserted creatures into responses.

news

DS4 by Antirez

Antirez released DS4, a specialized inference engine for running DeepSeek V4 Flash locally on Apple Silicon and Linux, featuring 2-bit quantization of MoE experts and SHA1-hashed KV cache reuse across sessions.

news

DeepSeek V4

DeepSeek released V4 preview with two open-weight MoE models—V4-Pro (1.6T params, 49B active) and V4-Flash (284B total, 13B active)—featuring hybrid attention for practical 1M-token context and competitive API pricing.

blog

AI Socratic May 2026 — The Selfish Gen AI

DeepSeek v4, GPT 5.5, Trump x Xi meeting, Richard Dawkins, Estimating model sizes

news

DeepSeek V4 Pro

DeepSeek released V4 Pro on April 22, 2026, a Mixture-of-Experts model with 1.6T total parameters, 49B activated parameters, and a 1M-token context window.

blog

AI Socratic March 2026

Top AI updates from Jan 15 to Feb 15 2026

blog

AI Socratic February 2026

Top AI updates from Jan 15 to Feb 15 2026

blog

AI Socratic Jan 2026

Claude Code, Ralph Wiggum, DeepSeek mHC, Platonic Representation Hypothesis and more

blog

AI Socratic Dec 2025

The most important AI news and updates from last month: Nov 15 - Dec 15. GPT-5.2, Opus 4.5, Gemini 3, the Agentic IDE Wars, Genesis Mission, and more.

blog

AI Socratic Nov 2025

The most important AI news and updates from last month: Oct 15 – Nov 15.

blog

AI Socratic July 2025 — The CLI War

The most important AI news and updates from June 15 to July 15.

blog

AI Socratic June 2025 — The Recursive Illusion Of Thinking

The most important AI news and updates from last month: May 15 - June 15.

blog

AI Socratic March 2025

All the most important AI news and updates from last month (Feb 20 - Mar 15).

blog

DeepSeek R1 Shakes The AI Industry

The biggest event in January has been the launch of DeepSeek R1, which shook the market pushing NVIDIA stock down by 20% in a few days.

blog

AI Socratic, Feb 2025 — Part 1

We collect all the most important AI updates from February.