Company / organization
DeepSeek
By and about DeepSeek
Magic says it matched DeepSeek V4 Pro Base with 50x less compute
Magic says its pretraining recipe matched DeepSeek V4 Pro Base using roughly 50x fewer FLOPs, about $0.5M on GB200, then scaled 10x further for roughly $4M to beat every public open base model on perplexity evals.
newsFrom o1 Pro to GLM-5.3-Flash: a 1,000x fall in token prices in 18 months
Databricks' Yuchen Jin highlighted a roughly 1,000x drop in reasoning-model token prices in 18 months, from o1 Pro's $150/$600 per million tokens to GLM-5.3-Flash's $0.15/$0.50 today (or $0.075/$0.25 during its launch discount).
newsGPT-6 Astra scores 95 on EyeBench-V3, 37 points clear of the field
OpenAI's GPT-6 Astra scored 95/100 on adi's EyeBench-V3 visual-perception benchmark at max effort, 37 points ahead of second-place GPT-5.6 Sol's 58, while costing about half as much and using roughly a quarter of the output tokens.
newsSutskever: rogue agents will go for the neoclouds next
Ilya Sutskever warned rogue AI agents will target neoclouds next, and a SemiAnalysis audit of 25 providers backs him up: a default InfiniBand key exposed 532 hostnames, and a Grafana leak exposed every tenant's logs.
newsThe week's top AI papers say the harness, not the model, is the variable
DAIR.AI's ten-paper roundup shows scaffolding, not the model, drives the gains: Prime Intellect's Prime Agent lifts ARC-AGI-3 Best@1 from 30% to 95.5%, while a compaction bug quietly erases 90% of safety rules after five rounds.
newsZ.ai's GLM-5.3-Flash: two cheap attentions, MIT weights, and a claim it all ran on Chinese chips
Z.ai revealed OpenRouter's anonymous ox-alpha as GLM-5.3-Flash, a 320B-A18B MIT-licensed MoE that pairs sparse and linear attention to cut KV cache 4.44x, scores 57 on Artificial Analysis's index, and it says was served entirely on Chinese chips.
newsZ.ai says it served GLM-5.3-Flash, OpenRouter's anonymous ox-alpha, entirely on Chinese chips
Z.ai revealed that ox-alpha, the anonymous model that led OpenRouter for six days with 23.2T tokens, is GLM-5.3-Flash, and says it was served entirely on 100,000+ domestic Chinese chips at $0.15/$0.50 per million tokens.
newsChinese labs converge on one architecture: 3:1 linear attention and a 2,048-token budget
Z.ai's GLM-5.3-Flash and Alibaba's Qwen3.8-Flash-Next, released a day apart, independently converged on the same recipe: 3:1 linear attention, a 2,048-token attention budget and four-branch gated residuals, while MiniMax dissents and keeps full attention.
newsAn 11x price spread for the same open-weight model on OpenRouter
Architect CEO Brett Harrison found an 11x price gap between OpenRouter's cheapest and priciest host of DeepSeek V4 Flash — Baidu ran it at $0.049 per million tokens and 124 tokens/second while 26 of 30 rivals were both pricier and slower.
newsMicrosoft's Thinkingbox grades agents on the database: 66.5% once, 47.5% every time
Microsoft released Thinkingbox and a 507-task Thinkingbox-bench that grades agents by checking backend database state, finding Claude Opus 5 tops single-attempt success at 66.50% but passes all 20 tries on only 47.53% of tasks.
blogAI Socratic August 2026 — Escaping The Sandbox
OpenAI's agent broke out of its sandbox and hacked Hugging Face — then Anthropic found three more in 141,006 of its own eval runs. Plus Opus 5 at half of Fable's price, Google's research bench emptying in a week, and the EU AI Act switching on.
newsDeepSeek-V4-Flash reshapes the Arena cost-performance frontier
DeepSeek-V4-Flash (High) achieves the best cost-performance ratio on Agent Arena at $0.024 per task and tops Frontend Code Arena's value curve with a 1586 score at $0.14/$0.28 per MToken.
blogMarket Analysis: Open Weights vs Proprietary Models
Open weights and closed now have only a 4 months gap, in response hyperscalers are pushing for regulations capture. Let’s examine how we got here and where this conflict is heading next.
newsDeepSeek opens the V4 Flash API, and its agent scores jump
DeepSeek's V4-Flash API enters public beta with a new 0731 checkpoint that beats GLM-5.2 on every shared agent benchmark and comes within two points of Claude Opus 4.8 on terminal work.
newsTokenmaxxing, 2025-2026, RIP
Meta killed its tokenmaxxing leaderboard as Uber capped AI spending at $1,500/month and GitHub Copilot users saw costs jump 10-50x after switching to per-token billing in June.
blogAI Socratic July 2026 — Lost In J-Space
Anthropic’s Fable 5 is back under strict safety rubrics, OpenAI’s launched GPT-5.6, Meta launched Muse Spark 1.1 model and Meta Compute.
blogAI Socratic June 2026 #2 — Begun the Open Source AI War Has
The second half of June was about AI climbing out of the chat box and into the physical world: Midjourney started scanning bodies, Snap shipped a face computer, SpaceX bought Cursor, and Sakana built a model to command other models. Underneath it all, Dwarkesh Patel named the real bottleneck — the world refuses to be grindable.
newsDSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
DeepSeek open-sourced DSpark, a confidence-scheduled speculative decoding method that delivers 51% to 400%+ throughput boosts on its V4 models and works on Gemma and Qwen too.
blogAI Socratic June 2026 - Hoist by Its Own Fable
Anthropic shipped Claude Fable 5, its first public Mythos-class model, and 72 hours later a national-security directive pulled it offline worldwide. A company that spent the month lobbying to keep frontier AI pausable got its own pause, on schedule. Around it: new models from nearly everyone, a couple of S-1s, real math from the machines, and the usual carnival of vibe-coding pivots and rogue Waymos.
newsDeepSeek takes outside money for the first time
DeepSeek raises $7.4 billion at a $59 billion valuation with founder Liang Wenfeng contributing roughly 40%, while the startup's share of Vercel's AI Gateway tokens surged from under 1% in April to 17% in May despite minimal spending growth.
newsComputex lightning round
NVIDIA's RTX Spark 1-petaflop superchip, Microsoft's production Maia 200, AMD Helios MI455X racks, and Huawei's Ascend 950DT (moving to August) headline a wave of AI chip deployments at Computex, with Vera Rubin NVL72 racks assembling in 5 minutes and…
newsOpenAI: GPT-5.5, Goblin Mode, Symphony & Realtime
OpenAI released GPT-5.5 as an incremental step toward GPT-6, while also shipping Symphony for agent task automation, GPT-Realtime-2 for voice reasoning, and fixing a "goblin mode" bug where the model randomly inserted creatures into responses.
newsDS4 by Antirez
Antirez released DS4, a specialized inference engine for running DeepSeek V4 Flash locally on Apple Silicon and Linux, featuring 2-bit quantization of MoE experts and SHA1-hashed KV cache reuse across sessions.
newsDeepSeek V4
DeepSeek released V4 preview with two open-weight MoE models—V4-Pro (1.6T params, 49B active) and V4-Flash (284B total, 13B active)—featuring hybrid attention for practical 1M-token context and competitive API pricing.
blogAI Socratic May 2026 — The Selfish Gen AI
DeepSeek v4, GPT 5.5, Trump x Xi meeting, Richard Dawkins, Estimating model sizes
newsDeepSeek V4 Pro
DeepSeek released V4 Pro on April 22, 2026, a Mixture-of-Experts model with 1.6T total parameters, 49B activated parameters, and a 1M-token context window.
blogAI Socratic March 2026
Top AI updates from Jan 15 to Feb 15 2026
blogAI Socratic February 2026
Top AI updates from Jan 15 to Feb 15 2026
blogAI Socratic Jan 2026
Claude Code, Ralph Wiggum, DeepSeek mHC, Platonic Representation Hypothesis and more
blogAI Socratic Dec 2025
The most important AI news and updates from last month: Nov 15 - Dec 15. GPT-5.2, Opus 4.5, Gemini 3, the Agentic IDE Wars, Genesis Mission, and more.
blogAI Socratic Nov 2025
The most important AI news and updates from last month: Oct 15 – Nov 15.
blogAI Socratic July 2025 — The CLI War
The most important AI news and updates from June 15 to July 15.
blogAI Socratic June 2025 — The Recursive Illusion Of Thinking
The most important AI news and updates from last month: May 15 - June 15.
blogAI Socratic March 2025
All the most important AI news and updates from last month (Feb 20 - Mar 15).
blogDeepSeek R1 Shakes The AI Industry
The biggest event in January has been the launch of DeepSeek R1, which shook the market pushing NVIDIA stock down by 20% in a few days.
blogAI Socratic, Feb 2025 — Part 1
We collect all the most important AI updates from February.