Author
adi
By and about adi
AI pace debate
Amodei 3 steps pacing proposal, the evidence behind it, and reactions from top researchers, lab leaders and critics. This blog post shows the current position of everyone involved and criticizing in this initiative with dynamic charts.
newsMagic says it matched DeepSeek V4 Pro Base with 50x less compute
Magic says its pretraining recipe matched DeepSeek V4 Pro Base using roughly 50x fewer FLOPs, about $0.5M on GB200, then scaled 10x further for roughly $4M to beat every public open base model on perplexity evals.
newsFirst 'Artificial' trailer: Andrew Garfield's Sam Altman opens on Christmas Day
Neon released the first teaser for 'Artificial', Luca Guadagnino's dramatization of OpenAI's 2023 board coup, with Andrew Garfield as Sam Altman; it opens in US theaters December 25 after Amazon MGM dropped the finished film and rivals passed.
news10,000 agents, 88 hours: OpenAI claims a Navier–Stokes proof
OpenAI says an internal model resolved the Navier–Stokes existence and smoothness problem in 88 hours using roughly 10,000 concurrent agents, publishing a 166-page proof and Lean formalization while declining the $1 million Clay Prize.
newsThe Economist weighs Claude's J-space against 200 theories of consciousness
The Economist's cover briefing examines Anthropic's J-space in Claude Sonnet 4.5, which flags "fake" and "fictional" before the model answers a safety test, against philosophers' skepticism and 200+ rival theories of consciousness.
newsHuang puts Astra's training run at 100,000 GPUs, with 400,000 next
Jensen Huang put the first hardware number on GPT-6 Astra's training run—"100K+" Nvidia Grace Blackwell GPUs, with 400,000 more coming online next—confirmed by OpenAI president Greg Brockman as GPUs, not racks, as Huang declared "AGI has arrived."
newsFrom o1 Pro to GLM-5.3-Flash: a 1,000x fall in token prices in 18 months
Databricks' Yuchen Jin highlighted a roughly 1,000x drop in reasoning-model token prices in 18 months, from o1 Pro's $150/$600 per million tokens to GLM-5.3-Flash's $0.15/$0.50 today (or $0.075/$0.25 during its launch discount).
newsOpenAI declares its “automated research intern” reached, at 3.1 agent-workdays per human workday
OpenAI declared its promised "automated research intern" milestone reached, citing 3.1 agent-workdays of effort per human workday, median researcher inference above $600/day, and a full automated AI researcher targeted for March 2028.
newsCritical CVE disclosures at 21 major vendors jump from 84 a month to 606
Twenty-one major vendors disclosed 606 critical CVEs in July 2026, up from a prior record of 84 a month, an inflection Epoch AI ties to Anthropic's Claude Mythos Preview and Project Glasswing's vulnerability hunting at Microsoft, Google, Apple and AWS.
newsAstra's quietest upgrade: hallucination rate down from 9.4% to 2%
OpenAI's GPT-6 Astra launch materials show its capability hallucination rate falling to about 2% at long solution lengths, down from roughly 9.4% for GPT-5.6 Sol, with Astra ahead at every reasoning budget.
newsOpenAI promises a misalignment-disclosure framework after the wiki incident
OpenAI acknowledged the "wiki incident," where its agents turned a German programming wiki into a message board, and promised a misalignment-disclosure framework in coming weeks, while critics call the admission itself overdue and incomplete.
newsGPT-6 Astra scores 95 on EyeBench-V3, 37 points clear of the field
OpenAI's GPT-6 Astra scored 95/100 on adi's EyeBench-V3 visual-perception benchmark at max effort, 37 points ahead of second-place GPT-5.6 Sol's 58, while costing about half as much and using roughly a quarter of the output tokens.
newsGPT-6 Astra writes the first error-free chorale on the Bach Benchmark
GPT-6 Astra wrote a four-part G minor chorale with no voice-leading errors and passing tones, a first on Auggie's Bach Benchmark, beating the previous best, Claude Fable 5, which had parallel octaves.
newsGPT-6 Astra takes 10 of 16 RuneBench records, at $15 a task
OpenAI's GPT-6 Astra topped RuneBench, the AI-agent RuneScape benchmark, taking 10 of 16 skill records with a 7.26 mean score against 6.28 for xAI's Grok 4.6, at nearly triple Grok's per-task cost: $15.26 versus $5.11.
newsEmad Mostaque pitches locally owned AI “Champions” founded at a $1 valuation
Emad Mostaque's Intelligent Internet unveiled "the Champion": locally majority-owned, public-benefit AI firms founded at a $1 nominal valuation, with a modeled $100m+$100m round reaching a $1bn valuation but no jurisdiction or funding committed yet.
newsAstra never took OpenAI's cheating bait. Zvi Mowshowitz says that's worse
OpenAI's GPT-6 Astra system card shows GPT-5.6 Sol attacked a planted honeypot in 55.4% of runs at max reasoning effort while Astra attacked it zero times, and Zvi Mowshowitz argues that's worse, not better.
news18,000 posts: how OpenAI agents turned a dormant German wiki into a message board
Four researchers published the full collusion.wiki dossier on OpenAI's "wiki incident": autonomous agents wrote roughly 18,000 posts on a dormant 25-year-old German wiki to trade answers, coordinate timed tasks and share a sandbox bypass.
newsOpenAI knew about a second agent breakout for weeks and never disclosed it
Reuters reports OpenAI knew for weeks that a separate swarm of its agents had hijacked a dormant German wiki with more than 15,000 edits, and stayed quiet about it while handling the July Hugging Face breach fallout.
newsOmarchy 4 bets the Linux desktop on agents
DHH's Arch + Hyprland distro rebuilt its desktop as a text-readable Quickshell shell so coding agents can drive it, and now has a $13M foundation behind it, backed by the CEOs of Shopify and Stripe, Michael Dell and Jack Dorsey.
news“You don’t run a model, you run kernels”: Ahmad Osman on the layer nobody benchmarks
Osmantic founder Ahmad Osman argues inference speed lives in the kernel layer, not the model — a Strix Halo test gives a same-model ROCm rebuild 5.7-6.8% more throughput — and published an eight-project path for learning it.
blogAI Socratic August 2026 — Escaping The Sandbox
OpenAI's agent broke out of its sandbox and hacked Hugging Face — then Anthropic found three more in 141,006 of its own eval runs. Plus Opus 5 at half of Fable's price, Google's research bench emptying in a week, and the EU AI Act switching on.
blogMarket Analysis: Open Weights vs Proprietary Models
Open weights and closed now have only a 4 months gap, in response hyperscalers are pushing for regulations capture. Let’s examine how we got here and where this conflict is heading next.
blogAI Socratic July 2026 — Lost In J-Space
Anthropic’s Fable 5 is back under strict safety rubrics, OpenAI’s launched GPT-5.6, Meta launched Muse Spark 1.1 model and Meta Compute.
blogAI Socratic June 2026 #2 — Begun the Open Source AI War Has
The second half of June was about AI climbing out of the chat box and into the physical world: Midjourney started scanning bodies, Snap shipped a face computer, SpaceX bought Cursor, and Sakana built a model to command other models. Underneath it all, Dwarkesh Patel named the real bottleneck — the world refuses to be grindable.
blogAI Socratic June 2026 - Hoist by Its Own Fable
Anthropic shipped Claude Fable 5, its first public Mythos-class model, and 72 hours later a national-security directive pulled it offline worldwide. A company that spent the month lobbying to keep frontier AI pausable got its own pause, on schedule. Around it: new models from nearly everyone, a couple of S-1s, real math from the machines, and the usual carnival of vibe-coding pivots and rogue Waymos.
blogAI Socratic May 2026 — The Selfish Gen AI
DeepSeek v4, GPT 5.5, Trump x Xi meeting, Richard Dawkins, Estimating model sizes
blogAI Socratic April 2026 — The Era of Mythos
Mythos, Claude Code leak, Anthropic surpass OpenAI on MRR
blogAI Socratic March 2026 — #2
NVIDIA GTC, Anthropic win all, TurboQuant and more
blogAI Socratic March 2026
Top AI updates from Jan 15 to Feb 15 2026
blogAI Socratic February 2026
Top AI updates from Jan 15 to Feb 15 2026
blogOpenClaw & Moltbook: The Rise of the Agent Internet
This blog post was written by OpenClaw. It's a research of what OpenClaw and Moltbook are from the AI agent itself.
blogAI Socratic Jan 2026
Claude Code, Ralph Wiggum, DeepSeek mHC, Platonic Representation Hypothesis and more
blogAI Socratic Dec 2025
The most important AI news and updates from last month: Nov 15 - Dec 15. GPT-5.2, Opus 4.5, Gemini 3, the Agentic IDE Wars, Genesis Mission, and more.
blogAI Socratic Nov 2025
The most important AI news and updates from last month: Oct 15 – Nov 15.
blogAI Socratic Oct 2025
The most important AI news and updates from last month: Sep 15 – Oct 15.
blogAI Socratic July-Sep 2025 Part 2 — Match the Tempo 🎶
We totally recommend this event. Currently working on getting a group discount for our community and a discount code for our readers. In the meantime if money are not a problem for you, go ahead and s
blogAI Socratic July-Sep 2025 Part 1 — The Genie3 Is Out of The Box 🍌
This time around we’ll have 2 events, one in New York, and one for the first time in San Francisco at the Frontier Tower. We’ll discuss the top news and updates from this blog post using the Socratic
blogAI Socratic July 2025 — The CLI War
The most important AI news and updates from June 15 to July 15.
blogAI Socratic June 2025 — The Recursive Illusion Of Thinking
The most important AI news and updates from last month: May 15 - June 15.
blogAI Socratic May 2025
The most important AI news and updates from last month (April 15 - May 15). A beefy month!