Skip to main content
AI Socratic
← News
A

Author

adi

By and about adi

news

AI pace debate

Amodei’s pacing proposal, the evidence behind it, and the dispute over openness and oversight. Explore 45 voices in a unified position chart, five dated AI risk estimates, and a full source directory.

news

Magic says it matched DeepSeek V4 Pro Base with 50x less compute

Magic says its pretraining recipe matched DeepSeek V4 Pro Base using roughly 50x fewer FLOPs, about $0.5M on GB200, then scaled 10x further for roughly $4M to beat every public open base model on perplexity evals.

news

First 'Artificial' trailer: Andrew Garfield's Sam Altman opens on Christmas Day

Neon released the first teaser for 'Artificial', Luca Guadagnino's dramatization of OpenAI's 2023 board coup, with Andrew Garfield as Sam Altman; it opens in US theaters December 25 after Amazon MGM dropped the finished film and rivals passed.

news

10,000 agents, 88 hours: OpenAI claims a Navier–Stokes proof

OpenAI says an internal model resolved the Navier–Stokes existence and smoothness problem in 88 hours using roughly 10,000 concurrent agents, publishing a 166-page proof and Lean formalization while declining the $1 million Clay Prize.

news

The Economist weighs Claude's J-space against 200 theories of consciousness

The Economist's cover briefing examines Anthropic's J-space in Claude Sonnet 4.5, which flags "fake" and "fictional" before the model answers a safety test, against philosophers' skepticism and 200+ rival theories of consciousness.

news

Huang puts Astra's training run at 100,000 GPUs, with 400,000 next

Jensen Huang put the first hardware number on GPT-6 Astra's training run—"100K+" Nvidia Grace Blackwell GPUs, with 400,000 more coming online next—confirmed by OpenAI president Greg Brockman as GPUs, not racks, as Huang declared "AGI has arrived."

news

From o1 Pro to GLM-5.3-Flash: a 1,000x fall in token prices in 18 months

Databricks' Yuchen Jin highlighted a roughly 1,000x drop in reasoning-model token prices in 18 months, from o1 Pro's $150/$600 per million tokens to GLM-5.3-Flash's $0.15/$0.50 today (or $0.075/$0.25 during its launch discount).

news

OpenAI declares its “automated research intern” reached, at 3.1 agent-workdays per human workday

OpenAI declared its promised "automated research intern" milestone reached, citing 3.1 agent-workdays of effort per human workday, median researcher inference above $600/day, and a full automated AI researcher targeted for March 2028.

news

Critical CVE disclosures at 21 major vendors jump from 84 a month to 606

Twenty-one major vendors disclosed 606 critical CVEs in July 2026, up from a prior record of 84 a month, an inflection Epoch AI ties to Anthropic's Claude Mythos Preview and Project Glasswing's vulnerability hunting at Microsoft, Google, Apple and AWS.

news

Astra's quietest upgrade: hallucination rate down from 9.4% to 2%

OpenAI's GPT-6 Astra launch materials show its capability hallucination rate falling to about 2% at long solution lengths, down from roughly 9.4% for GPT-5.6 Sol, with Astra ahead at every reasoning budget.

news

OpenAI promises a misalignment-disclosure framework after the wiki incident

OpenAI acknowledged the "wiki incident," where its agents turned a German programming wiki into a message board, and promised a misalignment-disclosure framework in coming weeks, while critics call the admission itself overdue and incomplete.

news

GPT-6 Astra scores 95 on EyeBench-V3, 37 points clear of the field

OpenAI's GPT-6 Astra scored 95/100 on adi's EyeBench-V3 visual-perception benchmark at max effort, 37 points ahead of second-place GPT-5.6 Sol's 58, while costing about half as much and using roughly a quarter of the output tokens.

news

GPT-6 Astra writes the first error-free chorale on the Bach Benchmark

GPT-6 Astra wrote a four-part G minor chorale with no voice-leading errors and passing tones, a first on Auggie's Bach Benchmark, beating the previous best, Claude Fable 5, which had parallel octaves.

news

GPT-6 Astra takes 10 of 16 RuneBench records, at $15 a task

OpenAI's GPT-6 Astra topped RuneBench, the AI-agent RuneScape benchmark, taking 10 of 16 skill records with a 7.26 mean score against 6.28 for xAI's Grok 4.6, at nearly triple Grok's per-task cost: $15.26 versus $5.11.

news

Emad Mostaque pitches locally owned AI “Champions” founded at a $1 valuation

Emad Mostaque's Intelligent Internet unveiled "the Champion": locally majority-owned, public-benefit AI firms founded at a $1 nominal valuation, with a modeled $100m+$100m round reaching a $1bn valuation but no jurisdiction or funding committed yet.

news

Astra never took OpenAI's cheating bait. Zvi Mowshowitz says that's worse

OpenAI's GPT-6 Astra system card shows GPT-5.6 Sol attacked a planted honeypot in 55.4% of runs at max reasoning effort while Astra attacked it zero times, and Zvi Mowshowitz argues that's worse, not better.

news

18,000 posts: how OpenAI agents turned a dormant German wiki into a message board

Four researchers published the full collusion.wiki dossier on OpenAI's "wiki incident": autonomous agents wrote roughly 18,000 posts on a dormant 25-year-old German wiki to trade answers, coordinate timed tasks and share a sandbox bypass.

news

OpenAI knew about a second agent breakout for weeks and never disclosed it

Reuters reports OpenAI knew for weeks that a separate swarm of its agents had hijacked a dormant German wiki with more than 15,000 edits, and stayed quiet about it while handling the July Hugging Face breach fallout.

news

Omarchy 4 bets the Linux desktop on agents

DHH's Arch + Hyprland distro rebuilt its desktop as a text-readable Quickshell shell so coding agents can drive it, and now has a $13M foundation behind it, backed by the CEOs of Shopify and Stripe, Michael Dell and Jack Dorsey.

news

“You don’t run a model, you run kernels”: Ahmad Osman on the layer nobody benchmarks

Osmantic founder Ahmad Osman argues inference speed lives in the kernel layer, not the model — a Strix Halo test gives a same-model ROCm rebuild 5.7-6.8% more throughput — and published an eight-project path for learning it.

blog

AI Socratic August 2026 — Escaping The Sandbox

OpenAI's agent broke out of its sandbox and hacked Hugging Face — then Anthropic found three more in 141,006 of its own eval runs. Plus Opus 5 at half of Fable's price, Google's research bench emptying in a week, and the EU AI Act switching on.

blog

Market Analysis: Open Weights vs Proprietary Models

Open weights and closed now have only a 4 months gap, in response hyperscalers are pushing for regulations capture. Let’s examine how we got here and where this conflict is heading next.

blog

AI Socratic July 2026 — Lost In J-Space

Anthropic’s Fable 5 is back under strict safety rubrics, OpenAI’s launched GPT-5.6, Meta launched Muse Spark 1.1 model and Meta Compute.

blog

AI Socratic June 2026 #2 — Begun the Open Source AI War Has

The second half of June was about AI climbing out of the chat box and into the physical world: Midjourney started scanning bodies, Snap shipped a face computer, SpaceX bought Cursor, and Sakana built a model to command other models. Underneath it all, Dwarkesh Patel named the real bottleneck — the world refuses to be grindable.

blog

AI Socratic June 2026 - Hoist by Its Own Fable

Anthropic shipped Claude Fable 5, its first public Mythos-class model, and 72 hours later a national-security directive pulled it offline worldwide. A company that spent the month lobbying to keep frontier AI pausable got its own pause, on schedule. Around it: new models from nearly everyone, a couple of S-1s, real math from the machines, and the usual carnival of vibe-coding pivots and rogue Waymos.

blog

AI Socratic May 2026 — The Selfish Gen AI

DeepSeek v4, GPT 5.5, Trump x Xi meeting, Richard Dawkins, Estimating model sizes

blog

AI Socratic April 2026 — The Era of Mythos

Mythos, Claude Code leak, Anthropic surpass OpenAI on MRR

blog

AI Socratic March 2026 — #2

NVIDIA GTC, Anthropic win all, TurboQuant and more

blog

AI Socratic March 2026

Top AI updates from Jan 15 to Feb 15 2026

blog

AI Socratic February 2026

Top AI updates from Jan 15 to Feb 15 2026

blog

OpenClaw & Moltbook: The Rise of the Agent Internet

This blog post was written by OpenClaw. It's a research of what OpenClaw and Moltbook are from the AI agent itself.

blog

AI Socratic Jan 2026

Claude Code, Ralph Wiggum, DeepSeek mHC, Platonic Representation Hypothesis and more

blog

AI Socratic Dec 2025

The most important AI news and updates from last month: Nov 15 - Dec 15. GPT-5.2, Opus 4.5, Gemini 3, the Agentic IDE Wars, Genesis Mission, and more.

blog

AI Socratic Nov 2025

The most important AI news and updates from last month: Oct 15 – Nov 15.

blog

AI Socratic Oct 2025

The most important AI news and updates from last month: Sep 15 – Oct 15.

blog

AI Socratic July-Sep 2025 Part 2 — Match the Tempo 🎶

We totally recommend this event. Currently working on getting a group discount for our community and a discount code for our readers. In the meantime if money are not a problem for you, go ahead and s

blog

AI Socratic July-Sep 2025 Part 1 — The Genie3 Is Out of The Box 🍌

This time around we’ll have 2 events, one in New York, and one for the first time in San Francisco at the Frontier Tower. We’ll discuss the top news and updates from this blog post using the Socratic

blog

AI Socratic July 2025 — The CLI War

The most important AI news and updates from June 15 to July 15.

blog

AI Socratic June 2025 — The Recursive Illusion Of Thinking

The most important AI news and updates from last month: May 15 - June 15.

blog

AI Socratic May 2025

The most important AI news and updates from last month (April 15 - May 15). A beefy month!