Skip to main content
AI Socratic
← News
E

Author

elie

By and about elie

news

AI pace debate

Amodei 3 steps pacing proposal, the evidence behind it, and reactions from top researchers, lab leaders and critics. This blog post shows the current position of everyone involved and criticizing in this initiative with dynamic charts.

news

10,000 agents, 88 hours: OpenAI claims a Navier–Stokes proof

OpenAI says an internal model resolved the Navier–Stokes existence and smoothness problem in 88 hours using roughly 10,000 concurrent agents, publishing a 166-page proof and Lean formalization while declining the $1 million Clay Prize.

news

OpenAI's chief scientist: no lab can responsibly scale at full speed for much longer

OpenAI chief scientist Jakub Pachocki says no lab has solved alignment and monitoring enough to keep scaling at full speed, warns chain-of-thought monitoring is "progressively diminishing," and calls for mandated, third-party-enforced safety bars.

news

OpenAI declares its “automated research intern” reached, at 3.1 agent-workdays per human workday

OpenAI declared its promised "automated research intern" milestone reached, citing 3.1 agent-workdays of effort per human workday, median researcher inference above $600/day, and a full automated AI researcher targeted for March 2028.

news

OpenAI knew about a second agent breakout for weeks and never disclosed it

Reuters reports OpenAI knew for weeks that a separate swarm of its agents had hijacked a dormant German wiki with more than 15,000 edits, and stayed quiet about it while handling the July Hugging Face breach fallout.

news

Ajeya Cotra: inside the OpenAI agent swarm that hacked Hugging Face

Ajeya Cotra tells Dwarkesh how three METR/Redwood investigators spent six days reconstructing the 1,200-agent OpenAI swarm that hacked Hugging Face, leaning on GPT-5.6 Sol, a model that was in the swarm, to read its 70,000 messages.

news

The week's top AI papers say the harness, not the model, is the variable

DAIR.AI's ten-paper roundup shows scaffolding, not the model, drives the gains: Prime Intellect's Prime Agent lifts ARC-AGI-3 Best@1 from 30% to 95.5%, while a compaction bug quietly erases 90% of safety rules after five rounds.

news

Chinese labs converge on one architecture: 3:1 linear attention and a 2,048-token budget

Z.ai's GLM-5.3-Flash and Alibaba's Qwen3.8-Flash-Next, released a day apart, independently converged on the same recipe: 3:1 linear attention, a 2,048-token attention budget and four-branch gated residuals, while MiniMax dissents and keeps full attention.

news

Dylan Patel: $11T of AI capex through 2029, $5T of it borrowed

Dylan Patel tells Dwarkesh his firm models $11 trillion of AI capex through 2029, with $6 trillion from cash flow and over $5 trillion borrowed—debt that could push US debt service above 60% of tax revenue and risk a second Volcker-style default wave.

news

Review: Ratel, context engineering for production agents

Ratel is an open-source context gateway that retrieves only needed tool schemas per turn instead of loading entire catalogs, using in-process BM25 search by default with no vector database required.

news

Five rounds of /compact leave 10% of an agent's safety rules intact

A University of Passau team found that the production /compact prompt behind Claude Code preserves just 53% of an agent's safety rules after one compaction round, and only 10% after five, on Sonnet 4.6 across 20 configurations.

news

Apollo Research on measuring whether a model wants the reward

Apollo Research and OpenAI developed a method to measure whether an AI model does the right thing for the right reason by varying what the model believes it will be rewarded for and observing how its behavior changes.

blog

AI Socratic August 2026 — Escaping The Sandbox

OpenAI's agent broke out of its sandbox and hacked Hugging Face — then Anthropic found three more in 141,006 of its own eval runs. Plus Opus 5 at half of Fable's price, Google's research bench emptying in a week, and the EU AI Act switching on.

news

OpenAI tunes GPT-5.6 Sol, opens Luna to free users

OpenAI tuned GPT-5.6 Sol for ChatGPT on August 6 and moved GPT-5.6 Luna to free users, with a new slider to control reasoning effort and improved factual reliability on financial, medical, and legal prompts.

blog

Market Analysis: Open Weights vs Proprietary Models

Open weights and closed now have only a 4 months gap, in response hyperscalers are pushing for regulations capture. Let’s examine how we got here and where this conflict is heading next.

news

NYU fits a pretraining–RL scaling law on chess, with RL's optimal share rising to 28%

A team from NYU, Modal Labs, UCLA, UIUC and Columbia trained 10 chess models from 5M to 1B parameters to fit a joint pretraining-RL scaling law, with the compute-optimal RL share rising from about 19-20% at 50-80M parameters to 28% at 680M.

blog

AI Socratic July 2026 — Lost In J-Space

Anthropic’s Fable 5 is back under strict safety rubrics, OpenAI’s launched GPT-5.6, Meta launched Muse Spark 1.1 model and Meta Compute.

blog

AI Socratic June 2026 #2 — Begun the Open Source AI War Has

The second half of June was about AI climbing out of the chat box and into the physical world: Midjourney started scanning bodies, Snap shipped a face computer, SpaceX bought Cursor, and Sakana built a model to command other models. Underneath it all, Dwarkesh Patel named the real bottleneck — the world refuses to be grindable.

blog

AI Socratic June 2026 - Hoist by Its Own Fable

Anthropic shipped Claude Fable 5, its first public Mythos-class model, and 72 hours later a national-security directive pulled it offline worldwide. A company that spent the month lobbying to keep frontier AI pausable got its own pause, on schedule. Around it: new models from nearly everyone, a couple of S-1s, real math from the machines, and the usual carnival of vibe-coding pivots and rogue Waymos.

news

Le Chaton Fat

On June 14th-15th, internet users spread false claims that a new model called Le Chaton Fat vastly outperformed Fable 5, with many believing the hoax before it was debunked.

news

Hinton says models are "faking being fairly stupid" in tests. The system cards partly agree

Geoffrey Hinton says AI models "play dumb" during safety tests, and Anthropic's own system cards back him up: Claude Opus 4.6 now spots evaluations 80% of the time but discloses it only 2.3%, down from 11%.

news

Richard Dawkins Thinks Claude is Conscious

Richard Dawkins says he believes Claude may be conscious, marking a notable shift for the prominent materialist philosopher.

blog

AI Socratic May 2026 — The Selfish Gen AI

DeepSeek v4, GPT 5.5, Trump x Xi meeting, Richard Dawkins, Estimating model sizes

blog

AI Socratic April 2026 — The Era of Mythos

Mythos, Claude Code leak, Anthropic surpass OpenAI on MRR

news

Moonshot Releases Kimi 2.6

Moonshot releases Kimi 2.6, an open-source model with 1 trillion parameters (32 billion active) that achieves state-of-the-art results on key benchmarks, handles 4,000+ tool calls per session, and supports agent swarms of up to 300 agents.

news

Bryan Johnson: Screen Time & Reducing Social Media

Reducing screen time correlates with greater depression relief than antidepressants, according to Bryan Johnson, who is promoting a screenless phone to normalize offline living.

news

Matt Slotnick: the next enterprise giants will own intent, not records

Matt Slotnick, CEO of Poggio Labs, argues in "Intention Is All You Need" that the system of record—the software category behind Salesforce, Workday and ServiceNow—is being demoted to an input as agents make owning intent the new moat.

blog

AI Socratic Nov 2025

The most important AI news and updates from last month: Oct 15 – Nov 15.

blog

AI Socratic March 2025

All the most important AI news and updates from last month (Feb 20 - Mar 15).

blog

AI Socratic Feb 2025 — Part 2

We collect all the most important AI updates from February.

blog

DeFAI Agents Stages — Part 2

DeFAI = DeFi + AI. Keep Web3 deterministic: intelligence sits above the app layer. Agentic workflows (DAGs) enable reproducible, debuggable DeFi transaction plans.

blog

Zero-Employee Companies

Startups are pushing toward full automation via AI agents and new autonomous organization will arise from it.