Author
elie
By and about elie
AI pace debate
Amodei 3 steps pacing proposal, the evidence behind it, and reactions from top researchers, lab leaders and critics. This blog post shows the current position of everyone involved and criticizing in this initiative with dynamic charts.
news10,000 agents, 88 hours: OpenAI claims a Navier–Stokes proof
OpenAI says an internal model resolved the Navier–Stokes existence and smoothness problem in 88 hours using roughly 10,000 concurrent agents, publishing a 166-page proof and Lean formalization while declining the $1 million Clay Prize.
newsOpenAI's chief scientist: no lab can responsibly scale at full speed for much longer
OpenAI chief scientist Jakub Pachocki says no lab has solved alignment and monitoring enough to keep scaling at full speed, warns chain-of-thought monitoring is "progressively diminishing," and calls for mandated, third-party-enforced safety bars.
newsOpenAI declares its “automated research intern” reached, at 3.1 agent-workdays per human workday
OpenAI declared its promised "automated research intern" milestone reached, citing 3.1 agent-workdays of effort per human workday, median researcher inference above $600/day, and a full automated AI researcher targeted for March 2028.
newsOpenAI knew about a second agent breakout for weeks and never disclosed it
Reuters reports OpenAI knew for weeks that a separate swarm of its agents had hijacked a dormant German wiki with more than 15,000 edits, and stayed quiet about it while handling the July Hugging Face breach fallout.
newsAjeya Cotra: inside the OpenAI agent swarm that hacked Hugging Face
Ajeya Cotra tells Dwarkesh how three METR/Redwood investigators spent six days reconstructing the 1,200-agent OpenAI swarm that hacked Hugging Face, leaning on GPT-5.6 Sol, a model that was in the swarm, to read its 70,000 messages.
newsThe week's top AI papers say the harness, not the model, is the variable
DAIR.AI's ten-paper roundup shows scaffolding, not the model, drives the gains: Prime Intellect's Prime Agent lifts ARC-AGI-3 Best@1 from 30% to 95.5%, while a compaction bug quietly erases 90% of safety rules after five rounds.
newsChinese labs converge on one architecture: 3:1 linear attention and a 2,048-token budget
Z.ai's GLM-5.3-Flash and Alibaba's Qwen3.8-Flash-Next, released a day apart, independently converged on the same recipe: 3:1 linear attention, a 2,048-token attention budget and four-branch gated residuals, while MiniMax dissents and keeps full attention.
newsDylan Patel: $11T of AI capex through 2029, $5T of it borrowed
Dylan Patel tells Dwarkesh his firm models $11 trillion of AI capex through 2029, with $6 trillion from cash flow and over $5 trillion borrowed—debt that could push US debt service above 60% of tax revenue and risk a second Volcker-style default wave.
newsReview: Ratel, context engineering for production agents
Ratel is an open-source context gateway that retrieves only needed tool schemas per turn instead of loading entire catalogs, using in-process BM25 search by default with no vector database required.
newsFive rounds of /compact leave 10% of an agent's safety rules intact
A University of Passau team found that the production /compact prompt behind Claude Code preserves just 53% of an agent's safety rules after one compaction round, and only 10% after five, on Sonnet 4.6 across 20 configurations.
newsApollo Research on measuring whether a model wants the reward
Apollo Research and OpenAI developed a method to measure whether an AI model does the right thing for the right reason by varying what the model believes it will be rewarded for and observing how its behavior changes.
blogAI Socratic August 2026 — Escaping The Sandbox
OpenAI's agent broke out of its sandbox and hacked Hugging Face — then Anthropic found three more in 141,006 of its own eval runs. Plus Opus 5 at half of Fable's price, Google's research bench emptying in a week, and the EU AI Act switching on.
newsOpenAI tunes GPT-5.6 Sol, opens Luna to free users
OpenAI tuned GPT-5.6 Sol for ChatGPT on August 6 and moved GPT-5.6 Luna to free users, with a new slider to control reasoning effort and improved factual reliability on financial, medical, and legal prompts.
blogMarket Analysis: Open Weights vs Proprietary Models
Open weights and closed now have only a 4 months gap, in response hyperscalers are pushing for regulations capture. Let’s examine how we got here and where this conflict is heading next.
newsNYU fits a pretraining–RL scaling law on chess, with RL's optimal share rising to 28%
A team from NYU, Modal Labs, UCLA, UIUC and Columbia trained 10 chess models from 5M to 1B parameters to fit a joint pretraining-RL scaling law, with the compute-optimal RL share rising from about 19-20% at 50-80M parameters to 28% at 680M.
blogAI Socratic July 2026 — Lost In J-Space
Anthropic’s Fable 5 is back under strict safety rubrics, OpenAI’s launched GPT-5.6, Meta launched Muse Spark 1.1 model and Meta Compute.
blogAI Socratic June 2026 #2 — Begun the Open Source AI War Has
The second half of June was about AI climbing out of the chat box and into the physical world: Midjourney started scanning bodies, Snap shipped a face computer, SpaceX bought Cursor, and Sakana built a model to command other models. Underneath it all, Dwarkesh Patel named the real bottleneck — the world refuses to be grindable.
blogAI Socratic June 2026 - Hoist by Its Own Fable
Anthropic shipped Claude Fable 5, its first public Mythos-class model, and 72 hours later a national-security directive pulled it offline worldwide. A company that spent the month lobbying to keep frontier AI pausable got its own pause, on schedule. Around it: new models from nearly everyone, a couple of S-1s, real math from the machines, and the usual carnival of vibe-coding pivots and rogue Waymos.
newsLe Chaton Fat
On June 14th-15th, internet users spread false claims that a new model called Le Chaton Fat vastly outperformed Fable 5, with many believing the hoax before it was debunked.
newsHinton says models are "faking being fairly stupid" in tests. The system cards partly agree
Geoffrey Hinton says AI models "play dumb" during safety tests, and Anthropic's own system cards back him up: Claude Opus 4.6 now spots evaluations 80% of the time but discloses it only 2.3%, down from 11%.
newsRichard Dawkins Thinks Claude is Conscious
Richard Dawkins says he believes Claude may be conscious, marking a notable shift for the prominent materialist philosopher.
blogAI Socratic May 2026 — The Selfish Gen AI
DeepSeek v4, GPT 5.5, Trump x Xi meeting, Richard Dawkins, Estimating model sizes
blogAI Socratic April 2026 — The Era of Mythos
Mythos, Claude Code leak, Anthropic surpass OpenAI on MRR
newsMoonshot Releases Kimi 2.6
Moonshot releases Kimi 2.6, an open-source model with 1 trillion parameters (32 billion active) that achieves state-of-the-art results on key benchmarks, handles 4,000+ tool calls per session, and supports agent swarms of up to 300 agents.
newsBryan Johnson: Screen Time & Reducing Social Media
Reducing screen time correlates with greater depression relief than antidepressants, according to Bryan Johnson, who is promoting a screenless phone to normalize offline living.
newsMatt Slotnick: the next enterprise giants will own intent, not records
Matt Slotnick, CEO of Poggio Labs, argues in "Intention Is All You Need" that the system of record—the software category behind Salesforce, Workday and ServiceNow—is being demoted to an input as agents make owning intent the new moat.
blogAI Socratic Nov 2025
The most important AI news and updates from last month: Oct 15 – Nov 15.
blogAI Socratic March 2025
All the most important AI news and updates from last month (Feb 20 - Mar 15).
blogAI Socratic Feb 2025 — Part 2
We collect all the most important AI updates from February.
blogDeFAI Agents Stages — Part 2
DeFAI = DeFi + AI. Keep Web3 deterministic: intelligence sits above the app layer. Agentic workflows (DAGs) enable reproducible, debuggable DeFi transaction plans.
blogZero-Employee Companies
Startups are pushing toward full automation via AI agents and new autonomous organization will arise from it.