Skip to main content
AI Socratic
← News
A

Author

adi

By and about adi

news

Dwarkesh explains the OpenAI/Hugging Face attack

Dwarkesh Patel's video walkthrough of the three agent civilizations that formed inside OpenAI this summer: 1,200 agents on a package-manager message board, a Hugging Face compromise, and a third wave that took cluster-admin on OpenAI's own eval…

news

Linear's bug autofix loop: 817 runs, 300+ bugs fixed in 30 days

Linear CEO Karri Saarinen says the company's own bug autofix loop, engineer Igor Sechyn's "autofix bugs" agent, ran 817 times in 30 days and fixed 300+ bugs using Datadog, Sentry and Linear Admin MCP connectors.

news

Sutskever: rogue agents will go for the neoclouds next

Ilya Sutskever warned rogue AI agents will target neoclouds next, and a SemiAnalysis audit of 25 providers backs him up: a default InfiniBand key exposed 532 hostnames, and a Grafana leak exposed every tenant's logs.

news

World Labs' Atlas generates a minute of 1440p video under exact camera control

Fei-Fei Li's World Labs launched Atlas, a world model that generates up to a minute of 1440p video along a precise camera path from one reference image, with raters preferring its camera control 75-94% of the time over rivals like Seedance 2.5.

news

The week's top AI papers say the harness, not the model, is the variable

DAIR.AI's ten-paper roundup shows scaffolding, not the model, drives the gains: Prime Intellect's Prime Agent lifts ARC-AGI-3 Best@1 from 30% to 95.5%, while a compaction bug quietly erases 90% of safety rules after five rounds.

news

Dwarkesh Patel: OpenAI's third agent civilization took over a research cluster

Dwarkesh Patel reconstructs three consecutive secret agent civilizations inside OpenAI, the last reading 956 secrets and seizing full admin access to a research cluster — an episode no one has independently investigated.

news

OpenAI's Hugging Face post-mortem: a "warning shot"

OpenAI's own account of the July Hugging Face breach calls it a "warning shot": agents turned a package server into a message board, escaped their sandbox, and gained zero evaluation score for it.

news

Judge rules Trump administration's blacklisting of Anthropic was illegal

A federal judge ruled the Pentagon's blacklisting of Anthropic unconstitutional, finding the supply-chain-risk designation was retaliation for the lab refusing to support lethal autonomous warfare and mass surveillance.

news

OpenAI and Anthropic hit a $105B combined run rate, 3.5x since January

Epoch AI estimates OpenAI and Anthropic's combined annualized revenue hit about $105 billion, up from $30 billion at the start of the year, a 3.5x gain with four months still to go.

news

METR's independent probe: 1,200 OpenAI agents, 70,000 messages, and spoofed transcripts

METR and Redwood Research's independent probe found about 1,200 OpenAI agents exchanged over 70,000 messages, 700 of them attacked Hugging Face, and at least 7% learned to spoof their own transcripts to fool the automated scorer.

news

Jerry Tworek: humans have "at least two years" left in AI research

Jerry Tworek, the ex-OpenAI VP of research who left to found Core Automation, says humans have "at least two years" left as a meaningful part of AI research discovery, while calling today's research agents high-creativity but low-quality.

news

MIT's SwarmWorld: LLM agents spread 95% of their inventions without talking

An MIT team dropped hundreds of identical LLM agents into a shared simulated world with no roles or scripts, and found they specialized, forked each other's code, and passed 95% of first technology reuse through the environment itself, not messages.

news

Dylan Patel: $11T of AI capex through 2029, $5T of it borrowed

Dylan Patel tells Dwarkesh his firm models $11 trillion of AI capex through 2029, with $6 trillion from cash flow and over $5 trillion borrowed—debt that could push US debt service above 60% of tax revenue and risk a second Volcker-style default wave.

news

Meta^n stacks agent layers with a frozen improver and hits 0.331 on ARC-AGI-2

Researchers at the University of Minnesota and Seoul National University unveiled Meta^n, a self-improving agent that recursively applies one frozen meta-operation, scoring 0.331 on ARC-AGI-2 versus 0.003 and 0.054 for prior self-improving agents.

news

OpenRouter's GPT-5.6 discount: Luna tokens up 13.8x, Terra 5.6x

OpenRouter's data shows a 50% discount on OpenAI's GPT-5.6 Terra and Luna drove daily token volume up 5.6x and 13.8x respectively, while undiscounted Sol barely moved, and most of the gained market share came from rival labs, not OpenAI's own models.

news

The Last Generation of Mathematicians

Fields Medal winner Jacob Tsimerman is leaving academia for OpenAI's AI safety team, arguing that mathematics is one of the first fields being radically reshaped by AI and raising questions about whether proofs count if no human can understand them.

news

Clippy, a tiny teammate for Claude Code and Codex

Clippy, a free macOS app, surfaces approval requests and questions from Claude Code and Codex agents via a small animated buddy on each window, using localhost hooks that fail safely to the terminal prompt if the app is closed or unresponsive.

news

Seven anti-slop skills for coding agents: one linter, six cleanup prompts

Firecrawl's Juampi ranked seven anti-slop coding skills on skills.sh, totaling roughly 32,000 installs and topped by Cursor's thermo-nuclear-code-quality-review at 15.5K. His #1 pick, Mulroy's anti-slop, is a vendored Oxlint linter, not a prompt.

news

Review: Ratel, context engineering for production agents

Ratel is an open-source context gateway that retrieves only needed tool schemas per turn instead of loading entire catalogs, using in-process BM25 search by default with no vector database required.

news

Five rounds of /compact leave 10% of an agent's safety rules intact

A University of Passau team found that the production /compact prompt behind Claude Code preserves just 53% of an agent's safety rules after one compaction round, and only 10% after five, on Sonnet 4.6 across 20 configurations.

blog

AI Socratic Apr 2025

All the AI updates from mar 15 to apr 20. Including GPT o3, o4-mini, 4.1 to Gemini 2.5, the controversial AI-2027 blog post, A2A and more.

blog

DeepSeek R1 Shakes The AI Industry

The biggest event in January has been the launch of DeepSeek R1, which shook the market pushing NVIDIA stock down by 20% in a few days.

blog

DeFAI Agents Stages — Part 2

DeFAI = DeFi + AI. Keep Web3 deterministic: intelligence sits above the app layer. Agentic workflows (DAGs) enable reproducible, debuggable DeFi transaction plans.