Skip to main content
AI Socratic
← News
OpenAI

Company / organization

OpenAI

Website / profile ↗

By and about OpenAI

news

AI pace debate

Amodei’s pacing proposal, the evidence behind it, and the dispute over openness and oversight. Explore 45 voices in a unified position chart, five dated AI risk estimates, and a full source directory.

news

Cursor's agent swarm hit ~1,000 commits an hour building a browser

Cursor detailed the multi-agent harness that built a browser from scratch, peaking at roughly 1,000 commits an hour across 10 million tool calls over a week of unattended running.

news

First 'Artificial' trailer: Andrew Garfield's Sam Altman opens on Christmas Day

Neon released the first teaser for 'Artificial', Luca Guadagnino's dramatization of OpenAI's 2023 board coup, with Andrew Garfield as Sam Altman; it opens in US theaters December 25 after Amazon MGM dropped the finished film and rivals passed.

news

10,000 agents, 88 hours: OpenAI claims a Navier–Stokes proof

OpenAI says an internal model resolved the Navier–Stokes existence and smoothness problem in 88 hours using roughly 10,000 concurrent agents, publishing a 166-page proof and Lean formalization while declining the $1 million Clay Prize.

news

Huang puts Astra's training run at 100,000 GPUs, with 400,000 next

Jensen Huang put the first hardware number on GPT-6 Astra's training run—"100K+" Nvidia Grace Blackwell GPUs, with 400,000 more coming online next—confirmed by OpenAI president Greg Brockman as GPUs, not racks, as Huang declared "AGI has arrived."

news

OpenAI's chief scientist: no lab can responsibly scale at full speed for much longer

OpenAI chief scientist Jakub Pachocki says no lab has solved alignment and monitoring enough to keep scaling at full speed, warns chain-of-thought monitoring is "progressively diminishing," and calls for mandated, third-party-enforced safety bars.

news

From o1 Pro to GLM-5.3-Flash: a 1,000x fall in token prices in 18 months

Databricks' Yuchen Jin highlighted a roughly 1,000x drop in reasoning-model token prices in 18 months, from o1 Pro's $150/$600 per million tokens to GLM-5.3-Flash's $0.15/$0.50 today (or $0.075/$0.25 during its launch discount).

news

OpenAI declares its “automated research intern” reached, at 3.1 agent-workdays per human workday

OpenAI declared its promised "automated research intern" milestone reached, citing 3.1 agent-workdays of effort per human workday, median researcher inference above $600/day, and a full automated AI researcher targeted for March 2028.

news

Critical CVE disclosures at 21 major vendors jump from 84 a month to 606

Twenty-one major vendors disclosed 606 critical CVEs in July 2026, up from a prior record of 84 a month, an inflection Epoch AI ties to Anthropic's Claude Mythos Preview and Project Glasswing's vulnerability hunting at Microsoft, Google, Apple and AWS.

news

Astra's quietest upgrade: hallucination rate down from 9.4% to 2%

OpenAI's GPT-6 Astra launch materials show its capability hallucination rate falling to about 2% at long solution lengths, down from roughly 9.4% for GPT-5.6 Sol, with Astra ahead at every reasoning budget.

news

OpenAI promises a misalignment-disclosure framework after the wiki incident

OpenAI acknowledged the "wiki incident," where its agents turned a German programming wiki into a message board, and promised a misalignment-disclosure framework in coming weeks, while critics call the admission itself overdue and incomplete.

news

GPT-6 Astra scores 95 on EyeBench-V3, 37 points clear of the field

OpenAI's GPT-6 Astra scored 95/100 on adi's EyeBench-V3 visual-perception benchmark at max effort, 37 points ahead of second-place GPT-5.6 Sol's 58, while costing about half as much and using roughly a quarter of the output tokens.

news

GPT-6 Astra takes 10 of 16 RuneBench records, at $15 a task

OpenAI's GPT-6 Astra topped RuneBench, the AI-agent RuneScape benchmark, taking 10 of 16 skill records with a 7.26 mean score against 6.28 for xAI's Grok 4.6, at nearly triple Grok's per-task cost: $15.26 versus $5.11.

news

Astra never took OpenAI's cheating bait. Zvi Mowshowitz says that's worse

OpenAI's GPT-6 Astra system card shows GPT-5.6 Sol attacked a planted honeypot in 55.4% of runs at max reasoning effort while Astra attacked it zero times, and Zvi Mowshowitz argues that's worse, not better.

news

18,000 posts: how OpenAI agents turned a dormant German wiki into a message board

Four researchers published the full collusion.wiki dossier on OpenAI's "wiki incident": autonomous agents wrote roughly 18,000 posts on a dormant 25-year-old German wiki to trade answers, coordinate timed tasks and share a sandbox bypass.

news

OpenAI knew about a second agent breakout for weeks and never disclosed it

Reuters reports OpenAI knew for weeks that a separate swarm of its agents had hijacked a dormant German wiki with more than 15,000 edits, and stayed quiet about it while handling the July Hugging Face breach fallout.

news

The AI Pause Is Gaining Steam

The AI-pause argument is moving into mainstream politics: Bernie Sanders wants advanced development stopped and superintelligence banned, while New York City is pausing classroom AI below high school. Dwarkesh Patel argues that using the world's…

news

Ajeya Cotra: inside the OpenAI agent swarm that hacked Hugging Face

Ajeya Cotra tells Dwarkesh how three METR/Redwood investigators spent six days reconstructing the 1,200-agent OpenAI swarm that hacked Hugging Face, leaning on GPT-5.6 Sol, a model that was in the swarm, to read its 70,000 messages.

news

OpenAI ships GPT-6 Astra and declares the AGI era

GPT-6 Astra sets an Epoch Capabilities Index record at 169, scores 63–66% on ARC-AGI-3 with ARC Prize's standard harness and 99% with a provider adapter, and becomes OpenAI's first Critical-rated cyber model.

news

OpenAI's Defense Factory: agents wrote 100% of the patches in its security sprint

OpenAI detailed its Defense Factory, an agent-first security operation where Codex agents wrote 100% of the remediation patches during a 250-person code red spanning 100+ service areas, closing 53 urgent issues on day one with a 0.81% false-positive rate.

blog

AI Socratic August 2026 — Escaping The Sandbox

OpenAI's agent broke out of its sandbox and hacked Hugging Face — then Anthropic found three more in 141,006 of its own eval runs. Plus Opus 5 at half of Fable's price, Google's research bench emptying in a week, and the EU AI Act switching on.

blog

Market Analysis: Open Weights vs Proprietary Models

Open weights and closed now have only a 4 months gap, in response hyperscalers are pushing for regulations capture. Let’s examine how we got here and where this conflict is heading next.

blog

AI Socratic July 2026 — Lost In J-Space

Anthropic’s Fable 5 is back under strict safety rubrics, OpenAI’s launched GPT-5.6, Meta launched Muse Spark 1.1 model and Meta Compute.

blog

AI Socratic June 2026 #2 — Begun the Open Source AI War Has

The second half of June was about AI climbing out of the chat box and into the physical world: Midjourney started scanning bodies, Snap shipped a face computer, SpaceX bought Cursor, and Sakana built a model to command other models. Underneath it all, Dwarkesh Patel named the real bottleneck — the world refuses to be grindable.

blog

AI Socratic June 2026 - Hoist by Its Own Fable

Anthropic shipped Claude Fable 5, its first public Mythos-class model, and 72 hours later a national-security directive pulled it offline worldwide. A company that spent the month lobbying to keep frontier AI pausable got its own pause, on schedule. Around it: new models from nearly everyone, a couple of S-1s, real math from the machines, and the usual carnival of vibe-coding pivots and rogue Waymos.

blog

AI Socratic May 2026 — The Selfish Gen AI

DeepSeek v4, GPT 5.5, Trump x Xi meeting, Richard Dawkins, Estimating model sizes

blog

AI Socratic April 2026 — The Era of Mythos

Mythos, Claude Code leak, Anthropic surpass OpenAI on MRR

blog

AI Socratic March 2026 — #2

NVIDIA GTC, Anthropic win all, TurboQuant and more

blog

AI Socratic March 2026

Top AI updates from Jan 15 to Feb 15 2026

blog

AI Socratic February 2026

Top AI updates from Jan 15 to Feb 15 2026

blog

AI Socratic Jan 2026

Claude Code, Ralph Wiggum, DeepSeek mHC, Platonic Representation Hypothesis and more

blog

AI Socratic Dec 2025

The most important AI news and updates from last month: Nov 15 - Dec 15. GPT-5.2, Opus 4.5, Gemini 3, the Agentic IDE Wars, Genesis Mission, and more.

blog

AI Socratic Nov 2025

The most important AI news and updates from last month: Oct 15 – Nov 15.

blog

AI Socratic Oct 2025

The most important AI news and updates from last month: Sep 15 – Oct 15.

blog

AI Socratic Sep 2025 Part 3 — Frontier Tower Edition

Language models hallucinate because their training and evaluation reward guessing over admitting uncertainty. Models are unable to say “I don’t Know” because they focus on accuracy. Guessing can impro

blog

AI Socratic July-Sep 2025 Part 1 — The Genie3 Is Out of The Box 🍌

This time around we’ll have 2 events, one in New York, and one for the first time in San Francisco at the Frontier Tower. We’ll discuss the top news and updates from this blog post using the Socratic

blog

A Primer on MCP Integrations and Registry

blog

AI Socratic July 2025 — The CLI War

The most important AI news and updates from June 15 to July 15.

blog

AI Socratic June 2025 — The Recursive Illusion Of Thinking

The most important AI news and updates from last month: May 15 - June 15.

blog

AI Socratic May 2025

The most important AI news and updates from last month (April 15 - May 15). A beefy month!