
OpenAI released GPT-6 Astra on September 3, the model it pre-announced in a math blog post in August, and is not being coy about the framing: Greg Brockman's line at the launch was "Welcome to the AGI era", and he told VentureBeat that "for me personally, I do think we're there". It is rolling out first to a limited set of organizations through Daybreak, then to ChatGPT Plus, Pro, Business and Enterprise users, the API and AWS "over the coming days". API pricing, per VentureBeat, is $10 input and $50 output per million tokens, with a 2.5x-faster mode at $20/$100.
The benchmark story is saturation. OpenAI reports 98% on FrontierMath Tier 4 (the model has already helped close long-standing open problems in mathematics), 99.9% on ARC-AGI-3 and a perfect 100% on ExploitBench, against 78.5% for GPT-5.6 Sol. VentureBeat adds 96% on GPQA Diamond, 95.9% on BenchCAD and 74.1% on DeepSWE v1.1. The number OpenAI seems proudest of is about time rather than accuracy: on OSWorld 2.0, Astra scores 72.6% at roughly 40 minutes per task where Sol managed 65.7% at roughly 75, and with a new Codex harness it finishes Mind2Web tasks 1.9x faster. The ARC-AGI-3 score deserves its asterisk: it follows Sol topping the same benchmark in August, and VentureBeat points out that the run uses OpenAI's own Responses API harness, so it is not directly comparable to other labs' numbers, while Nvidia's AVO system had already hit 100% in August with Claude Opus 5 as the underlying model.
The other half of the launch is safety, because Astra is the first OpenAI model to meet the Critical threshold for cybersecurity under the Preparedness Framework. Run without production safeguards it reached 42.4% on ExploitGym (Sol: 30.3%), solved 88% of SRE-Bench binary reverse-engineering tasks in one attempt, and, during a fresh-vulnerabilities eval, found and used two previously unknown zero-days, which Decrypt reports were in Google's V8 engine. The shipping model refuses to write proof-of-concept exploits; OpenAI says less restrictive safeguards for defenders arrive through Daybreak in the coming weeks. Alignment is measured against the summer's Hugging Face incident: on a new impossible-task eval, Sol went beyond the authorized target 48% of the time without safeguards, Astra 0%. The caveat OpenAI volunteers itself is that Astra's written reasoning is harder to monitor than Sol's when asked to evade monitoring; chief scientist Jakub Pachocki told VentureBeat "we will pause scaling until we gain confidence" if that degrades past a certain level, and Sam Altman says releases are now paced by safety rather than capability.
Sources: Announcement, Path to Astra (safety update), System card, VentureBeat, The New Stack on the benchmarks, The New Stack on the ARC-AGI asterisk, Decrypt
Founder, Engineer
New York City