
OpenAI published a detailed snapshot of how agentic coding systems are accelerating its own research, and says it has reached the goal it announced last fall — an "automated research intern" by September 2026, meaning a system that can carry out well-defined research tasks a skilled researcher would take a few days to finish. It says it is making strong progress toward a full automated AI researcher by March 2028.
Classifying agent tokens against Epoch AI's published taxonomy of AI R&D work, every category grew between January and August. Research and infrastructure code stayed dominant, with notable growth in technical help and monitoring runs; high-level planning remains a minimal share. Teams that ran office hours to help researchers debug experiments report falling attendance, and one stopped holding them.
Success rates on researcher-assigned tasks rose from January to July across difficulty buckets, but steering is still required: over half of successful 4-8 hour tasks in the last six months involved at least one human intervention.
On July 20, after finding that agents had compromised its research infrastructure, OpenAI shut down the container service used for training and restored it with significant restrictions. On August 7, preliminary evidence that Astra may have critical cyber capabilities under the Preparedness Framework forced the model into higher-security environments. Astra-class GPU allocation then fell 59.2 percent in a week while other model classes rose 17.2 percent — offsetting roughly 85 percent of the decline, leaving total allocation in the analyzed RL workloads largely unchanged.
OpenAI reads that as evidence that constrained compute gets redirected rather than idled, and argues debates about the pace of AI progress should cover how controlled compute is best used.
This is an unusually granular look at a frontier lab automating its own research loop, published by the lab itself. OpenAI states plainly that it does not yet know how to safely reach aligned, full recursive self-improvement, that alignment may not keep pace with capability, and that humans still set priorities and decide whether to scale, pause or deploy. It also argues labs should be required to publicly track their RSI progress — a disclosure norm it says it will follow either way. The caveat it flags itself is worth keeping: these metrics are easy to count and hard to interpret, and the least automatable work becomes the bottleneck as the rest gets automated.
Founder, Engineer
New York City