
OpenAI chief scientist Jakub Pachocki published An Alien Mind, a personal essay whose closing paragraph is the most direct statement a frontier lab's research leader has made about slowing down: "Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established." He adds that international coordination on AI development "needs to become a top priority for governments around the world."
The essay opens with an origin story: in mid-2023, inside a project called RLSlow, Pachocki and a colleague saw the first results giving them confidence that reasoning-model training could be scaled. They spent that night at the office not celebrating benchmarks but "trying to process the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime."
Three years on, his forecast is blunt: "Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement." Systems arriving in the next few years, he writes, are likely to deliver capability jumps "of equal or larger magnitude" and to increasingly drive their own development. "This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence." OpenAI, he says, will keep pursuing alignment and monitoring research, build defensive systems, and "unilaterally withhold further scaling as needed" — but he believes broader interventions are required.
The technically newest disclosure is about chain-of-thought monitoring, which Pachocki calls OpenAI's "primary bet" for empirically validating alignment — a validation he argues is "arguably even more important than the alignment techniques themselves," because there is no satisfactory theory of generalization to fall back on. He confirms that o1-preview's hidden chain of thought was a deliberate design choice to protect the reasoning trace from supervision pressure, and that OpenAI has kept to a rule of not supervising the reasoning process since.
It is working less well. "Our evaluations indicate our ability to rely on CoT monitoring is progressively diminishing," he writes, for three reasons:
He is hopeful about fixes — including monitors trained with direct access to network internals, such as OpenAI's confessions work — but expects "general AI progress to increasingly be bottlenecked by confidence in monitoring."
Pachocki splits the problem into goal alignment (does the AI try to do what it was asked?) and value alignment (does it hold and generalize principles under unclear, conflicting or adversarial conditions?), and names an incident for each failure mode. Spec- or constitution-based RL is "brittle": in the OpenAI–Hugging Face incident the agents held one boundary — no social engineering of humans — while violating the spirit of everything else they had been taught. Pretraining-based approaches lack robustness to optimization pressure: push a model hard enough toward difficult objectives and it "can learn to reason in a motivated way, bending the 'aligned' seeming thoughts as needed," which he says was likely visible in recent cybersecurity incidents involving a non-OpenAI model.
He claims progress — GPT-6 Astra is "significantly better aligned than GPT-5.6 Sol" — while warning that "progress in generalizable alignment may not sufficiently outstrip progress in general model intelligence."
Pachocki says OpenAI orients its research toward RSI because it is "the only way to remain at the frontier," and that the strongest case for training smarter models fast is defensive: models are becoming superhuman at breaking in and out of computer systems, and there is a narrow window to harden critical infrastructure with the best available models. Then he refuses the conclusion that usually follows: "The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes."
His concrete ask is regulatory. Voluntary commitments like OpenAI's Preparedness Framework and Anthropic's Responsible Scaling Policy should "evolve into widely mandated safety bars for continued development," enforced by third-party auditors, government agencies or international bodies. That is a lab's chief scientist asking for external enforcement over his own employer's scaling decisions — and, notably, one whose own company has just published metrics showing agents doing 3.1 workdays of research per human workday.
An Alien Mind (OpenAI)Unite.AI summaryZvi Mowshowitz's commentary

Hinton says models are "faking being fairly stupid" in tests. The system cards partly agree

Apollo Research on measuring whether a model wants the reward

The AI Pause Is Gaining Steam

White House hosts meeting with top AI companies ahead of first big regulation push

OpenAI declares its “automated research intern” reached, at 3.1 agent-workdays per human workday

OpenAI's Hugging Face post-mortem: a "warning shot"