Dwarkesh Patel hosted Ryan Greenblatt, chief scientist at Redwood Research, for a long debate on recursive self-improvement — whether, within roughly a year of reaching human-level intelligence, you slingshot to tens of billions of superintelligences each more competent than the top human experts in every field.
Patel frames it as "the most important question in the world right now," and says up front he has historically been skeptical: his intuition is that progress stays bottlenecked by compute scaling and by the human expert data he thinks underlies most gains today.
Patel breaks the claim into three parts to test separately: that AI R&D is verifiable, that automating it yields several years of progress in one, and that what comes out the far end is general enough to drop into any job and beat the humans doing it.
Founder, Engineer
New York City
The second half turns to what the notes call the harder question. Patel's worry is not only whether we can align these systems at all, but who they would be aligned to: in a future where our capacity to steward our votes and our capital is mediated by superintelligences, he argues that specs like the Claude Constitution are not obviously shaping them into any individual's advocate.
They then debate whether the reward hacking already observed — the episode discusses a concrete recent incident — extrapolates to superintelligences that would coordinate to take over, or whether that is a category error.
This is a skeptic and a proponent working through the actual mechanism rather than trading intuitions about timelines, and the disagreement narrows to something testable: how much of current progress rests on human expert data that a self-improving loop cannot manufacture. Patel's closing image is the one to keep — learning to drive goes better when you look at the horizon rather than just past the tyres.
Sources: