Timeline

Dwarkesh Patel's second interview with Ilya Sutskever declares the scaling era over

Sutskever said models 'generalize dramatically worse than people,' citing an example of an AI that fixes a bug, breaks it again, then reverts to the original error when corrected.

  • Ideas & essays
  • Major

Ilya Sutskever, the OpenAI co-founder who left to start Safe Superintelligence, gave a second long interview to Dwarkesh Patel in which he argued that the era of scaling pre-training compute — roughly 2017 to 2025, and the strategy his own earlier work had helped establish — was ending, and that the field was moving back into “the age of research,” where progress again depends on new ideas rather than larger runs.

His central claim was about generalisation rather than raw capability: current models, he said, “generalize dramatically worse than people,” despite scoring well on benchmarks. He illustrated this with an example of a coding model that fixes a bug, introduces a new one while doing so, is told to fix that, and then reintroduces the original bug — a loop he presented as evidence that strong evaluation scores can coexist with brittle, non-robust understanding. He distinguished this from sample efficiency (needing more examples than a human to learn something), framing the deeper problem as an inability to learn robustly from limited feedback the way people do. He said he had ideas about addressing this involving value functions but declined to elaborate, citing his company’s competitive position.

On timelines, Sutskever resisted a single dramatic breakthrough model, instead describing a path via continual learning: multiple specialised AI systems deployed into the economy, each improving through real-world use and sharing what they learn, rather than one centralised system built and then deployed complete. He offered a loose estimate of five to twenty years before systems could learn at anything like human efficiency and gave the image of a “superintelligent 15-year-old” — capable but still needing to learn on the job — rather than a finished superintelligence.

The interview was widely discussed because it came from someone with an unusually strong claim to have shaped the scaling paradigm, arguing publicly that it was reaching diminishing returns just as rivals were still racing to build ever-larger training runs, and it sharpened an existing disagreement inside the field over whether the next capability gains would come from compute or from new research.