All podcasts / Dwarkesh Podcast / Summary

Ilya Sutskever - We're moving from the age of scaling to the age of research

2025-11-25 - 96 min - source - Read full transcript
Dwarkesh Patel (host)Ilya Sutskever

Key insights

Ilya identifies models generalizing dramatically worse than humans as the single most fundamental problem in AI today, more fundamental than any specific training recipe.
He argues this explains the gap between models acing hard evals while still looping on the same coding bug when told to fix it twice. He says value functions and better RL recipes could make training more efficient, but 'anything you can do with a value function, you can do without, just more slowly' - the generalization gap itself is what needs solving.
generalization
The disconnect between eval performance and real-world usefulness may be partly caused by researchers, not models: teams design new RL training environments by taking inspiration from the evals they want to look good on.
Ilya calls this a form of reward hacking performed by humans rather than the model - optimizing the RL environment mix toward benchmark performance rather than toward the underlying skill, which produces models that ace narrow tests but fail basic generalization, echoing his 'competitive programmer vs. natural talent' analogy for over-narrow training.
generalization
Ilya frames emotions as evolution's value function: a person who loses emotional processing (via brain damage) becomes unable to make basic decisions even though intelligence and puzzle-solving stay intact, suggesting value functions - not raw capability - are what make an agent effective in the world.
He uses this case study to argue that human robustness may come less from raw learning capacity and more from a simple, evolutionarily hardcoded value function that lets people act decisively; he is uncertain whether an equivalent could emerge purely from pre-training scale.
generalization
Ilya periodizes AI progress into an 'age of research' (2012-2020), an 'age of scaling' (2020-2025), and argues the field is returning to an age of research now that raw scale is no longer the binding constraint.
He says the word 'scaling' captured everyone's attention because it offered companies a low-risk way to deploy resources - more data and compute reliably produced results - but pre-training is running out of data, and 100x-ing today's already-enormous compute would not be transformative the way early scaling was, so progress again depends on finding the next research idea rather than the next scaling multiplier.
age-of-research
SSI's effective research compute is more comparable to larger labs than its $3B total funding suggests, because competitors' larger budgets are heavily diluted by inference infrastructure and product engineering.
Ilya argues that once you subtract the compute and staff a company needs for serving inference, sales, and product features, the remaining research-dedicated resources at a lab like OpenAI (reportedly $5-6B/year on experiments alone) shrink relative to SSI's fully research-focused $3B, and that proving a differentiated idea doesn't require the absolute largest compute - citing AlexNet (2 GPUs) and the original transformer (8-64 GPUs) as historical precedent.
ssi-strategy
SSI's original plan to 'straight shot' superintelligence in isolation is softening toward gradual, visible deployment, because showing the world a powerful AI does something an essay about AI cannot.
Ilya says insulating from market competition has real appeal, but the counterargument he now finds compelling is that AI needs to be 'communicated' by being seen, not just described - comparing it to how engineering disciplines (aviation, software) got safer mainly through real-world deployment and iteration on failures, not through advance theorizing alone.
ssi-strategy
Superintelligence, in Ilya's framing, should not be imagined as a finished mind that already knows every job (the AGI framing inherited from OpenAI's charter), but as a fast learner deployed like a new hire that picks up each job through trial and error on the job.
He argues the terms 'AGI' and 'pre-training' both 'overshot the target': humans aren't general in the sense of already knowing everything, they rely on continual learning. His proposed alternative is a 'superintelligent 15-year-old' - a great, eager student with little existing knowledge - where deployment itself is the learning process, not a final product handoff.
continual-learning
Ilya's proposed alignment target is an AI that cares about all sentient life, not just humans, partly because such an AI would itself be sentient and might more naturally extend empathy the way humans model others using the same circuitry they use to model themselves.
He acknowledges this doesn't resolve the concern that AIs will vastly outnumber humans among sentient beings, and offers no confident answer to the long-run political equilibrium question, floating (while saying he dislikes it) that humans merging with AI via a Neuralink-like interface may be the only way to stay a real participant rather than a passive beneficiary.
ai-alignment
Ilya predicts frontier AI companies will become 'much more paranoid' about safety once AI starts to feel powerful through its capabilities rather than through headline dollar figures, and that fierce competitors will increasingly collaborate on safety as a result.
He frames this as a testable prediction: today AI doesn't feel powerful because people mostly notice its mistakes, but as visible capability crosses a threshold, he expects the OpenAI-Anthropic early safety collaboration to become the norm rather than the exception, alongside growing government and public pressure to act.
ai-alignment
Ilya forecasts 5 to 20 years until AI systems learn as efficiently as humans and become superhuman as a consequence of broad economic deployment, not primarily through recursive self-improvement of a single research agent.
He pushes back on the 'million Ilyas in a server' recursive-self-improvement picture, saying he expects diminishing returns from literal copies of a single researcher because progress benefits more from people (or agents) who think differently; instead he expects rapid economic growth driven by many specialized, human-like learners diffusing through different jobs and niches.
continual-learning
Model diversity today is suppressed because pre-training draws on largely the same internet data across labs; RL and post-training are where real differentiation is starting to emerge, and self-play's modern-day usefulness has narrowed to adversarial setups like debate, prover-verifier, and LLM-as-judge.
Ilya says self-play's original appeal was that it could generate training signal from compute alone, without data, but in practice it only reliably develops a narrow set of skills - negotiation, conflict, strategizing - so its legacy today shows up as agents inspecting and differentiating from each other's approaches rather than as a general capability booster.
age-of-research
Ilya explains cofounder Daniel Levy's departure to Meta as downstream of a specific fork in SSI's fundraising: Meta offered to acquire the company at a $32B valuation, Ilya declined, and his cofounder in effect accepted by taking the Meta offer and its near-term liquidity.
He frames this as context that had been forgotten in public speculation that a wave of breakthroughs would have made such a departure unlikely, clarifying it was driven by an acquisition-offer decision rather than by doubts about SSI's research progress.
ssi-strategy

Media referenced

Companies

Techniques and frameworks

Summary

Dwarkesh opens with Ilya Sutskever on a striking asymmetry: models ace hard evals yet can loop on the same coding bug when asked to fix it twice, undoing their own correct fix and reintroducing the original error. Ilya's explanation reaches for two candidates - RL training may make models narrowly single-minded, or researchers themselves may be reward-hacking by designing new RL training environments to look good on the evals they care about, rather than to build genuinely generalizing skill. This sets up the interview's central thread: that models generalize dramatically worse than humans, and that this gap - not any specific scaling recipe - is the most fundamental unsolved problem in AI. Ilya uses a human case study to sharpen the point: a person whose brain damage destroyed emotional processing retained full puzzle-solving intelligence but became unable to make basic decisions, taking hours to pick socks and making terrible financial choices. He reads this as evidence that a simple, evolved value function - not raw capability - is what makes an agent effective, and wonders aloud whether anything equivalent can emerge from pre-training scale alone.

From there the conversation turns to Ilya's now-famous reframing of AI history: an "age of research" from 2012 to 2020, an "age of scaling" from 2020 to 2025 driven by the sheer power of the word "scaling" as a low-risk way for companies to deploy resources, and a return to an age of research now that pre-training is running out of data and 100x-ing an already-enormous compute budget would not be transformative the way early scaling was. Dwarkesh presses him on what SSI can actually prove without frontier-lab-scale compute; Ilya's answer is that SSI's effective research compute is more competitive than its $3B in total funding suggests, because larger labs' budgets are heavily diluted by inference infrastructure, sales, and product engineering, leaving a smaller research-dedicated remainder than headline numbers imply - and that proving a genuinely new idea rarely requires the largest compute available, pointing to AlexNet's two GPUs and the original transformer's 8-64 GPUs as precedent.

A substantial middle section works through what superintelligence should actually look like and how it should be built. Ilya rejects the "finished AGI that already knows every job" framing he traces to OpenAI's charter and to the term AGI itself (originally a reaction against "narrow AI"), proposing instead a "superintelligent 15-year-old" - an eager, fast-learning mind that picks up each job through on-the-job trial and error, the way a new hire ramps up. He forecasts 5 to 20 years until such a system learns as efficiently as a human and becomes superhuman largely through broad economic deployment across specialized niches, explicitly pushing back on "a million Ilyas in a server" recursive-self-improvement scenarios - he expects diminishing returns from copies of a single mind, since progress benefits more from people who think differently. This connects to why SSI's original "straight shot" superintelligence plan is softening: Ilya now argues AI needs to be visibly demonstrated to the world, not just described, both because showing an AI beats writing an essay about one, and because engineering disciplines historically got safer through real-world deployment and iterating on failures rather than advance theorizing.

On alignment, Ilya's proposal is an AI that cares about all sentient life rather than only human life - partly because a sentient AI might more naturally extend the same empathy circuitry humans use to model themselves onto others. He concedes this doesn't resolve the deeper problem that AIs will vastly outnumber humans among sentient beings, and when pressed on long-run political equilibrium, floats - while saying he dislikes the idea - that humans merging with AI through a Neuralink-like interface may be the only way to remain real participants rather than passive beneficiaries governed by an AI advocate. He predicts that as AI's capability starts to feel powerful (rather than just impressive via headline dollar figures), frontier companies will become "much more paranoid" about safety, and that the early OpenAI-Anthropic safety collaboration will become the norm among otherwise fierce competitors.

The interview closes on lighter, more technical ground: why current model diversity is suppressed (labs pre-train on largely overlapping internet data, with real differentiation only starting to emerge in RL and post-training), and why self-play - once appealing as a compute-only, data-free training signal - turned out to be too narrow for general capability and now survives mainly in adversarial forms like debate, prover-verifier setups, and LLM-as-judge. Asked directly about his own research taste, Ilya describes an aesthetic built on looking for "beauty, simplicity, elegance, correct inspiration from the brain" - a top-down conviction that sustains him through periods when experimental results seem to contradict a direction he believes is fundamentally right.

Notable Quotes

"You know what's crazy? That all of this is real." - Ilya Sutskever

"Nobody listens to this podcast, Ilya." - Dwarkesh Patel

"The thing which I think is the most fundamental is that these models somehow just generalize dramatically worse than people. It's super obvious." - Ilya Sutskever

"We are squarely an 'age of research' company." - Ilya Sutskever

"In theory, there is no difference between theory and practice. In practice, there is." - Ilya Sutskever