The Godmother of AI on jobs, robots, and why world models are next | Dr. Fei-Fei Li
Key insights
Books referenced
- The Worlds I See - Fei-Fei Li - Referenced only as 'my book' when Fei-Fei says she wrote about the many researchers who inspired her; not named on air, but this is her 2023 memoir covering her ImageNet years.
Media referenced
- TED Talk on spatial intelligence and world models - other - Fei-Fei says she gave this talk in 2024, the point at which she first publicly formulated the world-models idea she had been developing since 2022 from her robotics and computer-vision research.
- New York Times op-ed on human-centered AI - article - Fei-Fei wrote this in 2018 while finishing her Google sabbatical, arguing AI needed a guiding framework anchored in human benevolence; it fed directly into founding Stanford HAI.
- Demis Hassabis interview on AGI and Newton - podcast - Lenny cites an unnamed recent Hassabis interview proposing a test for AGI: give a model all pre-1900s data and see if it can derive Einstein's breakthroughs; Fei-Fei extends the thought experiment to Newton and says today's AI still cannot do it.
Companies
- World Labs - Fei-Fei's ~18-month-old company (founded with Justin Johnson, Christoph Lassner, and Ben Mildenhall) building frontier spatial-intelligence/world models; ~30 people, mostly researchers and research engineers.
- Google Cloud - Fei-Fei was chief AI scientist there; she describes it as a place where major early AI breakthroughs emerged and where she worked alongside Jeff Dean and Jeff Hinton.
- Stanford HAI (Human-Centered AI Institute) - Fei-Fei co-founded it in 2018 with John Etchemendy, James Landay, and Chris Manning; it grew into the world's largest human-centered AI institute, spanning research, policy, and a national AI research cloud bill.
- Nvidia - Supplied the two consumer gaming GPUs the Toronto team used to train the 2012 ImageNet-winning neural network (AlexNet); scaled to today's massive GPU fleets.
- Scale AI - Founder Alexander Wang emailed Fei-Fei in Scale's early days describing how ImageNet inspired the company; cited as one of today's largest data-labeling companies for frontier labs.
- Sony - A virtual-production company collaborated with Sony using Marble-generated 3D scenes to shoot a video, cutting production time by roughly 40x versus traditional methods.
- Twitter (X) - Mentioned in the host's intro as a company where Fei-Fei served on the board.
- Figma - Podcast sponsor promoting Figma Make, its prompt-to-app vibe-coding tool.
- JustWorks - Podcast sponsor; international payroll/HR/benefits platform.
- Cinch - Podcast sponsor; customer communications API (RCS messaging).
Techniques and frameworks
- World models / spatial intelligence - Fei-Fei's framing for AI that can create, reason about, interact with, and navigate 3D/4D worlds from a prompt, as distinct from language models or flat video generation; she calls it the linchpin connecting embodied AI (robotics), visual intelligence, and language.
- The bitter lesson - Richard Sutton's thesis that simpler models plus more data beat complex models with less data; Fei-Fei calls it a 'sweet lesson' that validated her ImageNet bet, but argues it applies less cleanly to robotics because training data (3D actions) doesn't naturally exist the way text tokens do for language models.
- Prompt-to-world generation - Marble's core interaction: a sentence or image input generates an explorable, interactable 3D world, with an intentional 'dots rendering before textures' visualization added purely to make the generation process feel legible and delightful to users.
Summary
Lenny Rachitsky interviews Dr. Fei-Fei Li, the Stanford computer scientist and World Labs founder credited with sparking the current AI boom through ImageNet, on the day her company launches Marble, the first publicly available "world model" product. Much of the conversation traces the history Fei-Fei lived through firsthand: the AI winter of the 2000s, her realization around 2006 that AI's missing ingredient was internet-scale labeled data rather than better models, the resulting 15-million-image ImageNet dataset, and the 2012 moment when a Toronto team led by Geoffrey Hinton combined that data with two consumer GPUs to produce the first neural network to meaningfully crack object recognition. She draws a direct line from that "big data plus neural network plus GPU" recipe to ChatGPT, arguing the ingredients are unchanged, just scaled up by orders of magnitude.
The episode's technical core is Fei-Fei's case for world models as the next necessary leap beyond language models. She defines spatial intelligence as the ability to create, reason within, and interact with 3D and 4D worlds, distinguishing it sharply from passive video generation using Plato's cave allegory. Marble, built by a ~30-person team over roughly a year, lets users prompt a sentence or image into an explorable, interactable 3D world; a deliberately added "dots render before textures" visual effect turned out to delight users by making the generation process feel legible. Early use cases surfaced through user outreach rather than planning: a Sony-affiliated virtual-production shoot cut production time roughly 40x using Marble-generated camera-aligned sets, robotics researchers want it for synthetic training environments, game developers are exporting Marble meshes, and a psychology research team reached out to use it for exposure-therapy-style immersive environments.
Fei-Fei is candid about where current AI still falls short. She argues Richard Sutton's "bitter lesson" (simpler models plus more data wins) does not transfer cleanly to robotics, because robots need action data in 3D space as training input, and no natural internet-scale dataset of that form exists the way text does for language models - researchers are patching the gap with web video, teleoperation, and synthetic data. She also compares robots to self-driving cars: physical systems requiring two decades of maturation (citing Stanford's 2005 DARPA Grand Challenge as the starting point for the self-driving journey Waymo continues today), and notes that unlike cars, general-purpose robots must operate in 3D and deliberately touch objects, not just avoid them. On AGI itself, she is skeptical of the term as scientifically meaningful and offers concrete evidence current models fall short: they cannot reliably count objects across a video the way a toddler can, and cannot derive physical laws from raw data the way Newton or Einstein did, even given more data than those scientists had.
Beyond the technical narrative, Fei-Fei repeatedly returns to a humanist framing she says is underrepresented in Silicon Valley discourse: AI is not something happening to people, it is created by and impacts people, and its trajectory is a matter of collective responsibility, not inevitability. She describes founding Stanford's Human-Centered AI Institute in 2018 - after a 2018 New York Times op-ed argued for a human-benevolence-anchored governance framework - as a direct extension of this belief, and closes the episode answering a question she says she gets constantly while traveling: whether people in any profession (musician, teacher, nurse, farmer, accountant) have a role in an AI-saturated future. Her answer is an unqualified yes, paired with an insistence that human dignity and agency must remain central to how AI is built, deployed, and governed.
On her own career and World Labs specifically, Fei-Fei attributes her path - from a near-tenure position at Princeton to Stanford SAIL director, to Google Cloud chief AI scientist, to founding World Labs 18 months before this interview - to intellectual fearlessness and prioritizing mission and people over risk-minimization, advice she extends directly to young AI researchers weighing job offers today. She is candid that even with decades of research and institution-building experience, nothing fully prepared her for the intensity of talent competition and compensation escalation she has encountered building a frontier lab in the current AI landscape.
Notable Quotes
"I'm not a utopian... In fact, I'm a humanist. I believe that whatever AI does currently or in the future is up to us. It's up to the people." - Dr. Fei-Fei Li
"There's nothing artificial about AI. It's inspired by people. It's created by people. And most importantly, it impacts people." - Dr. Fei-Fei Li
"Robot is 3D things running in 3D world and the goal is to touch things." - Dr. Fei-Fei Li, contrasting robots with self-driving cars
"No technology should take away human dignity. And human dignity and agency should be at the heart of the development, the deployment, as well as the governance of every technology." - Dr. Fei-Fei Li
"I don't overthink of all possible things that can go wrong because that's too many. I feel like that's an important element - is not focusing on the downside, focusing more on the people, the mission, what gets you excited." - Dr. Fei-Fei Li