OpenAI's head of platform engineering on the next 12-24 months of AI | Sherwin Wu
Key insights
Books referenced
- There Is No Antimemetics Division - qntm - Wu's top fiction recommendation, read in two days; sci-fi about a government agency fighting things that erase memory of themselves.
- Breakneck - Dan Wang - Nonfiction pick on US-China dynamics; Wu cites its framing of the US as a 'lawyerly society' versus China as an 'engineering society.'
- Apple in China - Patrick McGee - Second nonfiction pick, praised for inside detail on Apple's manufacturing relationship with China.
- Structure and Interpretation of Computer Programs - Harold Abelson and Gerald Jay Sussman - College textbook Wu calls 'the wizard book'; its 1980 metaphor of programmers as sorcerers issuing incantations frames his whole discussion of AI coding agents.
- The Mythical Man-Month - Fred Brooks - Source of the 'software engineer as surgeon' analogy Wu adapts as his management philosophy for supporting top performers.
Media referenced
- Fantasia (The Sorcerer's Apprentice) - movie - Wu's analogy for ungoverned AI agents: Mickey sets the brooms loose and everything floods before the sorcerer returns to clean up.
- Her - movie - Referenced in passing when discussing AI agents conversing with each other, which Wu says feels like 'her happening in real life.'
- Jujutsu Kaisen (season 3) - show - Recent anime Wu watched in his limited free time; he says anime creates plots and universes Western media shies away from.
Companies
- OpenAI - Wu's employer; he leads engineering for the API and developer platform.
- Codex - OpenAI's coding agent; used by 95% of OpenAI engineers daily and reviews 100% of internal PRs.
- ChatGPT - OpenAI's consumer product, cited at roughly 800 million weekly active users and central to OpenAI's 'benefit all of humanity' mission.
- Cursor - Cited as an example of a startup that succeeded in a crowded AI-coding space despite OpenAI and other large labs also operating there.
- Quora - Wu's first job out of college, where he owned newsfeed code review and hated the volume of manual PR review.
- DX (getdx.com) - Episode sponsor; developer intelligence platform used by companies like Dropbox and Booking.com to measure AI's engineering impact.
- Sentry - Episode sponsor; application monitoring and AI debugging agent (Seer) that traces errors to commits and opens fix PRs.
- Datadog - Episode sponsor; home of the Eppo experimentation and feature-flagging platform for connecting product analytics to business impact.
- Airbnb - Referenced by the host as the origin of the team that later built Datadog's Eppo experimentation product.
- Ubiquiti - Home networking and security-camera brand Wu recently adopted and calls 'the Apple of home networking' due to its app quality.
- Zillow - Referenced for a book/study on home resale value, specifically that front-door replacement has the highest ROI of home improvements.
- Opendoor - Wu's previous employer, where he built pricing models for how much the company should pay for houses.
- FinTool - AI startup building agents for financial services; its founder Nicholas coined the phrase 'the models will eat your scaffolding for breakfast,' which Wu quotes.
Techniques and frameworks
- The bitter lesson (applied to AI product scaffolding) - Wu draws a direct parallel between the ML principle that more compute beats hand-built structure, and how AI product scaffolding (vector stores, agent frameworks) keeps getting superseded as base models improve.
- Build for where the models are going, not where they are today - Wu's standing advice to AI-product builders and his stated reason to sometimes discount direct customer feedback, since customers anchor on today's model limitations.
- METR-style task-length benchmark - Benchmark tracking how long an AI model can coherently complete a software-engineering task; Wu cites current frontier models at roughly multi-hour tasks 50% of the time and under an hour 80% of the time.
- Bottom-up 'tiger team' AI adoption - Wu's recommended pattern for enterprise AI rollouts: pair top-down executive buy-in with a dedicated, often non-engineer, enthusiast team that evangelizes and builds best practices internally.
Summary
Sherwin Wu, OpenAI's head of engineering for the API and developer platform, joins Lenny Rachitsky to describe just how far Codex adoption has gone inside OpenAI: 95% of engineers use it daily, virtually every merged PR gets a Codex review, and heavy users open 70% more PRs than everyone else. The bigger shift, he argues, is in the shape of the job itself. Engineers increasingly act as tech leads running 10-20 parallel agent threads rather than typing code, a change he frames through the "wizard" metaphor from the 1980s textbook SICP and the cautionary Sorcerer's Apprentice story: massive leverage, but real risk if you let the agents run unsupervised. An internal OpenAI team maintaining a 100%-Codex-written codebase, with no manual-code fallback, has found that most agent failures trace back to missing context rather than model limits, pointing teams toward encoding more institutional knowledge directly into docs and repo structure.
On management, Wu extends a long-standing practice of spending more than half his time with his top performers, drawing on the Mythical Man-Month's "surgeon" analogy: the manager's job is to clear organizational blockers so the strongest people can operate at maximum leverage, a dynamic he expects to intensify as AI further separates high performers from everyone else. He's more skeptical of blindly following customer feedback in a field this fast-moving, citing his own team's experience watching successive generations of AI scaffolding (prompt tricks, vector stores, now skills files) get "eaten" by improving base models. His standing advice is to build for where models are headed 12-24 months out, not for their present limitations.
Wu is candid that many companies see negative ROI from AI deployments, tracing this to purely top-down mandates that never build genuine bottom-up adoption; the fix he recommends is pairing executive sponsorship with an internal "tiger team," often staffed by enthusiastic non-engineers, that evangelizes and builds best practices. He's also bullish that the biggest AI opportunity outside Silicon Valley may lie in automating repeatable business processes, not open-ended engineering work, precisely because that space gets little attention on X or Twitter.
On platform strategy, Wu describes OpenAI's decision to keep the API open and avoid competing away app-layer businesses as core to its mission of spreading AI's benefits broadly, not just competitive positioning. His advice to founders anxious about being "squashed" by a big lab is that failed startups almost always die from lacking product-market fit, not platform competition. Looking ahead, he expects frontier models to push from roughly hour-long coherent task completion today toward much longer autonomous stretches within 12-18 months, alongside meaningful gains in speech-to-speech audio models, an area he considers underrated relative to text-based coding tools given how much of the world's work happens by talking rather than typing. He closes with a "one-person billion-dollar startup" thought experiment: if a single founder can build something massive, the second-order effect may be a wave of smaller, AI-built businesses selling each other bespoke software, potentially shrinking outlier venture returns even as it multiplies viable small companies.
Notable Quotes
"This is the worst the models will ever be." - Sherwin Wu (quoting OpenAI VP of Science Kevin Weil)
"The models will eat your scaffolding for breakfast." - Sherwin Wu (quoting Nicholas, founder of FinTool)
"Make sure you're building for where the models are going and not where they are today." - Sherwin Wu
"Never feel sorry for yourself." - Sherwin Wu, on his personal life motto