Dylan Patel - Deep dive on the 3 big bottlenecks to scaling AI compute
Key insights
Media referenced
- Dwarkesh's prior podcast interview with Dario Amodei - podcast - Referenced repeatedly - Dwarkesh explains he was arguing that Dario's stated 'two years from a data center of geniuses' timeline is inconsistent with Anthropic's comparatively conservative compute-purchasing posture.
- Dwarkesh's prior podcast interview with Elon Musk - podcast - Musk's space-GPU and TeraFab clean-room claims are the direct jumping-off point for the episode's closing sections on space data centers and Texas fab-building.
- SemiAnalysis newsletter/blog posts on AI power buildout and the memory crunch - article - Dylan references his own prior SemiAnalysis writing throughout - on gas turbine capacity from GE Vernova/Mitsubishi/Siemens, on the coming memory crunch he flagged roughly a year and a half before prices actually moved, and on TSMC's three-year CapEx total.
- The Information's reporting on Anthropic's gross margins - article - Cited as the source for Anthropic's sub-50% gross margin figure, used to back into how much of Anthropic's ~$20B ARR is being spent on compute rental.
Companies
- SemiAnalysis - Dylan Patel's company; sells supply-chain data, spreadsheets, reports, and API access on chip/data-center/power buildout to industry (~60% of revenue) and hedge funds (~40%).
- Nvidia - Discussed as the entity capturing an outsized share of TSMC N3/N2 wafer allocation and PCB/memory supply because it committed to demand earliest and most aggressively, unlike Google/Amazon.
- TSMC - The foundry whose N3 and (increasingly) N2 capacity is being crowded out from mobile/PC toward AI accelerators; discussed at length on Apple's shrinking share of its leading-edge nodes.
- ASML - Maker of the EUV lithography tools identified as the ultimate long-run bottleneck on AI compute scaling; discussed in detail on tool cost, production rate, and supply-chain complexity.
- Anthropic - Central case study for compute-constrained growth - conservative long-term contracting, ~$20B ARR at sub-50% margins, needing 5-6 gigawatts by year-end, and its 'commitment issues' meme relative to OpenAI.
- OpenAI - Contrasted with Anthropic as the more aggressive compute buyer, signing deals across many providers (CoreWeave, Oracle, SoftBank Energy, NScale) and ending the year with more accessible compute.
- Google - Discussed for its TPU pod topology (torus, thousands of chips), its slow 2025 wake-up to AI demand that then caused a scramble for extra TSMC wafer allocation, and its large Gemini Pro model enabled by its unipolar TPU compute fleet.
- Amazon - Discussed for Trainium/Graviton TSMC allocation dynamics, AWS's blended scale-up topology (partway between Nvidia's all-to-all and Google's torus), and its role serving Anthropic via Bedrock.
- Microsoft - Cited as a provider Anthropic increasingly leans on for last-minute/lower-tier compute access, and as the customer for Nebius's ship-engine-powered New Jersey data center.
- Meta - Cited for adding, in a single year, as much data center capacity as its entire 2022 fleet (serving WhatsApp, Instagram, Facebook, and AI combined).
- SK Hynix, Samsung, Micron - The HBM/DRAM memory makers; discussed for underinvesting in new fabs during the low-margin 2023 downturn and now scrambling (Micron buying a Taiwanese lagging-edge fab) to add capacity that won't come fully online until 2027-2028.
- CoreWeave, Oracle, Nebius, NScale, SoftBank Energy - Neoclouds/newer entrants cited as where OpenAI in particular sourced diversified, sometimes shorter-term compute capacity as hyperscaler allocation tightened.
- AMD - Cited for signing on to TSMC's brand-new N2 node in the same window as Apple (a scaling risk/bet) and for its own CPU/GPU chiplet packaging strategy.
- Huawei - Discussed as arguably capable of beating Nvidia's Rubin on raw capability if it had TSMC access; credited with cracked AI chip design, its own fabs, top AI talent, and an end-market, and with being first to ship a 7nm AI chip (Ascend) ahead of the TPU and A100.
- Tesla / SpaceX (Elon Musk's companies) - Discussed for Musk's TeraFab/clean-room ambitions, the abandoned-then-restarted Dojo wafer-scale chip, the Memphis/Colossus power-sourcing decisions, and the choice to fab robot chips on Samsung in Texas for supply-chain and geopolitical diversification from Taiwan.
- Apple - Case study for how AI demand is squeezing out a legacy customer - historically first to a new TSMC node and majority of its volume, now projected to fall to roughly half of N2 volume and eventually 'just any old customer' as TSMC prioritizes AI.
- Applied Materials, Lam Research, Carl Zeiss, Cymer - Other critical-tool and component makers in the EUV/fab supply chain; Zeiss (lenses/optics) and Cymer (EUV light source, ASML-owned) singled out as artisanal, hard-to-scale suppliers with under 1,000 specialized workers.
- GE Vernova, Mitsubishi, Siemens - The three combined-cycle gas turbine manufacturers whose ~60 gigawatts/year of production Dylan says is only one of many power-generation options, not the binding constraint people assumed.
- Bloom Energy - Fuel-cell power maker SemiAnalysis has been bullish on for its ability to scale production quickly even at somewhat higher cost than combined-cycle turbines.
- Nebius, Crusoe, Boom Supersonic - Cited as examples of unconventional behind-the-meter power sourcing (ship engines, aeroderivative jet-engine turbines) for data centers.
- Xiaomi, Oppo - Chinese smartphone makers reported (via SemiAnalysis's Asia analysts) to be cutting low-end/mid-range volumes roughly in half in response to memory price increases.
- DeepSeek, Moonshot AI (Kimi K2.5) - Cited as production inference workloads used to demonstrate that the real-world Hopper-to-Blackwell performance gap (~20x) is far larger than the raw FLOPS difference (~2-3x) once networking and system design are accounted for.
Techniques and frameworks
- Alchian-Allen effect - Economic principle Dwarkesh introduces and Dylan applies to AI: when a fixed cost (rising GPU prices) is added to both a cheaper and a pricier good (a weaker vs. a frontier model), the relative price gap between them shrinks, pushing buyers toward the higher-quality option on the margin.
- GPU total cost of ownership (TCO) model - The mechanical framework SemiAnalysis uses to price compute - data center cost, networking, chip/server cost, spare parts, and depreciation schedule combined to derive an hourly rental rate and implied gross margin for a given GPU generation.
- Dragonfly network topology - The network topology Nvidia, Google, and Amazon are all converging toward for future chip scale-up domains - partly fully-connected, partly not - to get scale-up domains into the hundreds/thousands of chips without full all-to-all contention.
- ClusterMAX - SemiAnalysis's rating system for 40+ neocloud and hyperscaler providers, used to differentiate them primarily on GPU deployment speed and failure-management quality rather than raw specs.
- InferenceX - SemiAnalysis's open-source model that searches the space of inference configurations (batch size, KV cache placement, chip count) to find optimal deployment points across chips and models.
Summary
Dwarkesh Patel interviews his roommate and SemiAnalysis founder Dylan Patel for a wide-ranging deep dive into what actually constrains AI compute scaling. The episode opens on the economics of the current buildout: hundreds of billions in hyperscaler CapEx and lab fundraising (OpenAI's $110B, Anthropic's $30B) is explained not as pure current-year compute spend but as prepayment for multi-year infrastructure - turbine deposits, data center construction, power purchase agreements. Dylan traces why Anthropic, despite deliberately conservative long-term contracting to avoid overcommitting, ended up compute-constrained relative to OpenAI's more aggressive multi-provider strategy, forcing it into costlier last-minute and revenue-share deals. A key throughline here is that an H100 GPU is worth more today than three years ago - counter to the standard depreciation bear case - because rising model quality and token value have outpaced the erosion from faster newer chips, a dynamic Dylan connects to the Alchian-Allen effect: as GPU prices rise, the relative price gap between frontier and lesser models shrinks, pushing buyers toward the best model available.
The conversation's structural core is Dylan's argument that the bottleneck constraining AI compute keeps migrating down the supply chain - from CoWoS packaging, to power, to (this year and next) fab clean-room construction, and ultimately, by the late 2020s, to ASML's EUV lithography tools. He walks through the physical mechanics of EUV in detail (tin-droplet plasma sources, multilayer mirror optics, nine-G reticle stage movement, sub-nanometer overlay accuracy) to explain why production can't simply be scaled up on demand, and calculates that even aggressive tool production growth caps total AI chip output near 200 gigawatts/year by 2030 - a number he says is compatible with, not far below, the most ambitious lab projections. A large middle section covers the parallel memory crunch: HBM's bandwidth advantage over commodity DRAM makes switching to cheaper memory impractical for AI workloads, and the resulting price increases are already showing up as roughly $100-150 in added iPhone bill-of-materials cost and a projected halving of low-end/mid-range smartphone volumes over the next two years.
The episode's most geopolitically charged stretch addresses US-China dynamics: Dylan frames the outcome as timeline-dependent (fast AI progress favors the US's compounding compute/revenue lead; slow progress gives China time to build a fully indigenized supply chain) and argues Huawei, if it had TSMC access, could plausibly out-execute Nvidia given its combination of chip design talent, in-house fabs, and end-market reach. The closing sections turn to alternative scaling paths: Dylan is skeptical of Elon Musk's space-GPU ambitions - not because space power isn't cheap, but because chips remain the binding constraint everywhere, and space adds deployment-delay and unreliable-interconnect costs that Earth-based facilities don't face - while he's comparatively bullish that power and land won't meaningfully constrain the US buildout this decade, citing over a dozen alternative generation technologies and unused grid headroom. The interview ends on humanoid robots (arguing most intelligence should stay centralized in the cloud rather than on-device) and a final risk assessment of Taiwan, where Dylan argues that even successfully evacuating TSMC's engineers wouldn't prevent a catastrophic multi-year setback to global AI compute capacity.
Notable Quotes
"There is some sense that we're in a fast takeoff. It's not like we're talking about a Dyson sphere by X date, it's more like the revenue is compounding at such a rate that it does affect economic growth." - Dylan Patel
"It's a snake eating its own tail meme because you can't make the tools without the chips from Taiwan, which you can't use without the tools in Taiwan." - Dylan Patel
"But Elon doesn't win by doing 20% gains. He never wins that way. Elon wins when he swings for the fences and does 10X gains." - Dylan Patel
"Fifty gigawatts of economic CapEx in the data center, and what gets built on top of that in terms of tokens is even larger. It might be $100 billion worth of AI value into the supply chain, held up by this $1.2 billion worth of tooling that simply cannot expand its supply chain quickly." - Dylan Patel
"You only buy the memory crunch if you believe AI is going to take off in a huge way." - Dylan Patel