All podcasts / Dwarkesh Podcast / Summary

Dylan Patel - Deep dive on the 3 big bottlenecks to scaling AI compute

2026-03-13 - 151 min - source - Read full transcript
Dwarkesh Patel (host)Dylan Patel

Key insights

The ultimate long-run bottleneck on scaling AI compute is EUV lithography tool production from ASML, not power or data centers.
ASML makes roughly 70 EUV tools this year, ~80 next year, and even under aggressive expansion only reaches ~100/year by 2030 - implying roughly 700 cumulative tools by decade's end. At ~3.5 tools needed per gigawatt of AI chip capacity, that caps total addressable AI chip output near 200 gigawatts/year, a number Dylan says is 'completely compatible' with Sam Altman's stated ambition of a gigawatt a week by 2030.
semiconductor-bottlenecks
This year's and next year's binding constraint is physical fab clean-room construction, not the lithography tools themselves.
Fabs take two to three years to build versus under a year for a data center (Amazon has built one in eight months). Dylan is skeptical Elon Musk's 'move fast, it can be dirty' approach transfers to semiconductor fabs, where all air must be replaced roughly every three seconds and particle tolerances are extreme - he thinks Musk could build the clean room shell reasonably fast but not develop the process technology itself.
semiconductor-bottlenecks
ASML has never raised EUV tool prices faster than it has improved the tools' capability, which is unusual restraint for a monopoly supplier and helps explain why no one has successfully arbitraged its scarce capacity.
Tool prices rose from ~$150M to ~$400M while throughput and overlay accuracy more than doubled. Dylan contrasts this with Nvidia and memory vendors, who have taken pricing power aggressively, and argues that attempting to pre-buy and resell EUV allocation (the way some funds did with gas turbine capacity) likely wouldn't work because ASML would refuse to sell to speculative buyers.
semiconductor-bottlenecks
An H100 GPU is worth more today, in dollar terms, than it was three years ago - the opposite of the standard depreciation narrative.
Because model quality and token value have risen faster than newer chips have eroded relative performance, the value an H100 can extract running a modern model (e.g., GPT-5.4, which is cheaper to serve and produces more valuable tokens than GPT-4) now exceeds what it could extract three years ago. This undercuts bearish arguments (e.g., Michael Burry's) that short depreciation cycles make data-center CapEx financially unsound.
compute-pricing-economics
Because most AI compute is locked into multi-year contracts, nearly all new pricing power sits with whoever controls the incrementally added capacity each year, not the existing installed base.
Providers like CoreWeave have over 98% of compute on 3+ year contracts and can't flex pricing on it, but every year's newly added capacity transacts at that year's much higher price. Dylan traces where that margin flows upstream - to chip vendors (Nvidia), memory vendors, and eventually equipment makers like ASML - as each layer secures its own long-term supply commitments.
compute-pricing-economics
Anthropic's historically conservative approach to signing long-term compute contracts left it structurally compute-constrained relative to OpenAI, forcing costlier workarounds.
Dario Amodei stated a deliberate preference for conservative compute commitments to avoid bankruptcy risk if revenue growth stalled. With Anthropic's revenue instead accelerating (~$4-6B added per month), it now has to acquire capacity through revenue-share deals via Bedrock/Vertex/Foundry or short-term/spot deals at a markup Dylan estimates around 50%, while OpenAI - having signed aggressively across many providers including unconventional ones like SoftBank Energy - ends the year with materially more accessible compute.
compute-pricing-economics
The AI-driven memory crunch is a direct, quantifiable tax on consumer electronics.
Dylan estimates the DRAM in a 12GB iPhone went from ~$50 to ~$150 as memory prices roughly tripled, adding roughly $100-150 to Apple's bill of materials once NAND is included - likely passed on as ~$250 in higher consumer price. The effect is worse for low-end/mid-range phones where memory is a larger share of cost and margins are thinner; SemiAnalysis's Asia analysts report Xiaomi and Oppo already cutting low/mid-range volumes roughly in half, with total smartphone shipments projected to fall from ~1.1B toward 500-600M within two years.
memory-crunch
HBM's per-edge-area bandwidth advantage over commodity DRAM (roughly 2.5 TB/s vs. 64-128 GB/s) makes switching AI accelerators to cheaper, higher-capacity DDR memory impractical, because inference workloads are bandwidth-bound, not capacity-bound.
Chip I/O is constrained by the physical edge (shoreline) of the die, and HBM packs roughly an order of magnitude more bandwidth into that same shoreline than DDR. Even though DDR would yield ~4x more bits per wafer, using it would leave a chip's compute mostly idle waiting on memory reads/writes, so the 'obvious' fix of using more abundant commodity memory doesn't actually solve the underlying constraint.
memory-crunch
Whether the US or China 'wins' the AI compute race depends on how fast AI capability timelines are, not on any fixed technological gap.
Fast timelines favor the US: its current compute and revenue lead compounds (Anthropic/OpenAI reinvesting revenue into more capacity) faster than China can build an indigenized supply chain. Slow timelines (out to 2035) favor China, which could build a fully vertically indigenized semiconductor stack instead of the West's more fragmented US/Japan/Korea/Taiwan/Europe supply chain - though as of the interview China still lacks that full indigenization even in DUV, let alone EUV.
us-china-ai-race
Huawei is arguably capable of building a better AI accelerator than Nvidia's Rubin if it had equal access to TSMC's leading-edge process.
Dylan argues Huawei uniquely combines cracked chip design, its own fabs, top-tier AI research talent with a larger concentrated pool in China, its own token/product end-market, and networking expertise - the pieces Nvidia has but Huawei may execute even better on. He estimates that absent the 2019 TSMC ban, Huawei would likely already be TSMC's largest customer, surpassing Apple.
us-china-ai-race
Elon Musk's space-GPU thesis is weakened less by energy economics - power genuinely is near-free in space - than by the same chip-manufacturing bottleneck constraining Earth-based data centers, plus added costs unique to space deployment.
Because chips, not power, are the binding constraint through the 2020s, moving GPUs to space doesn't add compute capacity, it just relocates existing scarce chips. Space adds its own costs: testing, deconstructing, launching, and re-testing chips delays deployment by months (costly when compute is most valuable in its first six months), and unreliable space-laser interlinks (versus mass-produced pluggable optical transceivers) make inter-satellite networking for large sparse-MoE models especially hard.
power-and-scaling-alternatives
Power and land are unlikely to constrain US AI buildout this decade because dozens of alternative generation methods, beyond the three major combined-cycle turbine makers, can each individually scale to tens of gigawatts.
Dylan cites 16+ tracked power-generation vendors (aeroderivative jet-engine turbines, reciprocating diesel-style engines, repurposed ship engines, Bloom Energy fuel cells, solar-plus-battery) with hundreds of gigawatts of combined orders already in SemiAnalysis's data, plus the possibility of unlocking roughly 20% of the US grid's peak-reserved capacity via utility-scale batteries, since that headroom sits idle outside a few peak-demand hours per year.
power-and-scaling-alternatives

Media referenced

Companies

Techniques and frameworks

Summary

Dwarkesh Patel interviews his roommate and SemiAnalysis founder Dylan Patel for a wide-ranging deep dive into what actually constrains AI compute scaling. The episode opens on the economics of the current buildout: hundreds of billions in hyperscaler CapEx and lab fundraising (OpenAI's $110B, Anthropic's $30B) is explained not as pure current-year compute spend but as prepayment for multi-year infrastructure - turbine deposits, data center construction, power purchase agreements. Dylan traces why Anthropic, despite deliberately conservative long-term contracting to avoid overcommitting, ended up compute-constrained relative to OpenAI's more aggressive multi-provider strategy, forcing it into costlier last-minute and revenue-share deals. A key throughline here is that an H100 GPU is worth more today than three years ago - counter to the standard depreciation bear case - because rising model quality and token value have outpaced the erosion from faster newer chips, a dynamic Dylan connects to the Alchian-Allen effect: as GPU prices rise, the relative price gap between frontier and lesser models shrinks, pushing buyers toward the best model available.

The conversation's structural core is Dylan's argument that the bottleneck constraining AI compute keeps migrating down the supply chain - from CoWoS packaging, to power, to (this year and next) fab clean-room construction, and ultimately, by the late 2020s, to ASML's EUV lithography tools. He walks through the physical mechanics of EUV in detail (tin-droplet plasma sources, multilayer mirror optics, nine-G reticle stage movement, sub-nanometer overlay accuracy) to explain why production can't simply be scaled up on demand, and calculates that even aggressive tool production growth caps total AI chip output near 200 gigawatts/year by 2030 - a number he says is compatible with, not far below, the most ambitious lab projections. A large middle section covers the parallel memory crunch: HBM's bandwidth advantage over commodity DRAM makes switching to cheaper memory impractical for AI workloads, and the resulting price increases are already showing up as roughly $100-150 in added iPhone bill-of-materials cost and a projected halving of low-end/mid-range smartphone volumes over the next two years.

The episode's most geopolitically charged stretch addresses US-China dynamics: Dylan frames the outcome as timeline-dependent (fast AI progress favors the US's compounding compute/revenue lead; slow progress gives China time to build a fully indigenized supply chain) and argues Huawei, if it had TSMC access, could plausibly out-execute Nvidia given its combination of chip design talent, in-house fabs, and end-market reach. The closing sections turn to alternative scaling paths: Dylan is skeptical of Elon Musk's space-GPU ambitions - not because space power isn't cheap, but because chips remain the binding constraint everywhere, and space adds deployment-delay and unreliable-interconnect costs that Earth-based facilities don't face - while he's comparatively bullish that power and land won't meaningfully constrain the US buildout this decade, citing over a dozen alternative generation technologies and unused grid headroom. The interview ends on humanoid robots (arguing most intelligence should stay centralized in the cloud rather than on-device) and a final risk assessment of Taiwan, where Dylan argues that even successfully evacuating TSMC's engineers wouldn't prevent a catastrophic multi-year setback to global AI compute capacity.

Notable Quotes

"There is some sense that we're in a fast takeoff. It's not like we're talking about a Dyson sphere by X date, it's more like the revenue is compounding at such a rate that it does affect economic growth." - Dylan Patel

"It's a snake eating its own tail meme because you can't make the tools without the chips from Taiwan, which you can't use without the tools in Taiwan." - Dylan Patel

"But Elon doesn't win by doing 20% gains. He never wins that way. Elon wins when he swings for the fences and does 10X gains." - Dylan Patel

"Fifty gigawatts of economic CapEx in the data center, and what gets built on top of that in terms of tokens is even larger. It might be $100 billion worth of AI value into the supply chain, held up by this $1.2 billion worth of tooling that simply cannot expand its supply chain quickly." - Dylan Patel

"You only buy the memory crunch if you believe AI is going to take off in a huge way." - Dylan Patel