All podcasts / Invest Like the Best / Summary

Dylan Patel - The Infinite Demand for Tokens, Claude Mythos, and Supply Constraints

2026-04-23 - 46 min - source - Read full transcript
Patrick O'Shaughnessy (host)Dylan Patel

Key insights

SemiAnalysis's own Claude Code spend rocketed from tens of thousands of dollars a year to a $7 million annualized run rate, now over 25% of the firm's roughly $25 million salary base and on pace to exceed 100% of salary spend by year end.
Patel says spend began inflecting in January after an internal 'Claude psychosis' moment spread from one non-technical executive to the rest of the firm, driven by an enterprise Anthropic contract. He frames this as visceral, firsthand evidence of the demand-side story he tracks professionally.
token-demand-explosion
Willingness to pay for frontier tokens is effectively unbounded because value generated scales with usage, not price paid.
Patel says Anthropic could double or even significantly raise Opus pricing and he would keep paying, because the information and analysis those tokens produce is worth far more to his business than the token cost. He argues this same logic applies broadly to any customer whose token use is generating outsized economic value.
token-demand-explosion
Individual employees are now replicating work that previously required entire specialized teams, for a few thousand dollars of tokens.
Examples cited: one person built a GPU-accelerated chip-imaging tool (hosted on CoreWeave) that a prior employer needed a full team to build and maintain; an economist alone built a 2,000-task AI-capability benchmark and 'Phantom GDP' metric that would have taken a 200-person team a year; another employee mapped every U.S. power plant and transmission line in a few weeks, rivaling a company with 100 employees and a decade of work.
ai-implementation-cost-collapse
The scarce resource in the economy has flipped from execution to idea selection: implementation is now cheap and fast, so the bottleneck is choosing which ideas are worth the token spend and then selling what gets built.
Patel argues that historically 'execution was very, very difficult' while ideas were comparatively easy to generate; now that AI collapses implementation cost and time, the constraint becomes picking good ideas and having the capital and distribution to capture value from what AI builds.
ai-implementation-cost-collapse
Patel warns of a 'permanent underclass' forming among people and firms that fail to use more tokens, generate value from them, and capture that value.
He distinguishes the 'boring' approach (working fewer hours because AI absorbs part of the job) from the 'cool' approach (working the same hours but producing far more output and capturing outsized returns), and frames failure to do the latter as an existential competitive risk as AI adoption becomes table stakes.
ai-implementation-cost-collapse
Anthropic's internal model, referred to as 'Mythos,' represents roughly two years' worth of capability jump on benchmarks and is being deliberately withheld or selectively released (for example to major banks for cybersecurity use) rather than shipped broadly; the public Opus 4.7 release was intentionally weakened on some capabilities per its own model card.
Patel says Anthropic internally moved from an L4-equivalent software engineer model to an L6-equivalent in about two months, and that the company is worried enough about Mythos's real-world impact that it priced selective cyber access at 5-10x normal token cost rather than releasing it generally.
model-scaling-and-mythos
Anthropic's gross margins are likely above 72% and possibly much higher, because revenue has grown from roughly $9 billion to a $40-45 billion run rate far faster than its compute footprint, meaning the company rations demand through rate limits and price rather than competing on cost.
Patel calculates this floor by assuming all incremental compute Anthropic acquired went to inference (a conservative assumption, since some clearly went to R&D for Mythos and Opus 4.7). He notes this stands in sharp contrast to a leaked ~30%-gross-margin figure from Anthropic's funding documents earlier in the year.
model-scaling-and-mythos
Memory supply (DRAM and NAND) can only grow capacity in the low-double-digit-to-30%-per-year range, so true incremental supply responding to the current demand shock will not arrive until 2027 or 2028 at the earliest; Patel expects DRAM prices to double or triple again from already-elevated levels.
He frames this as demand destruction via pricing rather than rationing: because memory makers cannot meaningfully accelerate capacity builds, the only way to reallocate scarce supply toward AI is to price out lower-value uses.
compute-supply-bottlenecks
TSMC's capex trajectory, roughly $56-57 billion this year and potentially approaching $100 billion by around 2028, whipsaws the entire downstream semiconductor equipment supply chain (ASML, Lam Research, Applied Materials, and smaller suppliers), and even niche, previously overlooked inputs like copper foil and glass fiber for PCBs are sold out with buyers making prepayments.
Patel separately flags CPUs, not just GPUs and ASICs, as an underappreciated bottleneck: reinforcement-learning training environments run on CPU (not accelerator silicon), and deployed inference-serving applications also route their final output through CPU infrastructure, so CPU demand is rising alongside accelerator demand.
compute-supply-bottlenecks
Patel says the hardest thing to forecast is not supply, which SemiAnalysis understands well, but 'tokenomics': how AI usage and adoption diffuse through the economy and what real value they create, since standard GDP accounting misses most of it.
He and his team have repeatedly underestimated near-term revenue (crazy January estimates for February were smashed by Anthropic, and the pattern repeated into March), and argues the economic value AI tokens create for someone like himself or Patrick is far larger than what shows up in official statistics, a gap he calls 'phantom GDP.'
model-scaling-and-mythos
Patel predicts large-scale public protests against Anthropic and AI broadly within about three months, driven by AI already polling less popular than ICE or politicians, and argues lab leaders are worsening this by talking publicly about future capabilities instead of showing present, uplifting use cases.
He points to comment sections cheering the (reported) Molotov cocktail attacks on Sam Altman's house as an early warning sign, and argues the average person has no personal connection to anyone at Anthropic or OpenAI, so they default to viewing the labs as a small, opaque group about to automate jobs and destroy society. His prescription: leadership should stop doing interviews that alienate audiences (citing Altman's Tucker Carlson appearance), stop emphasizing future capability, and focus messaging on present-day benefits.
ai-backlash-and-perception

Media referenced

Companies

Techniques and frameworks

Summary

Dylan Patel returns for a second conversation with Patrick O'Shaughnessy, and the throughline is the widening gap between AI token demand and the world's ability to supply it. Patel opens with SemiAnalysis's own numbers as a case study: the firm's AI spend went from tens of thousands of dollars a year to a $7 million Claude Code run rate, now more than a quarter of its ~$25 million salary base and trending toward exceeding it by year end. He walks through several examples of individual, often non-technical, employees using Claude Code to replicate work that used to require whole specialized teams, a chip-imaging tool that once needed an Intel team to build, a 2,000-task AI-capability benchmark and a "Phantom GDP" metric built solo by an economist, and a full U.S. power-grid mapping project that rivals a hundred-person incumbent, all for a few thousand dollars of tokens each.

That experience frames Patel's central argument: implementation, historically the scarce and difficult part of building anything, has collapsed in cost and time, which flips the bottleneck to choosing the right ideas and being able to sell or capture value from what gets built. He warns that people and firms who do not ramp up token usage, generate value from it, and capture that value risk falling into a "permanent underclass," while those with the best access to frontier tokens and the capital to leverage them compound an advantage. This ties into a discussion of Anthropic's internal model, referred to as "Mythos," which Patel describes as roughly a two-year capability jump that the company has deliberately withheld or selectively released (for example to banks for cybersecurity), while intentionally shipping a weakened public version, Opus 4.7. He estimates Anthropic's gross margins are likely above 72%, given revenue growth from roughly $9 billion to a $40-45 billion run rate has vastly outpaced compute growth, meaning demand is being rationed through price and rate limits rather than cost competition.

On the supply side, Patel details bottlenecks stacking up across the entire hardware chain: memory capacity can only grow 20-30% a year, so true relief from the current shortage will not arrive until 2027-2028, and he expects DRAM prices to double or triple again. TSMC's capex is on a path from roughly $56-57 billion this year toward a possible $100 billion by 2028, a trajectory that whipsaws downstream equipment makers like ASML, Lam Research, and Applied Materials, and even reaches into overlooked inputs like copper foil and glass fiber, which are already sold out with prepayments. He also flags CPUs as a hidden bottleneck, since reinforcement-learning training environments and deployed inference applications both depend on CPU infrastructure that gets little of the attention GPUs and ASICs receive.

Patel is candid that the hardest thing to model is not supply, which SemiAnalysis tracks closely, but "tokenomics," how AI usage diffuses through the economy and what value it actually creates, a gap he thinks standard GDP statistics miss entirely. The conversation closes on a darker note: Patel predicts large-scale public protests against Anthropic and AI more broadly within about three months, pointing to AI's already-low public favorability and arguing that lab leaders undermine themselves by discussing future capabilities in the abstract rather than showing concrete, present-day benefits that ordinary people can relate to.

Notable Quotes

"If you don't use more tokens, you'll never escape the permanent underclass." - Dylan Patel

"What used to matter a lot was execution was very, very fucking difficult and ideas were cheap. Now ideas are cheap and plentiful, but execution is very easy." - Dylan Patel

"It's clearly by every subjective metric, amazing. But where is the phantom GDP?" - Dylan Patel

"I think we'll see large-scale protests against AI in three months." - Dylan Patel

"First of all, Sam Altman and Dario have to stop getting on interviews. They're so uncharismatic." - Dylan Patel