All podcasts / No Priors / Summary

Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud

2026-05-01 - 43 min - source - Read full transcript
Sarah Guo (host)Elad Gil (host)Tuhin Srivastava

Key insights

Over 95% of tokens in Baseten's dedicated-inference business now run custom or post-trained models, not stock open-source weights.
Srivastava says almost every customer either fine-tunes a model on its own data or recompiles it for performance; 'no one is just running the vanilla open source weights.' The market has shifted from choosing a model API to deploying a custom, specialized model.
custom-models-post-training
Baseten acquired a post-training research team because post-training and inference performance turned out to be deeply coupled problems, not separate ones.
The acquired team was itself a former Baseten customer building post-trained models who realized they needed to become an inference company. Srivastava says how a model is trained increasingly determines how it must be quantized and served, so Baseten now treats post-training and inference as two sides of the same loop and is building training APIs to make continual learning routine.
custom-models-post-training
Srivastava's advice for moving to custom models: prove value with a best-in-class off-the-shelf model at real scale before investing in post-training.
He paraphrases an old joke - 'no GPUs pre-product-market-fit' - as now being 'no post-training pre-product-fit.' Customers should have a proven, scaled user signal worth optimizing before specializing a model, not specialize first and look for value after.
custom-models-post-training
The application layer survives model commoditization because durable value sits in proprietary workflow signal, not in model weights.
Using Abridge, an ambient clinical scribe deployed across most US hospitals, as the example: the defensible signal is the clinicians' edits to AI-drafted notes and what happens with those notes downstream inside the EMR - workflow data frontier labs structurally cannot access, which can then be used to post-train specialized reward models.
application-layer-moat
Serving AI-native application companies gives Baseten an indirect read on enterprise requirements without having to sell into enterprises directly.
Customers like Abridge, OpenEvidence, and Decagon sell into large enterprises and healthcare systems at scale, so their infrastructure requests (data retention, deployment region, GPU choice, latency, model transparency) function as a live translation of what enterprise buyers actually need - positioning Baseten well to serve healthcare and other regulated verticals indirectly.
application-layer-moat
Publicly discussed AI compute shortages understate how tight the real crunch is.
Baseten runs 90 clusters across 18 clouds at roughly mid-90s percent utilization. Many newer capacity providers are inexperienced at running data centers and don't understand inference-specific SLAs, which Srivastava says leaves only three or four 'gold tier' clouds the company actually trusts, on top of raw compute scarcity.
compute-supply-and-capacity-economics
GPU capacity contract terms shifted hard toward multi-year prepaid commitments in the six months before this conversation.
Securing roughly 1,000-plus B200s from a reputable cloud now typically requires a three-to-five-year contract with 20-30% of total contract value prepaid. That favors buyers with cheap cost of capital, which Srivastava says is part of why capacity-heavy AI infrastructure companies would consider going public sooner.
compute-supply-and-capacity-economics
Raw GPU access is a commodity, but inference bundled with software is sticky - and that stickiness, not compute ownership, is Baseten's real moat.
Srivastava says none of Baseten's top 30 customers have ever churned, with roughly 400% annual net dollar retention, and contrasts this with 'GPUs as a service,' which he says customers already treat as interchangeable across providers.
compute-supply-and-capacity-economics
Falling inference cost increases total inference spend rather than shrinking it, a practical version of Jevons paradox.
As cost per token drops, developers insert more intelligence into products and let agents run longer instead of stopping at a fixed, cheaper answer. Srivastava says this pattern holds across nearly all Baseten's customers and expects it to continue for years, calling inference 'the last market' - even in an AGI scenario, inference is what remains to be sold.
compute-supply-and-capacity-economics
Srivastava sees no real evidence of embedded backdoors or bias in Chinese open-source models, beyond isolated early cases people caught quickly.
He argues network-level isolation makes data-exfiltration concerns largely theoretical for self-hosted models, and that DeepSeek-class models can run at roughly 20% the cost of comparable closed models with comparable or better latency and reliability - so restricting access would be a real loss to US AI progress.
open-source-and-china-geopolitics
Chinese government subsidies for open-source model development effectively pass through as an indirect subsidy to US enterprises that adopt those models.
Srivastava frames the current dynamic as the Chinese government absorbing costs that would otherwise fall on US companies, while still arguing it's strategically important for the US to build its own credible open-source frontier in case that supply disappears or the dynamic reverses.
open-source-and-china-geopolitics
Scaling past a flat organization required deliberately hiring leaders against a narrow, explicit rubric, which felt counterintuitive to an engineering-heavy team.
Baseten stayed very flat until roughly 12-18 months before this conversation; Srivastava says bringing in dedicated leadership felt like 'overhead' at first, and that needing to be involved in everything is 'a bit of a cop out as a founder.' The rubric itself is narrow - first-principles thinkers with a high work bar, low ego, no hero culture - which he credits with keeping unwanted turnover low, since people who need heavy management tend to self-select out.
scaling-team-and-operations-culture

Companies

Techniques and frameworks

Summary

Tuhin Srivastava, founder and CEO of Baseten, joins Sarah Guo and Elad Gil to explain how the company scaled roughly 30x in the past year toward more than $1B in projected revenue, and what that growth reveals about where AI inference spending is actually going. The headline number: over 95% of tokens running through Baseten's dedicated-inference business are custom or post-trained models, not stock open-source weights. Customers fine-tune with proprietary data, recompile for performance, or both - a shift Srivastava says reflects the market moving past "which model API do I call" into "which specialized model do I run." This is also why Baseten acquired a post-training research team: the company found that how a model is post-trained increasingly determines how it must be quantized and served, making post-training and inference two sides of the same engineering problem rather than separate stages.

A recurring thread is why the application layer survives even as frontier labs get more capable. Srivastava's answer is that durable value lives in workflow-embedded signal a lab structurally cannot reach - his central example is Abridge, an ambient clinical scribe used across most US hospitals, where the moat is what clinicians edit in AI-drafted notes and what happens with that note three steps downstream in the EMR. That signal can be used to post-train specialized reward models that a general-purpose frontier model has no path to replicate. He extends this to Baseten's own position: because its fastest-growing customers (Abridge, OpenEvidence, Decagon, Writer, Gamma) sell into large enterprises, their infrastructure requirements function as a real-time translation of what enterprise buyers actually need - data retention rules, deployment regions, latency tolerances, GPU choices - without Baseten having to sell into the enterprise directly.

On compute supply, Srivastava argues the public narrative about a GPU shortage understates the real problem. Baseten runs 90 clusters across 18 clouds at roughly mid-90s percent utilization, and beyond raw scarcity, many capacity providers are operationally unreliable - inexperienced at running data centers and unfamiliar with inference-specific SLAs - leaving perhaps three or four "gold tier" clouds the company actually trusts. Contract terms have hardened over the prior six months: locking in 1,000-plus B200s from a reputable provider now typically requires a three-to-five-year term with 20-30% of total contract value prepaid, which rewards buyers with cheap cost of capital and is part of why he thinks capacity-heavy AI infrastructure companies might go public sooner rather than later. He's explicit that raw GPU access itself is commodity - "GPUs as a service is not sticky" - while inference bundled with software is; Baseten's top 30 customers have never churned, with roughly 400% annual net dollar retention.

The conversation covers the geopolitics of Chinese open-source models directly. Srivastava says he's seen no real evidence of embedded backdoors or bias beyond a couple of isolated early cases that were quickly caught, and argues that network-isolated, self-hosted deployment makes most exfiltration concerns theoretical. He frames DeepSeek-class models, which he says can run at roughly 20% the cost of comparable closed models with comparable or better latency, as effectively a subsidy the Chinese government is passing through to US enterprises that adopt them - while still arguing it's strategically important for the US to build its own credible open-source frontier in case that supply disappears.

The last stretch turns to how Baseten scaled its team and culture through 30x growth. Srivastava describes a company that stayed very flat until 12-18 months before this conversation, and says bringing in dedicated leaders felt counterintuitive to an engineering-first instinct - needing to be involved in everything, he argues, is "a bit of a cop out as a founder" that signals the wrong people are in place. Baseten's hiring rubric is deliberately narrow: first-principles thinkers with a high bar for work, low ego, and no hero culture, which he credits with keeping unwanted turnover low. He closes by describing Baseten's operational intensity - a running joke about an office siren for P0 incidents - as core to running infrastructure other companies treat as mission-critical, and predicts that falling inference costs will keep expanding total AI spend rather than shrinking it, since cheaper intelligence gets used more, not less.

Notable Quotes

"GPUs as a service is not sticky - customers see that as commodity. Inference with the software layer included is incredibly sticky." - Tuhin Srivastava

"It is truly - I think we're kind of in a world that is the last market. Even if there's AGI, all that's left is inference." - Tuhin Srivastava

"If you network bound these models, they're not magically going to be able to cross those network boundaries." - Tuhin Srivastava

"No post-training pre-product-fit." - Tuhin Srivastava