Baseten CEO Tuhin Srivastava on the AI Inference Crunch, Custom Models, and Building the Inference Cloud
Key insights
Companies
- Baseten - The AI inference cloud Srivastava co-founded and runs; discussed throughout as growing roughly 30x in a year toward over $1B in revenue.
- Abridge - Ambient clinical scribe customer used as the central example of a workflow-embedded data moat frontier labs cannot access.
- OpenEvidence - Fast-growing AI-native healthcare customer cited alongside Abridge and Decagon as a top-scale Baseten customer.
- Decagon - AI-native customer-support company cited as an example of a new application-layer company selling AI directly to enterprises.
- Cursor - Referenced as one of Baseten's fastest-growing, largest-scale AI-native customers.
- Writer - Cited as an enterprise-facing AI-native customer whose infrastructure requirements indirectly reveal enterprise buyer needs.
- Gamma - Cited alongside Writer as an enterprise-facing AI-native customer.
- Stripe - Used as an analogy for building infrastructure for the most demanding customers first and growing into the broader enterprise market later.
- AWS - Referenced for scale-limitation comparisons Baseten hit at its own hyperscaler dependencies, and in an anecdote about AWS executives' pagers going off mid-meeting.
- Nvidia - Discussed as the dominant GPU and CUDA ecosystem player that is hard to displace in the near term given its supply chain and developer ecosystem advantages.
- Groq - Referenced regarding inference-specific chip architectures (LPUs) as part of a possible multi-chip future.
- Meta - Referenced as the maker of Llama, credited with shifting the open-source model landscape a few years ago.
- Mistral - Referenced as an early open-source model provider in the market's evolution.
- DeepSeek - Central example in the discussion of Chinese open-source model cost, performance, and geopolitical risk; cited as running at roughly 20% the cost of comparable closed models.
- Moonshot AI - Referenced (as its Kimi model) among the frontier open models Baseten customers choose to run.
- Canopy Labs - Referenced as the maker of the Orpheus open-source text-to-speech model that customers run on Baseten.
- OpenAI - Referenced as a closed frontier lab; Srivastava cites OpenAI and Anthropic figures publicly calling inference capacity strategic.
- Anthropic - Referenced as one of the closed frontier labs still ahead on raw model capability.
- Google - Referenced as one of the closed frontier labs still ahead on raw model capability.
Techniques and frameworks
- Post-training / RL fine-tuning on proprietary reward signal - Customers customize open models on their own data and workflow-derived reward signals to build specialized, defensible models.
- Quantization tied to training method - Srivastava describes how a model was trained increasingly determines how it must be quantized for efficient inference, linking post-training and inference engineering.
- KV cache-aware routing - Routing inference requests to preserve and reuse key-value cache for latency and cost gains.
- Speculative decoding - Newer speculative-generation techniques Baseten adopts to speed up token generation.
- Disentangling prefill and decode - Treating the prefill and decode phases of inference as separate optimization problems rather than one pipeline.
- Async batch inference - Batching inference requests asynchronously to drive up utilization and margin, framed as core to what an 'inference cloud' should offer.
- Multi-cloud runtime fabric - Baseten's technology spanning 90 clusters across 18 clouds to abstract reliability, latency, and failover from customers and rapidly onboard new capacity providers.
Summary
Tuhin Srivastava, founder and CEO of Baseten, joins Sarah Guo and Elad Gil to explain how the company scaled roughly 30x in the past year toward more than $1B in projected revenue, and what that growth reveals about where AI inference spending is actually going. The headline number: over 95% of tokens running through Baseten's dedicated-inference business are custom or post-trained models, not stock open-source weights. Customers fine-tune with proprietary data, recompile for performance, or both - a shift Srivastava says reflects the market moving past "which model API do I call" into "which specialized model do I run." This is also why Baseten acquired a post-training research team: the company found that how a model is post-trained increasingly determines how it must be quantized and served, making post-training and inference two sides of the same engineering problem rather than separate stages.
A recurring thread is why the application layer survives even as frontier labs get more capable. Srivastava's answer is that durable value lives in workflow-embedded signal a lab structurally cannot reach - his central example is Abridge, an ambient clinical scribe used across most US hospitals, where the moat is what clinicians edit in AI-drafted notes and what happens with that note three steps downstream in the EMR. That signal can be used to post-train specialized reward models that a general-purpose frontier model has no path to replicate. He extends this to Baseten's own position: because its fastest-growing customers (Abridge, OpenEvidence, Decagon, Writer, Gamma) sell into large enterprises, their infrastructure requirements function as a real-time translation of what enterprise buyers actually need - data retention rules, deployment regions, latency tolerances, GPU choices - without Baseten having to sell into the enterprise directly.
On compute supply, Srivastava argues the public narrative about a GPU shortage understates the real problem. Baseten runs 90 clusters across 18 clouds at roughly mid-90s percent utilization, and beyond raw scarcity, many capacity providers are operationally unreliable - inexperienced at running data centers and unfamiliar with inference-specific SLAs - leaving perhaps three or four "gold tier" clouds the company actually trusts. Contract terms have hardened over the prior six months: locking in 1,000-plus B200s from a reputable provider now typically requires a three-to-five-year term with 20-30% of total contract value prepaid, which rewards buyers with cheap cost of capital and is part of why he thinks capacity-heavy AI infrastructure companies might go public sooner rather than later. He's explicit that raw GPU access itself is commodity - "GPUs as a service is not sticky" - while inference bundled with software is; Baseten's top 30 customers have never churned, with roughly 400% annual net dollar retention.
The conversation covers the geopolitics of Chinese open-source models directly. Srivastava says he's seen no real evidence of embedded backdoors or bias beyond a couple of isolated early cases that were quickly caught, and argues that network-isolated, self-hosted deployment makes most exfiltration concerns theoretical. He frames DeepSeek-class models, which he says can run at roughly 20% the cost of comparable closed models with comparable or better latency, as effectively a subsidy the Chinese government is passing through to US enterprises that adopt them - while still arguing it's strategically important for the US to build its own credible open-source frontier in case that supply disappears.
The last stretch turns to how Baseten scaled its team and culture through 30x growth. Srivastava describes a company that stayed very flat until 12-18 months before this conversation, and says bringing in dedicated leaders felt counterintuitive to an engineering-first instinct - needing to be involved in everything, he argues, is "a bit of a cop out as a founder" that signals the wrong people are in place. Baseten's hiring rubric is deliberately narrow: first-principles thinkers with a high bar for work, low ego, and no hero culture, which he credits with keeping unwanted turnover low. He closes by describing Baseten's operational intensity - a running joke about an office siren for P0 incidents - as core to running infrastructure other companies treat as mission-critical, and predicts that falling inference costs will keep expanding total AI spend rather than shrinking it, since cheaper intelligence gets used more, not less.
Notable Quotes
"GPUs as a service is not sticky - customers see that as commodity. Inference with the software layer included is incredibly sticky." - Tuhin Srivastava
"It is truly - I think we're kind of in a world that is the last market. Even if there's AGI, all that's left is inference." - Tuhin Srivastava
"If you network bound these models, they're not magically going to be able to cross those network boundaries." - Tuhin Srivastava
"No post-training pre-product-fit." - Tuhin Srivastava