All podcasts / No Priors / Summary

No Priors Ep. 143 | With ElevenLabs Co-Founder Mati Staniszewski

2025-12-11 - 42 min - source - Read full transcript
Sarah Guo (host)Mati Staniszewski

Key insights

ElevenLabs scaled to $300M ARR and 350 employees in about three years by running roughly two businesses under one platform.
The company splits close to 50/50 between a self-serve creative platform (5M+ monthly active users doing narration, dubbing, voiceovers) and an enterprise sales-led agent platform serving a few thousand customers from Fortune 500s to fast-growing AI startups.
voice-ai
The founding insight came from watching dubbed films in Poland, where every character shares one flat narrator voice.
Staniszewski and co-founder Piotr Dabkowski (both Polish) grew up watching foreign films dubbed with a single monotone narrator covering every character, which convinced them AI could instead preserve the original speaker's voice, emotion, and intonation across a language switch - not just translate the words.
voice-ai
ElevenLabs organizes around sequential, problem-specific 'labs' rather than one flat product org.
Each lab (voice, then agents, then music) combines researchers, engineers, and operators around a single hard problem, does the deep research first, builds a simple product layer on top, then expands into a fuller product suite before the next lab spins up.
org-design
The company uses an explicit three-month rule to decide whether to wait for research or ship a product workaround.
Product teams are free to build ahead of research; if closing a capability gap is estimated at more than three months of research work, the product team builds something now, if it's under three months, they wait for the research to land.
research-product-strategy
Model quality alone does not determine perceived output quality - the specific voice matters enormously.
Citing Artificial Analysis benchmarks, Staniszewski notes that comparing model A to model B with different voices can shift perceived quality dramatically even when the underlying models are close, because humans are highly sensitive to voice given it is our oldest interaction mode.
voice-ai
Evaluating voice AI is still an unsolved problem, partly because the industry lacks a vocabulary to label how something was said.
Traditional data-labeling vendors could transcribe what was said but not the emotion, accent, or delivery, so ElevenLabs had to build its own internal capability to describe audio qualitatively, in addition to running enterprise-specific 'voice concierge' teams that help customers pick the right voice for their use case.
voice-ai
ElevenLabs' business model differentiates itself from both consulting firms and narrow point-solution competitors.
Unlike Palantir-style forward-deployed consulting (broad digital-transformation scope) or a single-use-case agent company like Sierra, ElevenLabs positions itself as an open platform for companies that want to deploy multiple voice/agent experiences (support, sales, internal tooling) themselves, backed by ElevenLabs' own forward-deployed engineering.
research-product-strategy
Staniszewski argues foundation-model labs like Google and OpenAI will keep underinvesting in audio because winning requires architectural breakthroughs, not just scale, and few people can do that work.
He estimates only 50-100 researchers globally have the skill to produce the necessary audio model breakthroughs, of whom ElevenLabs employs roughly 10, and argues that focus on that narrow expert pool, not headcount, explains why ElevenLabs beats larger labs on text-to-speech and speech-to-text benchmarks.
ai-agents
Technology advantages in AI are not permanent moats, which he says makes investors uncomfortable but is the accurate model.
Research or product leads might last a year or might last a decade, but they are not infinitely defensible; what compounds over time is the ecosystem built around a research head start - distribution, the voice/integration library, and workflows customers build on top.
research-product-strategy
Real-time, fused speech-to-speech translation ('Babel fish' dubbing) is roughly one to two years away, but cascaded pipelines remain the enterprise default for now.
Cascaded speech-to-text, LLM, text-to-speech architectures give more structure, visibility into each step, tool-calling reliability, and are less prone to hallucination, so Staniszewski expects them to be the right choice for reliable enterprise use for at least another year even as fused speech-to-speech models become more expressive.
ai-agents
High-traction agent use cases are shifting from reactive customer support to proactive commerce, immersive IP, and tutoring.
Examples given: Meesho and Square moving from reactive support/ordering into proactive shopping assistants that navigate customers to products; Epic Games licensing Darth Vader's voice for a live, interactive Fortnite experience; and Chess.com and MasterClass building AI tutors that let users learn in a specific expert's voice and practice interactively (e.g., a live negotiation with Chris Voss).
ai-agents
Staniszewski expects AI tutoring, not companions, to be voice AI's biggest long-term unlock, but argues it should complement, not replace, human interaction.
He is personally more excited by a 'Jarvis'-style personal-assistant use case than social AI companions, and while he expects AI tutors to consume a large share of learning time, he believes an explicit, tech-free chunk of human-to-human time should remain part of education for emotional guidance and peer interaction.
future-of-interaction

Media referenced

Companies

Techniques and frameworks

Summary

Sarah Guo interviews Mati Staniszewski, co-founder and CEO of ElevenLabs, about scaling the voice AI company from its 2022 founding to $300M in ARR and 350 employees, split roughly evenly between a self-serve creative platform (narration, dubbing, voiceovers) and an enterprise agent platform (customer support, proactive commerce, internal tooling). Staniszewski traces the company's origin to a personal frustration: dubbed foreign films in his native Poland use one flat narrator voice for every character, which convinced him and his co-founder that AI could instead preserve a speaker's original voice, emotion, and intonation across a language switch, rather than merely translating words.

Much of the conversation covers how ElevenLabs organizes to combine deep research with product velocity. The company builds sequential "labs" - a voice lab, then an agent lab, then a music lab - each a small combined team of researchers, engineers, and operators assembled around one hard problem, doing the research first and layering a product on top before the next lab spins up. Internally, teams follow a three-month rule: if closing a capability gap through research would take longer than roughly three months, a product team ships a workaround now rather than waiting. Staniszewski also argues that evaluating voice AI quality remains genuinely unsolved industry-wide - benchmarks are highly voice-dependent, and even data-labeling vendors historically lacked the vocabulary to describe how something was said (emotion, accent, delivery), not just what was said, forcing ElevenLabs to build that labeling capability itself alongside dedicated "voice concierge" teams for enterprise customers.

On competitive positioning, Staniszewski contrasts ElevenLabs with both consulting-style forward-deployed engineering (Palantir, his own prior employer) and narrow point-solution agent companies like Sierra: ElevenLabs positions itself as an open platform for customers who want to build several voice and agent experiences themselves, backed by its own engineering support. He argues foundation-model labs like Google and OpenAI will keep underinvesting in audio specifically because winning requires architectural breakthroughs rather than raw scale, and estimates only 50-100 researchers globally can do that work, of whom ElevenLabs employs roughly 10 of the best. He is candid that none of this creates a permanent moat: technology advantages might last a year or a decade, but what compounds is the ecosystem - voice libraries, integrations, distribution - built around a research head start.

The back half of the episode surveys emerging use cases and the near-term technical roadmap. Reactive customer support is the most mature use case, but Staniszewski sees momentum shifting toward proactive commerce assistants (cited examples: Meesho, Square), interactive IP licensing (ElevenLabs voiced Darth Vader for a live, playable Fortnite experience with Epic Games), and AI tutoring (Chess.com lessons in Hikaru Nakamura's or Magnus Carlsen's voice, a MasterClass negotiation practice session with Chris Voss). He also describes ElevenLabs' work with Ukraine's Ministry of Digital Transformation building what he calls a first "agentic government." On the roadmap, he expects real-time, fused speech-to-speech translation (the "Babel fish" vision) is roughly one to two years out, but expects cascaded speech-to-text/LLM/text-to-speech pipelines to remain the more reliable enterprise default for at least another year because they are more structured, tool-callable, and less prone to hallucination.

Closing on the future of interaction, Staniszewski says he's more excited by a "Jarvis"-style personal assistant than by social AI companions, though he expects companions to be a large market addressing loneliness. He predicts education will be voice AI's biggest unlock - an on-demand personal tutor, potentially voiced by figures like Richard Feynman or Einstein delivering historical lectures - but argues explicit human-to-human, technology-free time should remain part of learning for emotional guidance and peer interaction, not be fully displaced by AI tutors.

Notable Quotes

"Actually most advantages in technology, like they could last you a year or they could last you 10, but they're not like infinitely defensible." - Mati Staniszewski

"If you watch a movie in Polish language, a foreign movie in Polish language, all the voices, whether it's a male voice or a female voice, are narrated with one single character." - Mati Staniszewski

"We think there's maybe 50 to 100 researchers in audio space that could do it. We think we have probably 10 of them in the company that are some of the best ones." - Mati Staniszewski

"If a product team thinks we should deliver value to the customer by doing something different, they can... rough rule of thumb is like three months. If we think it's going to be longer than three months, we will probably build it. If it's less than that, we probably won't." - Mati Staniszewski