No Priors Ep. 143 | With ElevenLabs Co-Founder Mati Staniszewski
Key insights
Media referenced
- The Hitchhiker's Guide to the Galaxy - other - Referenced for its 'Babel fish' concept as the aspirational end state for real-time, in-person spoken-language translation.
- Lex Fridman Podcast interview with Narendra Modi - podcast - Cited as an example of ElevenLabs dubbing technology preserving the original speaker's voice, emotion, and intonation across a language switch.
- MasterClass: Chris Voss negotiation lesson - other - ElevenLabs built an interactive layer letting learners call a voice agent and run a live practice negotiation with FBI negotiator Chris Voss after the lesson.
Companies
- ElevenLabs - Staniszewski's company: $300M ARR, 350 employees, founded 2022, building foundational audio (text-to-speech, speech-to-text, orchestration) plus creative and agent platform products.
- Palantir - Staniszewski's prior employer; used as the reference point for a forward-deployed-engineering business model versus a narrow point solution.
- Sierra - Named as a phenomenal but narrower, use-case-specific (customer support agent) competitor versus ElevenLabs' broader open platform.
- Meesho - Large Indian e-commerce company where ElevenLabs shifted customer support from reactive (refunds, tracking) to a proactive voice shopping assistant.
- Square - Partner that started with voice ordering and is expanding it into a full discovery/checkout experience.
- Epic Games / Fortnite - ElevenLabs voiced Darth Vader for a live, interactive in-game experience that millions of Fortnite players could talk to.
- Chess.com - Built an AI tutoring experience letting users learn chess in the voice/style of players like Hikaru Nakamura and Magnus Carlsen.
- Ukraine Ministry of Digital Transformation - Building what Staniszewski calls a first 'agentic government' - proactive citizen services and AI-assisted internal ministry operations, coordinated by engineering leads embedded in each ministry.
- Artificial Analysis - Benchmark provider cited for evidence that swapping the voice used, independent of model quality, has an outsized effect on how good a model is perceived to be.
Techniques and frameworks
- Problem-first lab structure - ElevenLabs organizes around sequential 'labs' (voice lab, then agent lab, then music lab) - small combined teams of researchers, engineers, and operators assembled around one problem, doing deep research first, then building a product layer, before spinning up the next lab.
- Three-month build-vs-wait rule - Internal rule of thumb: if closing a capability gap via research would take longer than about three months, a product team builds a workaround now; if less than three months, they wait for the research.
- Cascaded vs. fused speech architecture - Cascaded pipelines (speech-to-text, then LLM, then text-to-speech) stay more reliable and tool-callable for enterprise use today; fused speech-to-speech models are more expressive but can hallucinate, and are the longer-term research bet.
Summary
Sarah Guo interviews Mati Staniszewski, co-founder and CEO of ElevenLabs, about scaling the voice AI company from its 2022 founding to $300M in ARR and 350 employees, split roughly evenly between a self-serve creative platform (narration, dubbing, voiceovers) and an enterprise agent platform (customer support, proactive commerce, internal tooling). Staniszewski traces the company's origin to a personal frustration: dubbed foreign films in his native Poland use one flat narrator voice for every character, which convinced him and his co-founder that AI could instead preserve a speaker's original voice, emotion, and intonation across a language switch, rather than merely translating words.
Much of the conversation covers how ElevenLabs organizes to combine deep research with product velocity. The company builds sequential "labs" - a voice lab, then an agent lab, then a music lab - each a small combined team of researchers, engineers, and operators assembled around one hard problem, doing the research first and layering a product on top before the next lab spins up. Internally, teams follow a three-month rule: if closing a capability gap through research would take longer than roughly three months, a product team ships a workaround now rather than waiting. Staniszewski also argues that evaluating voice AI quality remains genuinely unsolved industry-wide - benchmarks are highly voice-dependent, and even data-labeling vendors historically lacked the vocabulary to describe how something was said (emotion, accent, delivery), not just what was said, forcing ElevenLabs to build that labeling capability itself alongside dedicated "voice concierge" teams for enterprise customers.
On competitive positioning, Staniszewski contrasts ElevenLabs with both consulting-style forward-deployed engineering (Palantir, his own prior employer) and narrow point-solution agent companies like Sierra: ElevenLabs positions itself as an open platform for customers who want to build several voice and agent experiences themselves, backed by its own engineering support. He argues foundation-model labs like Google and OpenAI will keep underinvesting in audio specifically because winning requires architectural breakthroughs rather than raw scale, and estimates only 50-100 researchers globally can do that work, of whom ElevenLabs employs roughly 10 of the best. He is candid that none of this creates a permanent moat: technology advantages might last a year or a decade, but what compounds is the ecosystem - voice libraries, integrations, distribution - built around a research head start.
The back half of the episode surveys emerging use cases and the near-term technical roadmap. Reactive customer support is the most mature use case, but Staniszewski sees momentum shifting toward proactive commerce assistants (cited examples: Meesho, Square), interactive IP licensing (ElevenLabs voiced Darth Vader for a live, playable Fortnite experience with Epic Games), and AI tutoring (Chess.com lessons in Hikaru Nakamura's or Magnus Carlsen's voice, a MasterClass negotiation practice session with Chris Voss). He also describes ElevenLabs' work with Ukraine's Ministry of Digital Transformation building what he calls a first "agentic government." On the roadmap, he expects real-time, fused speech-to-speech translation (the "Babel fish" vision) is roughly one to two years out, but expects cascaded speech-to-text/LLM/text-to-speech pipelines to remain the more reliable enterprise default for at least another year because they are more structured, tool-callable, and less prone to hallucination.
Closing on the future of interaction, Staniszewski says he's more excited by a "Jarvis"-style personal assistant than by social AI companions, though he expects companions to be a large market addressing loneliness. He predicts education will be voice AI's biggest unlock - an on-demand personal tutor, potentially voiced by figures like Richard Feynman or Einstein delivering historical lectures - but argues explicit human-to-human, technology-free time should remain part of learning for emotional guidance and peer interaction, not be fully displaced by AI tutors.
Notable Quotes
"Actually most advantages in technology, like they could last you a year or they could last you 10, but they're not like infinitely defensible." - Mati Staniszewski
"If you watch a movie in Polish language, a foreign movie in Polish language, all the voices, whether it's a male voice or a female voice, are narrated with one single character." - Mati Staniszewski
"We think there's maybe 50 to 100 researchers in audio space that could do it. We think we have probably 10 of them in the company that are some of the best ones." - Mati Staniszewski
"If a product team thinks we should deliver value to the customer by doing something different, they can... rough rule of thumb is like three months. If we think it's going to be longer than three months, we will probably build it. If it's less than that, we probably won't." - Mati Staniszewski