All podcasts / No Priors / Summary / Transcript

No Priors Ep. 143 | With ElevenLabs Co-Founder Mati Staniszewski

2025-12-11 - source: youtube-captions

00:00:05Hi listeners, welcome back to No Priors.Today I'm here with Mi Stendes, theco-founder and CEO of 11 Labs, which wasfounded to change the way we interactwith each other [music] and withcomputers with voice. Over three shortyears, they've skyrocketed to [music]more than 300 million in run rate. Mottiand I talk about the future of voiceeducation, customer experience and theother applications of this voice as wellas how to build a multi-segment [music]from self-s serve to enterprise andcombined research and product company.Welcome Marty.>> S thanks for having me>> and thank you for doing this at 7 in themorning.>> Our pleasure. Thank you for doing thatat 7:00 in the morning. It's great we wewe got to finally do this together. Uh Ithink a lot of our listeners will haveused or played with 11 at some point butfor everybody else can you justreintroduce the company?>> Definitely we uh at 11 Labs we aresolving how humans and technologyinteract how you can create seamlesslywith that technology. Um what this meansin practice is we build foundationalaudio models. So models in a space tohelp you create speech that soundshuman, understand speech in a muchbetter way or orchestrate all thosecomponents to make it interactive andthen build products on top of thatfoundational models. And we have ourcreative product which is a platform forhelping you with

00:01:21narrations foraudiobooks with voiceovers for ads ormovies or dabs of those movies to otherlanguages. in our agent uh platformproduct which is effectively an offeringto help you elevate customer experiencebuilt an agent for personal AI educationnew ways of immersive immersive media uhbut all is kind of under under light ofthat mission of solving how we caninteract with technology on our terms ina better way>> you started the company in 2022>> that's right>> and you've had amazing like rocket shipgrowth since then I'm sure it's felt upand down different ways I want to askyou about that can you give a sense ofwhat the scale of the company is today.>> So we've grown to 350 people globally.We started from from Europe. We startedas a remote company and are still firstremote first but have hubs around theworld with London being the biggest, NewYork being second biggest, Warso, SanFrancisco and now Tokyo and and one inBrazil. We are at uh 300 million in inin ARR which is uh roughly 50/50 betweenself-s serve so a lot of subscriptionand creators using our creative platformand then approaching 50 uh% on theenterprise side using our agentsplatform uh uh work and that's on thesalesled classic salesled side and weserve more than 5

00:02:36million monthlyactivives on that on that on thatcreative uh side of the work and then onthe enterprise side we have few thousandcustomers from Fortune 500s to some ofthe fastest AI growing startups.>> I think this is such a you're an amazingfounder, but I also think it's such aninteresting company because it is umvery unintuitive to I think many peopleand investors in particular. I don'tknow if you faced this at the beginning,but I we were both there in 2022.There's a there's a class of companiesthat allow creation in some way when welook at your like first business beyondthe research itself. Uh, and I would put11 and Midjourney and Sunno and Hunen inthis category. And I think there's likethis overall sense of like who reallywants to do this? Um, what was yourinitial read of like how many peoplewant to make voices or what made youbelieve that was going to be muchbroader than like if I look at dubbingfor example like it's not a huge market.I think first piece was which is as youmentioned there's like a very>> it's very tricky to do both the productand the research. I'm in a in a luckyposition that I that my co-ounder and Iknown each other for 15 years. I thinkhe's the smartest person I know and hasbeen able to create a lot of thatresearch work to be able to create thatfoundation to then elevate that

00:03:51experience. But both of us are from fromPoland originally and the originalbelief came from Poland. It's a it's avery peculiar thing. But if you if youwatch a movie in Polish language, aforeign movie in Polish language, allthe voices, whether it's a male voice ora female voice are narrated with onesingle character. So you have like aflat delivery for everything in a movie.>> A terrible experience.>> It is terrible experience and it's stillyou like if you grow up the the as soonas you learn English, you like switchout and you don't want to watch contentin this way. Um and it's crazy that itstill happens until today in this wayfor majority of of of content. combiningthat and I worked with Palunteer,Michael founder worked at Google, weknew that that will change in the futureand that all the information will beavailable globally. And then as westarted digging further, we realized>> in in every language in a high qualityway. That was the starting point and theand the big the big thing was likeinstead of having it just translated umcould you have the original voice,original emotions, original inonationcarried across? Mhm.>> So like uh imagine having this podcastbut say people could switch it over toSpanish and they still hear Sarah, theystill hear Matty and and the same voice,the same the same delivery. Um which iskind of exactly what we did with Lexback when he interviewed Narendra Modiand you could

00:05:06like kind of immerseyourself in that story a lot better.Mhm.>> Um so that was the original uh uh kindof insight and um and we we then starteddigging further which is that just somuch of the technology we interact withwill will change whether this is how youcreate. It's still relatively tricky tobring voice alive. You you you you wouldyou need to go through the expensiveprocess of hiring a voice talent havinga studio space having expensive toolingto then actually adjust it. The toolingisn't intuitive to be able to do this.So like all that creation process willand should change to make it easier fornew people with keenness to bring thatto life. Then a lot of the technologywasn't possible for you to be able to umrecreate a specific voice or be able tocreate that in that high quality way.And then of course as we dived intofurther and and shifted away from thestatic piece, the whole interactivepiece is still crazy in the way itfunctions where most of us seen thistechnological evolution over last overlast decades. But you still will spendmost of your time on the keyboard. Youwill look at the screen and and thatinterface feels broken. It should bewhere you can communicate with thedevices through through speech throughthe most natural interface there is one

00:06:21that kind of started when the humanityhumanity started and um and we realizewe want to we want to solve that and Ithink now fast forward from 2022 I feellike many people will carry that belieftoo that voice is the interface of thefuture as you think about the devicesaround us whether it's smartphones wheit's computers whether it's robotsspeech will be one of the key ones but Ithink 2022 it wasn't and um and as Ifyou think about the market for thecreative side or whether for for thatinteractive side, it was like very clearit will be a huge a huge huge one.>> So even when you think about uh just theresearch part of your business and thenyou have products for at least twodifferent markets and then you have thislarger mission. A lot has changed in thelast 5 or 10 years but it used to belike a very strongly held traditionalbelief of like one must do one thingwell in a startup and there's no otherpath. like you're treating this like aninteraction company, a platform company.How did you think about sequencing likethe research and the product effort?>> Does that make sense? Or like thinkingabout new markets?>> And maybe wrapped up in that questiontoo is just like well where are we inquality on on voice as well? Because ifif I I would sort of claim like if themodels are not good enough for certainuse cases at all like it kind of doesn'tmake sense. Do product

00:07:36>> and I think that's right. It's it'salmost exactly like when we when westarted originally what we what we didwas try to actually use existing modelsthat were in the market and kind ofoptimize them for our first use case wasactually starting with combination of ofnarration and dubbing and then on thatcreative side and um we realized prettyquickly that the models that existedjust produced such a robotic and and andnot not good speech that people didn'twant to listen to it and that's wheremicrofunders genius came in where he wasable to assemble the team and and do alot of the research himself to actuallycreate new version of of creating thatwork. But like to your question, I thinkthe the way we are kind of organizedinternally and how we think aboutsequencing a lot of that was looking atthe first problem and then creatingeffectively a lab around that problemwhich is like a combination of mightyresearchers, engineers, operators to goafter that problem. And the firstproblem was the problem of voice. So howcan we recreate the the voice and likeyou say it needs to have that researchexpertise to be able to do that well. Sowe started with effectively a voice labwhich was that mission of can we narratethe work in in in in better way. It wasa combination of roughly five peoplethat were that were doing that work andthen sequence the research first

00:08:51andthen build a simple layer on top of thatwork to to allow people to use that workand then kind of expand it from therewith a holistic suite for creating afull audiobook and then creating a fullmovie narration movie dab. Um and thenwe move to the next problem which isrealization that okay we have solved thevoice great for making content soundhuman>> the first problem for that to be usefulfor us to interact with the technologyyou need to solve how you bring theknowledge on demand into that. M>> so we effectively started then thesecond team which was a second lab uh anagent lab effectively which was a teamthat would combine researchers engineersand operators once more uh which wouldtry to fix okay we have text to speechhow do they now combine this with lambsand speechto text and orchestrate allthose components together whileintegrating that with other systems tomake it easier and then similarly youknow you you you kind of expand fromlooking just at the voice layer into howthose systems work together and heretoo. You need the research expertise todo that in a low latency way, efficientway, accurate way. Um, but at the sametime, there's that product layer thatstarts forming that it's not only theorchestration that matters. It's alsothe integrations of how you link up tothe legacy systems, how you buildfunctions

00:10:06around it or how you deploythat in production and test, monitor,evaluate over time.>> Do you feel like you were creating newuse cases when you built the tools? Dopeople know that they wanted to do thisalready? Um because one argument likethat I remember hearing was like ah likeyou know enterprises don't know what todo with voice how many people reallywant to do it and then you're servingessentially like perhaps the likecreator publisher side of your business.Yeah,>> it's definitely a combination of likeinitiatives that we believe will happenin the world and then like response to alot of that like as I think back we youknow of course voice the internal voicelab or agents lab then kind of thatkickstarted so many of the other labs inresponse to the problems we started amusic lab because people wanted tocreate music with 11 labs was a fullylicensed model where people wanted touse and create speech but they wanted toadd music in a in a simple way. Wewanted to deliver that and then ofcourse that kind of came togetherthrough how do we combine music, audio,sounds. Uh we are now integratingpartner models from image and video intothat suite is how could you combine allof that in in one and a lot of that wasin response to the market saying likehey we would love this and then you willhave completely different use cases evenin that space. Let's say dabbing.Dubbing is a use case that

00:11:21we didn'tfeel there was like a a big push for forthat but we we knew that in the idealworld in the future you will be able tohave that content delivered naturallyaround the languages still carrying thatum and I still think actually thismarket will be immense because it's notgoing to be only the static delivery inmovies but if you travel around theworld and want to communicate in realtime like the full bubblefish idea fromhitchhiker's guide of galaxy this willhappen it will be like the biggest>> uh like the whole breaking down languagebarriers that are the barriers tocommunication to creation like all ofthat will break and and and that will belike the foundational real time dabbingconcept. So super excited about thatpart. And similarly on on the on theagent side, you um you you you are likesome obvious things that of coursecustomers that we work with or partnerswill will want to want to integratewhich is we want integrations with XYZsystems. But then there are like otherparts that might not be as easy topredict of as you interact withtechnology you of course want tounderstand what's happening but you alsowant to understand how the things arebeing said and bring that into the foldwhich would be something we try toprioritize on our side. So then thepeople when they actually interact withthe technology they realize oh express athing is actually so so much moreenjoyable and beneficial and helpful. SoI want to ask a

00:12:36question about thiswhich relates to quality. Um uh you[clears throat] know I work with aseries of companies where we're>> selling a product to uh the buyers aregenerally not machine learningscientists. Right. Right.>> And even the the scientific communitydoes not have the like full suite ofeval benchmarks to understand everydomain. Well there a well-known problembut I imagine for a lot of yourcustomers it's not like they like knowhow to choose good voice. So how do youhow do you deal with that problem? likeis it like a hey I make a clone and likethat sounds like me and I believe it I'mgoing to try all of these differentoptions or or you know actually are youteaching people to do eval?>> It's a great question because I thinkthere are like two big problems. One islike how do you benchmark the generalspace in audio where like you say it'slike so dependent on the specific voicelet alone like if you are training intointeractive then it's like even moretricky. Um and then the second piecewhich is as you are working on specificuse case how you select a voice. So I'lltake the f second front first which isuh we have like a voice siliereffectively with as we work with withenterprises we we we deploy that personto work with them and help them navigatethat person is like a voice coach has anincredible voice themselves and uh

00:13:52andnow we have like a team under thatperson that like will partner to helpyou find what's the right branding>> and now you have like the celebritymarketplace>> and now we have a celebrity marketplaceto like help you even get iconic talentin there like sir Michael pain thatpiece was important because of coursethe voice will depend on the use casethat you are trying to build thelanguage all of that will will have animpact of what's the right voice foryour customer base. So we haveeffectively a um a voice person helpingthose companies and some companies willbe very opinionated on what they want.So they will sometimes select itthemselves sometimes give us a brief ofhey we want a voice that soundsprofessional neutral is coming. Werecently had a company, one of the oneof the biggest European companies thatwanted uh that gave us a brief which isvery original uh that uh they wanted asrobotic voice as possible.>> Okay.>> It was counterintuitive.>> Um but for>> you like we can't do that anymore[laughter]>> almost. But we were like trying to gobackwards of like how do we do that? ButI think we we got a good result. Uh uhbut recently we had a company in inJapan where um Japan and Korea wherethey wanted to serve different voicesdepending on the customer that's callingin. They have a>> older populationing and a very youngerpopulation.

00:15:07The younger one they wantedlike one of the famous voices in themarket that's very excitable and happy.Uh and for the older one they wantedlike a calm slow speaking one. We help alot with that. So that's on the voicepiece and I do think it's going to be aa big important>> like a personalized choice and then itcan even be dynamic in a customer.>> Yes. Okay. Exactly. Exactly. And thenmaybe in the future it's like going tobe like fully depending on yourinteraction. You will like have a voicecreated as we understand the preferencesof what people want. So you know likelet's say you're in the evening and youare tired and you want a slightlydifferent or maybe not. Maybe that'slike the best uh focus time that youhave like a voice that's that's givingthat energy and probably it's adifferent when you wake up and gives youthe morning news of what's happening orwhat's the weather. So like all of thosecould be different. Yesterday we had awe had a dinner with some of our ourpartners uh and one of them the firstthing they said is like hey I have a newrequest for you. I want a New York voicewith a Long Island voice uh uh accentwhich I never knew is a thing and it'sterritory is supposedly a thing. So uhso we have that and then on the firstpiece I don't I think it's unsolvedproblem still where I think you have agood benchmarks of course in LMS I thinkin image space they are pretty good invoice space you you have of course thespeech quality but then so much

00:16:22ofwhether you like or not the speechdepends on the voice that just if youcompare a model A to model B and youserve them different voices even if thequality is very different the voiceitself can just make that sort ofdifferent we've seen this I don't knowif you know artificial analysisbenchmarks. I think they're pretty good.Just switching the voice makes thatmakes such a such a big impact.>> That's so interesting. Yeah. And Iwonder if um as you said this is uh themost dominant interaction mode we've hadfor millennia of all all of humanhistory, right? And so>> and bias is of serving but I think so.>> We're just very sensitive to it. Um andI think people are going to be verysensitive to uh their their ownpersonalization as well.>> 100%. I think there's also a third piecewhich maybe is not directly to your toyour to your note but we've alsorealized that you have uh so you havethe benchmarks you have like how do Ifind the right voice for my audience buteven the understanding of how youdescribe audio data is still lagging inthe industry like when we initiallystarted we of course went into thetraditional players for them to help uslabel not only what was said so liketranscription but also how it was saidlike what are emotions use accent andmost people just weren't able to do thatworkffect effectively because you kindof need to hear and have

00:17:37like a littlebit of a skill set of like how would Idescribe this specific delivery. So weneeded to create that ourselves. So Ithink there is that piece as well oflike how do you effectively interpretthe data of audio in a in a in a in amore qualitative basis. That's>> that's that's yeah trickier. Can youtalk about what's happening on uh theagent platform side like what ischallenging for you know businesses oreven creators that are trying to buildagents and what the what maybe what thesurprising or high traction use casesare. I think everybody's kind of awareof the idea of like agent-based customersupport but I imagine you're doing manythings beyond that.>> Yeah. So the exactly customer support isprobably the one that's like kicking offthe quickest and and that's the the onethat like we see overtaken so many usecases whether it's where I work withCisco or Twilio or Tel Digital all ofall of them are kind of elevating thatto a high extent. I think the secondexciting piece within that domain whichis happening is the shift fromeffectively a reactive customer support.I have a problem. I'm reaching out thecustomer support into more of like aproactive>> part of the experience customer support.So to make it explicit, uh we work withthe biggest e-commerce uh um shop inIndia, Misho, where they

00:18:53started workingon the customer support side where Iwant the uh refund. I want to see thetracking of the the package to actuallyhaving an agent be a front part of theexperience. So, if you go to thewebsite, you can you have um you havethe the widget, you can engage itthrough voice, and you can ask it, hey,can you help me navigate to item X, itemuh Y, or can you explain what's theright thing for me to give up for a giftfor this period of time? And then itwill actually help you based on yourquestions, based on what is on theoffer, show you those items, navigate tothe right parts of the piece, maybe goall the way through the checkout. And Ithink this will be a phenomenal thing oflike elevating the full experience wherethat's more of an assistant across thewhole thing. We kicked off our work withSquare that enable sort of businesses todo that work. Exactly the same patternstarted with voice ordering. Uh how cannow this be part of the full discoveryexperience too where you get items shownto you. You can have a lot moreexplanation which I think will be aphenomenal piece where where effectivelyfrom the beginning to the end. So that'sone category. The second one is thewider shift from static to immersivemedia where there's just so muchincredible stories and IP that todayexist in effectively one way of deliveryand now you'll be able to interact withthat content

00:20:08in a completely new way. Weuh I think one of the incredible usecases was working with Epic Games. Weworked with them on bringing the voiceof Darth Vader and Darth Vader intoFortnite where millions of players couldinteract with Darth Vader life in thegame where you had like a fullexperience of of Darth Vader in a in ain in a new way. And I think this willbe a theme across whether it's talkingto a book, talking to the character thatyou like to the whole the whole spaceshifting. And then I think the onethat's that I'm most excited about forthe world and for the shift is going tobe education where you will just be ableto have like effectively a personaltutor uh on your headphone and you likeactually study something in a in a in anamazing way. I'll give you like twoquick examples. One is uh we recentlyworked uh with chess.com. I'm a I'm ahuge fan of chess. I'm a true chess fan.Okay, great. So you can learn chess butyou can have Hikaru Nakamura or MagnusCarlson be your your teacher of how youdeliver that which is amazing or evenBotus sisters or it's like all all theplethora of different players thatengaged with that which I think is greatand then maybe a last one which is amaster class who we worked with to

00:21:23shiftfrom you can of course have the contentgo through step by step>> um but you can also have like aninteractive experience and the bestexample of that was working with ChrisBossthe FBI negotiator, one of the topnegotiators who has a masterclasslesson, but then you can actually callhim and have a practice negotiation,which is crazy.>> Yeah. Got to get that hostage out. We'lldefinitely try it.>> Yeah. Um>> can I add one more? I think the one onelast one which combines all of themtogether which I which I realized justrecently is uh which was crazy. Sorecently I went to uh to Ukraine wherewe are working with ministry oftransformation where they areeffectively creating a first agentgovernment.>> And the crazy thing is they have all ofthose>> government>> agentic government. So they want to likerechange of how they run all theministries.>> Okay.>> And it soundslike a big ambitious goal and lofty.>> No, I think the baseline is like here.So actually I'm I'm by that immediately.Yeah. And the crazy thing is I thinkthey are like so ahead in actually doingthat>> and I think they are like uh uh twoconcrete things there. One they theykind of combine all those use cases. Sothey we we are looking into how they canhave effectively customer support ofgovernment whether it's asking aboutbenefits or employment

00:22:39about process ofof how you leave uh the country. All ofthat be run through effectively adigital app. Then two how you can haveproactive way of informing citizens ofthings that might be happening. but thenhaving education system that also runthrough like this personal uh uhtutoring experience and all of that ishappening. So that was that wasincredible to see and the second amazingthing was that the way they've done it.So they have the digital transformationpiece but they have engineering leadersin each of the ministries that leadthose efforts and then bring them backto that one central piece. So that islike incredible to see and and and alsoproud to be able to be working with withthem on on that shift. that despiteeverything that's happening they're likeso>> that's amazing that's really encouragingum can I ask you a business modelquestion here because looking at thestrategic landscape um actually I havemany questions here um one of theobservations I'd have is if I look atone of these like rich voice and actionagent experiencesthere a lot of uh let's say fortune 500global 2000 leaders who listen to thepod uh they I think a lot of them aregoing to buy the idea of like I wantthis amazing um automatic like real timeavailable 24/7 every language experiencefor my

00:23:54customer that's consistent andhigh quality. The ways I might get thereinclude working with a Palunteer or alarge consulting firm, uh, working with11 or a like platform technology companyor or like an open AI or something,right? Let's talk about that. Uh, orworking with a sort of more usecaseoriented company like Sierra, right?[clears throat] How do you think abouthow people are making that decision orhow they should make that decision? theso so my past is also in Palunteer so Istarted exactly kind from from that sideand we do blend a lot of the forwarddeployed engineering inside of thecompany too as I think about the kind ofour offering and and the customersmaking that choice if you're lookingjust as as a like onepointed solution uhand only that one then likely we aren'tthe best choice if you are looking todeploy that across a plethora ofdifferent experiences so be it customersupport but then you also want internaltraining then you might want to elevateyour sales sales part and actuallyincrease the top line with newexperiences of how you engage customersbeyond that kind of reactive piece. Mhm.>> Um then it's a great platform to buildand then we effectively as we engagewith customers combine that platformwork with uh with our

00:25:09engineeringresources to help those companies deployon that or which we also seeincreasingly in um in Fortune 500s G2G2000s where they will want to buildparts of the things themselves becausethey already have a lot of theinvestments in that platform while thenengage us on some of the the new onesand combine those and and and I thinkthat our model and the way it'sdifferent to to a lot of the use casespecific ones is that our platform isrelatively open where you can use piecesof that platform and not all of them umfor for those different use cases.Palunteer of course will will or or someof the consulting companies will have alot more resources to go in the widerdigital transformation journey. In ourcase, it's like very specificconversational agents. It's like if youare looking for new interface withcustomers, that's the the best way. Andum and companies like Sierra phenomenalof course on on how they are thinkingabout the the specific pointed uh uh usecase and the maybe the other piece is uhlike as we think about our workdepending on how you are what you areoptimizing for. So we we have a lot ofinternational partners. If you have likea a wider geographic user base, great.That's what we optimize for. Our voices,our languages, our support forintegrations internationally are just

00:26:24somuch broader. There's frequently a piecethat you will look into depending onyour exact scope, this will be this willbe a big factor. But I would summarizethat if you're looking for a solutionacross the set of different use casesthat you want our engineering help anddeploy that, then we are the rightsolution and probably the best solution.I want to talk a little bit about maybelike opening eye and the foundation LLMfoundation model companies. One of thereasons a lot and I called this podcastno priors is because we're like okaypeople are making a lot of assumptionsall the time about how the market isgoing to work and lo and behold likemany of those assumptions end up beingnonsense actually and you you you can'tyou have to very much decide your ownnarrative at this point in time. Ithink, correct me if I'm wrong, like in2022 and 23, you probably heard a lot ofpeople say like Google can do this andOpenAI can do this and like why do youget to persist working on voice anywayas a general capability? What's theanswer? That also adds adds a kind ofanother element to to to that the coupleof the other previous questions wherewhether it's agents work whether it'sthe creative work deploy the value inthose in those work you need a verystrong product layer you need theintegrations you need to help peopledeploy the work which is the most commonpiece but our superpower and our

00:27:39focusfor a long time was building thefoundationalmodels to actually make that experienceseamless and as I think about a lot ofthe companies in the market they willoptimize for a lot of other things andthat that will be like thedifferentiator um in our case where wewill make the whole experienceespecially with voice seamless humancontrollable in a in a much better way>> and so fundamentally you would arguethat like the labs just aren't going tofocus on this and haven't>> exactly so I think most of thosecompanies and that's the thing about thelong term it's going to be incredibleresearch and incredible product thatmeets customers where they are and workbackwards from there. I don't think thethe labs will focus on building thatproduct layer that's so important.>> But I think the you know part of thequestion that you're asking is like howuh or and and and why they haven't doneeven the research part>> to the quality that that we've been ableto as here I'm also biased but we arehappily beating them on benchmarks withtext to speech or speech to text or theorchestration mechanisms and here creditto my co-founder and the team uh thatthey've been able to do it. It's just amighty researchers just continuing theirwork. But I think the main part that Ithink is different

00:28:54in audio space isthat you don't need the scale as much asyou need the architecturalbreakthroughs, the model breakthroughsto really to really make a a dent. Andum and we've been able to do that coupleof times and I think the number ofpeople doesn't matter but the peoplethat you do does. We think there's maybe50 to 100 researchers in audio spacethat could do it. We think we haveprobably 10 of them um in the companythat um that are some of the best ones.And I think this like obsession of justthose people working across and thenactually giving the full focus on thecompany on making them actually work onthat and bringing their work toproduction, seeing how the usersinteract back was was so important. Sothat's the that's I think how we how webeen able to create models um betterthan some of the the the top companiesout there. But you know the truth isit's like to large extent is why theyweren't able to do it is also like aninteresting we we don't know it's likeit's a it's a it's uh they are like theyhave such an incredible talent theretoo.>> How do you think at the same time aboutum like open source models? anyone youask in the company I think will say thatsame and that's like a narrative wethink about it's in the long term modelswill commoditize or the differencesbetween

00:30:09will be negligible for some usecases they will still matter for mostlike the long like the most use casesthey they won't um>> and they'll be broadly available and>> they will be broadly available exactlyand we don't know where that is whetherit's two years three years four yearsbut it's it's going to happen at somestage then of course you will have afine tuning layer that will matter a lotum on on top of those models but likethe base models I think will get prettygood. Um and that's why for us theproduct piece is so important from thecompany perspective but also from thevalue perspective because if you have amodel that's great but to actuallyconnect your business logic andknowledge to um to be able to have theright interface for creating a an ad foryour work or a completely new materialthat's uh that's a very differentexercise um but open source models aregetting if I split into two like more ofthat async content narration I thinknarration is pretty much open source isgreat, commercial models are great, thedifferences are are getting smaller onthe on the out of the box quality. Whatmost of the models haven't figured outum and I think we we wear is how to makethem controllable.>> So that's the kind of the narrationpiece I think the whole interactionpiece of how you orchestrate thecomponents together whether that'scascaded speechto text and

00:31:24lamb text tospeech approach or whether in the futureit's a fused approach where you trainthem together. I think this is is goodfor customer support or customerexperience but it's still away from likeconversation like we have and likepassing that Turing test. So I thinkthis is still like a at least a yearlike a within a year and then you'llhave like real time dubbing kind ofvariation of like real-time translationconversation and I think that's maybelike more two years within two yearsaway. You know, a very uncomfortablebelief that I I feel comfortable havingthis belief, but I think is uncommon inthe market right now is that actuallymost advantages in technology, like theycould they could last you a year or theycould last you 10, but they're not likeinfinitely defensible. And if you thinkabout that from a model qualityperspective or a product perspective,they allow you to like serve thecustomer better and build momentum andbuild scale for some period of time. Andactually that's really powerful overtime, right? But it's not like a cleanforever answer. And so I think thatmakes I don't know business people andinvestors uncomfortable.>> And I mean it's it's it's it's very trueas well. [laughter]>> The way we I mean the way you thinkabout it research is head start.

00:32:40Thisgives us we can give advantage to thecustomer earlier and it's six 12 monthsof advantage. That is also a way for usto build a right product layer for youto get best of that research. Frequentlywe do that in parallel. So the momentthe research is out there you have theproduct because we know our initiativeswe know what the product is. That'sright. So you have research product inparl that extends that. But the kind ofthe thing that will really give thatlong-term value is the ecosystem thatyou create around whether that's the runand distribution whether that's thecollection of voices you can have thecollection of integrations you can buildthe workflows that you can build. Umlike I think that's that's the way wekind of sequence that in our mind thatresearch product ecosystem that we builtand um and research all it is is a is ahead start and being able to likeaccelerate the future a little bitcloser. I think that's a really powerfulinsight especially if you know theresearch and the research team and thecompany team believe that as wellinternally>> it's it's it's I think the piece that wewhat's like interesting for us is u andI think this is like the the the bigquestions for all companies that doresearch in product is do you wait forresearch or do you do like a a productchange uh or even not only researchproduct companies like do you wait forsomeone else to do the research becausethe timeline

00:33:55for that isn't clear is it3 months 6 months 12 months don't knowexactly what it will do which is thehard choice of like do I invest intoproduct layer or do I just wait more forthe research so like in our case weinternally let all the product teams theresearch initiative so we can paralyzethat work uh but we don't hold them thatif if a product team thinks we shoulddeliver value to the customer by doingsomething different they can and roughrule of thumb is like three months if wethink it's going to be longer than threemonths we will probably build it if it'sless than that we probably won't>> can you talk about some of the researchthat you're doing now and and how youthink about like the cadence of deliveryand what's worth working on.>> We have now number of of differentinitiatives across the audio space andthere are there are kind of two bigbuckets and and and roughly they willrelate to that creative and agent side.On the creative side what this means uhwith the texttospech models that arecontrollable. Uh we then addedspeechtoext model that transcribes in ahigh accurate way but across a lowresource languages as well. So coveringalmost 100 languages. then created amusic model, a fully licensed musicmodel. Um, and as you think about thefuture is how those models will alsointeract with some of the visual space.So that's uh a lot of effort and how youcan get a best of audio and thenpotentially combine that with existingvideo that you

00:35:10have to to to really havethe best delivery. And then on the agentside, it's of course how you optimizethe real-time speech to text, real-timetext to speech. We just released ourspeechto text model scribe v2 which isunder 150 milliseconds 93.5% accuracyacross the top 30 languages on on flaresand it's only top 30 here because weserve so many others but most of thepeople don't so uh so it's uh so it'sbeating beating the the all the modelson on benchmarks but as you think aboutthe future it's also the orchestrationpiece of how you bring speechto text lmand text to speech we are releasingwe'll be releasing over the next coupleof months a new orchestration mechanismthat will lower the end to end part umwe think in a great way. But secondthing which is what is so hard is it'snot going to only allow you to combinethose pieces but add also the emotionalcontext of the conversation so you canactually uh respond with the model andwe think and more expressive in a in abetter way. And in the future andsomething we're investing is paralyzinga speechtospech more fused approach aswell. And of course depending on the usecase if you are like enterprise reliableuse case the cascaded approach is theapproach for the next year too>> has more structure yeah>> more structure you have more uhvisibility into each of

00:36:25the steps it'sreliable you can I'll call tools ifyou're think more expressive and canhallucinate speech to speech might bethe choice and maybe over time you'llsee them the kind of go one over overanother depending on the on the industrybut that's like a huge investment on ourside which is where the foundation ofall the platform and and and the mainpart that we are continually investingin is is is kind of plethora ofdifferent models that combine the bestof audio with some of the best of theother modalities together.>> I want to take uh our last few minutesand ask you a few questions about justthe future that I think you'll have areally good point of view on given youthink about voice and audio all thetime. What do you think of AIcompanions?>> I think they will be a big thing andexist in a big way. not something I'mpersonally excited about or somethingthat we spend much uh time on but uh Ithink the whole line of of what's aassistant companion character that youenjoy as part of experience will kind ofblurry and blend to a large extent>> they can be very common but you're notlike enthusiastic personally about it>> I'm more excited about like more of umthe Jarvis version of that or like likemore of like I have a super assistantsuperpower. It's like>> versus the social

00:37:40version>> versus the [clears throat] socialversion that's like I think it's it justwould be like such an incredible unlockand and it also like is in a somethingblending in that person context like Iwould love to start the day and likesomeone that understands me and likestart and tell me what's like relevantto me and open the blinds and then liketell me about the weather and thesunshine is and play music straightaway.>> It's going to happen.>> It's going to happen. That I'm excitedfor. I think the companion um use caseswill will will mention solvingloneliness and in that part I thinkthat's one way maybe there are likedifferent ways of engaging people back Ido I do think there will be like aninteresting future even if you thinkabout education where you will havesuperpower with learning from AI tutorsbut I think on the flip side of that andI think this will like that's mypersonal take you will have educationgood percent of time spent with AItutors but then explicit percent of timespent and without any technology humanto human.>> So you can you can kind of learn thatpart too.>> Yeah, I think this is the correct model.Um both in terms of like emotionalguidance and coaching and um you know uhguard rails, right? As well as likepeerto-peer. Um>> exactly. What do you think about umdictation or what happens in

00:38:55terms ofhow we like control uh technology thatisn't necessarily personified as well>> or does it just all become personified?>> I think not all personified. I thinklike some you know communicating withinoven and and home probably will likestay pretty static and>> or code I might just>> Yeah, exactly. like you don't probablyneed that much of of like additionalemotional input>> but uh but I think it's yeah it's goingto be huge part where like in a way whatI hope will happen is you will haveability to like stay more immersed inthe real life with the devices goinginto back into the pocket back into someversion of a um attached element umassuming that's that's in in the rightsetting and um and that kind of acts onyour behalf and in many ways like let'ssay dictation it's as Karpati saysdecade of agents. Let's let's let's callit a decade. Then you'll have a decadeof of robots. If you are interactingwith robots, of course, voice will bethe input and the output as one of thekey interfaces. So, you will need thatdictation as a as a huge part. Butsimilarly,>> I think the robot's going to bepersonified.>> Yeah. 100% 100%. Yeah. No, like I thinkI think most of the use cases will bepersonified.>> Okay. Last one. What's like one thingthat you've seen already exist today orif you project out

00:40:10a few years willchange about how we interact withcontent maybe it's like personalizedvoice content or um just somethingpeople are going to do with with AIvoice that they don't do today or thatnot everybody knows about.>> I think this still the biggest one thathasn't yet kicked into the the thesystem is like how education will bedone. I think this go like I thinklearning with AI will with voice whereyou it's like on your headphone or in aspeaker it's just going to be such a bigthing where you have like your ownteacher on demand and who understandsyou very personified and kind ofdelivers the right content through yourlife I think this will be one of thebiggest use cases uh uh and I don'tthink it happened yet I think we havesee of course some of the commercialpartners but like schools universitieshow that's deployed in a safeguarded wayin a way that like supports the otherpart of the education, the social partof education. I think all of that willwill evolve and maybe there's a coolversion of that where you have likeRichard Fineman or Albert Einsteindeliver those lecture notes or otherteachers that you love. It's it's itwill be sick.>> It's a great note to end on. Thanks fordoing this, Marty.>> Thanks so much.

00:41:20>> Find us on Twitter at No Prior Pod.Subscribe to our YouTube [music] channelif you want to see our faces. Follow theshow on Apple Podcasts, Spotify, orwherever you listen. That way, you get anew episode every week. And sign up foremails [music] or find transcripts forevery episode at no-briers.com.