Intro0:00
Okay, and I think we're live. So welcome to another Latent Space Lightning Pod. This time very, very timely because we're actually collaborating on a voice AI course with Kwindla from Daily. Welcome.
Hey, thanks. Glad to be here and really excited about the course.
I am too. As you know, I'm a fan of voice AI in general. I think, um, you know, I, I was one of the first few to like sing with ChatGPT's advanced voice mode, and we actually didn't really know if that would ever come out as an API, and now it's an out as an API.
And since then, there's just been a proliferation of voice products. I think the voice models keep improving. You've come out with Pipecat, which started as kind of like an open source orchestration framework that was from the Daily folks, but now is being adopted by a lot of, uh, others.
So what's your inspiration in, in doing this voice AI course?
I answer the same questions so many times, and I love doing it because this stuff is so much fun and there's so much like interesting engineering stuff to rabbit hole on. But it seemed really clear that it would be awesome to get a bunch of people together, talk about all this stuff, use some of the materials we've been kind of building over the last few months, and kind of see where it goes.
And the inspiration was Hamel and Dan's like LLM Woodstock course last fall.
Mm-hmm.
Which, you know, I know you were part of. I got a lot out of it. It was super fun, and lots of companies came and jumped in and helped everybody out, like with seeing all of the tools that you should know about if you're building stuff.
So I was like, "Well, let's try to do that for voice AI."
Yeah, totally. So I think Hamel's course honestly took-- was the biggest thing for Maven 'cause now Maven's like the default-
Yeah
... voice platform. Even though there are many, many like it, Maven seems to be popular. Um, and, I mean, okay, so you have a lot of partners and maybe let's just go over like the syllabus, right? Like what do you think that people need to know in voice?
Syllabus1:59
I think of it as three big things. The first is like, what's the landscape? What models do you use? Why? Like what are the best practices for low latency, real-time conversational multi-modality? And it turns out that if you're building like agents that are multi-modal, multi-turn, real-time, it's just a totally different shape of programming prob- problems and best practices than most other kinds of AI development even.
There's overlap in like how do you prompt stuff, how do you eval stuff, but the, the code just looks different. And so kind of best practices and models landscape, that's the first big thing. The second thing is like how do you run this stuff in production?
Again, because the shapes of the code are different, deploying a real-time agent to production and then scaling it and making sure you can run evals and you have monitoring and observability, like a whole new set of stuff there.
And to some extent, we're all figuring out this stuff as we go along, but there are some emerging best practices. There are people doing this stuff in production, so I think we can, we can jumpstart people's path from prototype to production.
And then the third thing for me is what's next. And this is a little bit selfish because like this was-- first I-- the first thing that I was like, I, I wanna do in this course that I don't know enough about is I wanna do voice-based programming.
Ha.
And I know you have thought a ton about that, so I was like, "I have to pull swyx into this." But the other things in my like what's next bucket are what are the models that are coming out, you know, later this year from the big labs, from the open source community.
We'll have people from OpenAI and from DeepMind and from NVIDIA and from AWS doing office hours. What's happening with speech-to-speech models and APIs? Super hot topic, not widely used yet in production. Uh, we'll talk about that a bunch during the course.
We'll also talk about that sort of in, in what's next. Running some models locally. I'm really excited about like models getting good enough at medium size that you can run them locally, like on device or on robots or on local laptops.
And then real-time video. Like I really think real-time video is gonna hit the same inflection point that voice did by the end of the year or early next year. So we've got a couple of, uh, folks who are really focused on real-time video participating in the course and doing some credits and some office hours.
What do you think-- Sorry, w- w- there's, there's a lot there. Just double-clicking in reverse order, like what is real-time video really mean? Are, are you talking to like an avatar?
Real-time Video4:19
Yeah. Yeah, yeah. It's like, just like you can now have a voice-based conversation with an LLM-
Like Tavis
... you know, from, from folks like Tavis, you can have a video conversation with a transformer-based model or a collection of transformer-based models behind an API. And it's, you know, it changes the dynamic if you have that whole video channel.
Yeah. I, my, I, I tried this out and like it's... Who is this? I guess it's Charlie Witty. Okay. Yeah, yeah. I, I tried this out and I can't believe that there's a ton of traction behind this, 'cause it's kinda very robotic.
We ser- we are the net- Daily is the network layer for Tavis and some others, and I can tell you there is a ton of traction behind it.
Yeah. Um, my sense is like it's, it's, uh, for coaching and...
What are the verticals that kinda use this stuff?
Yeah. There's a bunch of coaching and kind of enterprise teaching and learning education stuff in-
Yeah
... inside big companies.
Yeah.
I also think we are starting to see the beginnings of some, some social stuff. I mean, a little bit to everybody's surprise, the voice AI monetization pull actually came from enterprise and B2B use cases. So a little bit to everybody's surprise, the inflection point with voice AI in terms of monetizable use cases came on the business side.
Uh, it's like telephony and customer support and, uh, you know, a bunch of like vertical SaaS with voice on top. I think real-time video may actually hit on the consumer side first. When it, when it's right, when it's on the right side of the uncanny valley- It's really, really compelling, and I think we're just starting to see some of that.
Yeah. Um, I'm starting to see, like, I think there's another platform, not Tavis, but, um, someone else doing more consumery video conversations. Are, are we talking about AI companions or is there some other product consumer that is interesting?
Well, I think we're gonna have friends that are video in all our group chats.
I see.
And the, the next TikTok is gonna be not just hyper-personalized recorded content, but hyper-personalized interactive content.
Yeah. Yeah, yeah. And, and, you know, for what it's worth, this, this is more, much more constrained domain for video than generalized video generation, which is a whole different thing. Like, here you're, maybe you're taking an avatar, maybe you're taking a clone.
You have, like, some kind of bone structure thing, and then you're animating that with, um, some expressions and, and the mouth movements, and that's about it.
Yeah. There's a whole bunch of different approaches. You can go, like, single image to generate an avatar. You can take source video and you can kind of fine-tune a model. You can do rigged characters and have an LLM drive the rigged characters.
So people are experimenting in sort of every way with video. It feels early days, but-
Yeah, yeah
... it's happening fast.
It's happening fast. Okay, so then recursing up, uh, there's just sections about-- You mentioned some open models that are just, you know, voice stuff. What are the main things and trends you're eyeing here? There's, there's sort of the end-to-end speech models like the KyuTime Mushis.
Open Models7:17
There are, like, the couple guys doing-- Never, never mind. I, I, I don't, I don't really know what the Parakeet and Dia type models are optimized for except for, for like, I guess podcast generation.
Well, there's kind of the opposite ends of the spectrum there. Like, Dia is this amazing project by, like, one and a half Korean students. And it took, you know, it took at least my Twitter feed by storm like two weeks ago, and it's all the way on the, like, dynamic end of the spectrum.
Like, it's totally unhinged. And then Parakeet is NVIDIA's new speech model, speech transcription model that's number one on the leaderboards, and it's just, like, very enterprise tuned, like really, really rock solid, reliable at fairly small number of weights, like 600 million parameters or something.
So you've got, like, people building on the, "I need to use this in enterprise production" side of things, and you've got people building open source models that are like, "Let's just see how, let's just see how it goes.
Let's see what we can do."
Yeah, sorry. I-
And they're both really valuable
... I made the link because I think the Dia guys said that they borrowed from Soundstorm and Parakeet in some ways. I, I don't know. I, I don't even know if there's code published, but they just took the papers and implemented them, which is the hard part.
Yeah. NVIDIA does a great job publishing papers. It's super impressive if you follow NVIDIA research. They have a lot of GPUs, those researchers. They write good papers.
Definitely do. Um, yeah. What, what else is, um, interesting among all this? Like, is this, is this sort of a Pipecat centric? Are we gonna spend a lot of time on, I guess, this like guardrail scripting? I think, I think really, like, building smart voice is the-- A really quick demo is super easy, but then adding capabilities to it, making it reliable is really hard.
That's a lot of the genesis of the course is like how do you get from a demo to something that you could actually use in production?
Mm-hmm.
And there'll be a lot of Pipecat just 'cause I'm doing a lot of the work to kind of glue the course together, but it's not Pipecat centric or Pipecat only. There are a bunch of folks who are participating in the course who have their own stuff going on, like Freddy from Hugging Face is gonna talk about Fast RTC and hang out in the Discord and kind of be around.
I really love that work. Uh, Sean from OpenAI is gonna talk about their WebRTC and 1-800-CHATGPT layer, and so is Dominic, gonna talk about the OpenAI Agents framework, which they did a really elegant like voice layer on top of that original Agents framework in the most recent release.
Mm-hmm.
We have like VAPI, which is a, a kind of higher up the stack, like everything together platform that I recommend often to a lot of people who start with Pipecat but kind of want more batteries included. There's a brand new competitor VAPI, I think it's fair to say, called LayerCode, uh, that has $1,000 of credits for students in the course that I got an early demo of.
I think you know those guys, and we were like, "Oh, well you should, you should hang out in the course too. You've got good ideas."
Do I know those guys? Have we-
They told me that you know them.
Uh, I, I may have. I, I know a lot of people, but
Yeah, I know.
Maybe, maybe I met them individually without knowing the name of the-
LayerCode is very new, so actually that, that may be the case.
Yeah.
It's like Damian Tanner.
Yeah, okay. I do know him. Yeah. Uh, interesting, 'cause when Damian was at Deepgram... Was it, was it Deepgram? Well, he, he went over to LayerCode and I was like, "Oh, this like seems like an established company." He did not say that or there was no indication that this is a new company.
I think it's still in private beta. I think it's that new, which is fun.
Okay.
There is so much stuff launching in voice and I, I've really tried to pull in like lots of people that I know because, like, you can't cover everything yourself. I mean, you do an incredible job covering like everything across all of AI, but we, we can't all be you.
Well, there, there should not be. You know, I think the world needs specialists, and then we need to be able to lean on their expertise and, and then, you know, build products. Like I, I-- There, there should not be that many of me who waste their time being journalists and, and wide on everything and shallow on, on a lot of them.
Voice & LLMs11:15
So, okay. I'll, I'll just dive in, double-click a little bit into how voice currently interacts with LLMs because you mentioned this thing about how the Agents framework from OpenAI adds a voice layer, and I think that's something that Pipecat does as well.
Everyone wants to stream. Everyone wants to, uh, do semantic detection. It's really hard, and I don't know if we have the right interfaces to do this without adding a ton of latency that makes it very unnatural.
We basically spent all of 2024 trying to get to the 80/20 point on, you know, the half a dozen hard problems here that you were outlining. It's like how do you do the low latency networking of the audio to somewhere in the cloud where you can do the orchestration?
How do you do turn detection so you know when the user's done talking and the LLM should start talking? How do you do context management that's kind of optimized for the real-time multi-turn- You know, your LLMs are stateless, right?
So you have to maintain state, and that is a whole interesting rabbit hole to go down. How do you make it so that you can swap out different parts of the system because models are evolving so fast in, in every category, the models are leapfrogging each other, so you don't wanna only be tied to a single model?
And then how do you do things like function calling, which were never designed to even work in a multi-turn context, much less happen asynchronously or happen at low latency or, you know, whatever, how- whatever hole you have, you have a round peg or vice versa with things like function calling and structured data output and all that stuff.
So I think a lot of us who were working on voice last year were like, "Well, let's just get it so it works pretty well. And once we have all the pieces working pretty well, there's gonna be something really interesting here."
And it felt like that sort of happened at the end of last year, and now there are these kind of 2025 problems, which is how do you write really great evals? How do we give really great feedback to the people at the labs who are training these models, who, you know, a lot of what we're doing, it's become clear is pretty out of distribution.
The multi-turn, multimodal native audio stuff is all like it's frontier of the frontier models. And how do you like level up from the 80/20 point for things like turn detection? Like we have a-- we'll, we'll talk a lot in the class, I hope, about this open source, open data, open training code, uh, native audio turn detection model that a lot of folks, including me and the Pipecat community, have been working on.
So there's like the 2024 problems, and now there's a new set of higher, you know, in the Maslow's hierarchy, 2025 problems.
Yeah. Yeah, I, I remember last year talk, um, asking you publicly about the semantic turn detection. Now, like assuming that is solved, which is, you know, it's not, not really, but it's, it's getting better. Do people have... Do, does this Pipecat or any other framework support thinking while the other person's speaking?
So there's a question about what the model can do, and there's a question about what you can put together from the orchestration pieces. And none of the, none of the SOTA models can kind of think while also taking input.
The, that Kyutai Mochi model you mentioned, which was my, my, my favorite academic paper last year, is a step towards a truly bidirectional streaming in both directions, thinking all the time LLM. That work has not been kind of scaled up to, to, to, you know, large LLM size.
So what we try to do in the Pipecat world is we do a bunch of things in parallel in separate Python tasks and have abstractions for kind of putting those things together or having those things interact in useful ways.
And you can do that with like state machine abstractions, you can do it with like multi-agent abstractions, all of which I think are just kind of abstractions on top of the idea that you often actually want multiple LLMs talking to each other.
One of my favorite Pipecat demos is this, uh, two connections to Gemini multimodal live-
Yeah, talking to each other
... playing a word guessing game with you.
Yeah.
One, one LLM is the judge and the moderator of the game, the other LLM is the player you're cooperating with. Increasingly, people are building stuff like that, and increasingly you want to have great abstractions in your agent framework to do that kind of thing.
Yeah. Amazing. And then, uh, one- always want to have one thing about telephony and other integrations. What are the trends in people accessing these things? So you have got the, you got the 1-800 ChatGPT, you've got, um, the little microphone in every, next to every text box.
Telephony15:30
What else are people doing? What's hard?
I mean, 99% of the monetizable voice AI use cases today are telephony.
So like Twilio?
Like Twilio. A lot more people are using Pipecat with Twilio than with anything else-
What's, what's the number two with Twilio?
... whatsoever to see. I'm sorry. Say that again.
What's, what's number two after Twilio? I'm just-
After Twilio, it gets more regional.
Okay.
So like Plivo in India, where Twilio doesn't have great coverage, like can't, for regulatory reasons, can't do phone numbers and stuff. So there's a bunch of places all over the world where there are regional winners.
Yeah.
Twilio has great infrastructure. It's hard not to, you know, buy with Twilio if you, if, if you're in a place Twilio serves.
Right. So like the web voice API connections, like the retails of the world, they're much smaller than compared to the Twilio ones, where you just have the phone line.
Yeah. And all of the, you know, the, the VAPI Retail, Bland, Sierra only, like lo- I, I sort of mixed a bunch of categories there, but, but most of those, even if you've tried them from their website, most of their revenue, almost all of their revenue comes from inbound or outbound telephony-
Yeah. Yeah. Amazing
... today.
Yeah, yeah. Um, I'm always curious, like just because, um, I would like to see more integration, more voice everywhere, you know, like I, I, I'm a supporter of, of like the B and like Limitless and like sort of just ambient listening devices.
I briefly looked into trying to have my, make my own home speaker thing that just listens at all times and, you know, takes notes. It's, it's surprisingly hard. Like even if you take the hackable like Raspberry Pi based, um, what is it called?
Hardware17:06
Home Assistant thing, it's very elementary. Like you can bas- you can just about tell it to turn on your lights.
Yeah. It feels like, uh, compiling a Linux kernel in like 1999. It's like that, that's the level where we are for the home automation hardware do it yourself stuff. But I do think hardware is super interesting and hardware with voice is super interesting.
I, I wanna do-- I, we actually haven't scheduled it yet, but you know, you and I should like recruit a bunch of people to do a hardware session for the class.
I mean, I, I can get the, the B folks anytime. Um-
I've been playing with little robots. For the people who are watching on YouTube, this is my voice controlled, latest voice controlled robot with a camera on it.
Okay. What does it do?
It drives around and it recognizes you, and if you've had a conversation with it before, it remembers what you talked to it about.
No way. Okay, you gotta do a little demo of that. That's awesome.
It, it started out as a hackathon project at one of the, one of the Antarctic hackathons.
Didn't you get the, the dogs too? You got the dog?
Yeah, I got the dog, but the dog, we need to get Sean, uh, Sean Dubois' help on the dog because that's a sort of ESP32 and I just haven't managed to like compile a WebRTC stack for it. But Sean has one.
The robot I just held up is a Raspberry Pi, so you know, it's, it's easy mode. It runs Python.
Yeah, yeah. Excellent. Excellent. Um, cool. Um, that's a, you know, that's, that's a brief survey of the course. There's, there's a lot of other stuff. Anything else that you think we should highlight but haven't?
I just wanna highlight how much of a community effort we want it to be. So like we want people to hang out in the Discord, trade ideas like the-- It's a course, but like in my mind, the course is just a Trojan horse for, for a community, like a festival.
Community18:47
It is. It is. And, uh, you know, you'll be at Google IO and the AI Engineer Summit at World's Fair. Got that right?
Yeah.
And, uh, doing, doing a lot more voice, uh, vo-
You know, we should actually credit AI Engineer World's Fair as the other thread that led to this. Like I used AI Engineer World's Fair as a forcing function for-
Oh, yeah
... writing this like 16,000-word guide on voice AI.
I have the book too. It's upstairs, but yeah.
And the, like the feedback on that was like, "Let's do reading groups about this. Let's talk about this in the Discord." And I was like, "That's interesting. Let's, let's do that." And kind of that pinged around and-
Mm-hmm
... became this course.
Yeah. Excellent. What is the Discord? Is it, is it the Pipecat Discord?
I think we'll just do it in the Pipecat Discord 'cause like that way I don't have to set up another Discord, but-
Yeah
... we'll figure it out in the next 24 hours.
Too many Discords. Too many Discords. All right. Excellent. Um, well that's it. Uh, you know, everyone go sign up for the course and I'll put the link in the show notes. Any other, uh, any last calls to action or words of wisdom?
Ping me if you have an idea for a session you wanna run.
Wrap-up19:53
Wow. Okay. You're still taking sessions. You already have, you already have too many, man.
We have 28, so yeah, we do have too many but, you know.
What are you going for here?
Um, no, good, good, good. Uh, I mean, you know, I think it's a special time. I think voice is definitely one of these. The timing is right to work on these things because voice is here but it's not evenly distributed and I think more people can, can, can enhance their existing AI applications with voice just 'cause it actually makes it more, much more magical than-
And it's the future
... just typing it in the chat window.
I mean, uh, U- UX is gonna be, you know, 50%, 60%, 75% voice in the future. I 100% believe that and I did not believe that, you know, two years ago. Uh, but the trend line is just really, I think-
Yeah
... clear.
I mean like this is our first year of having a full voice track, you know, and, uh, and if I were to subdivide it, I don't know how I would subdivide it, but clearly, you know, this is growing.
Excellent. All right. Thanks, Kwindla.
Thank you. Bye-bye.






