LALatent SpaceMay 6, 2025· 20:51

Voice AI Masterclass — Kwindla Hultman Kramer and swyx

Shawn Wang and Kwindla Hultman Kramer announce a voice AI masterclass course, diving into the landscape of models like Dia and Parakeet, production deployment challenges, and future trends such as speech-to-speech and real-time video. Kwindla explains that telephony (Twilio) drives 99% of current monetizable voice AI, while open source frameworks like Pipecat enable low-latency multi-modal apps. The course covers turn detection, context management, and evals, with 28 sessions featuring partners like OpenAI, Google, and NVIDIA. The goal is to jumpstart builders from prototype to production, with real-time video expected to hit its inflection point by year-end.

  1. 0:00Intro
  2. 1:59Syllabus
  3. 4:19Real-time Video
  4. 7:17Open Models
  5. 11:15Voice & LLMs
  6. 15:30Telephony
  7. 17:06Hardware
  8. 18:47Community
  9. 19:53Wrap-up

Powered by PodHood

Transcript

Intro0:00

Swyx0:01

Okay, and I think we're live. So welcome to another Latent Space Lightning Pod. This time very, very timely because we're actually collaborating on a voice AI course with Kwindla from Daily. Welcome.

Kwindla Hultman Kramer0:14

Hey, thanks. Glad to be here and really excited about the course.

Swyx0:18

I am too. As you know, I'm a fan of voice AI in general. I think, um, you know, I, I was one of the first few to like sing with ChatGPT's advanced voice mode, and we actually didn't really know if that would ever come out as an API, and now it's an out as an API.

And since then, there's just been a proliferation of voice products. I think the voice models keep improving. You've come out with Pipecat, which started as kind of like an open source orchestration framework that was from the Daily folks, but now is being adopted by a lot of, uh, others.

So what's your inspiration in, in doing this voice AI course?

Kwindla Hultman Kramer0:50

I answer the same questions so many times, and I love doing it because this stuff is so much fun and there's so much like interesting engineering stuff to rabbit hole on. But it seemed really clear that it would be awesome to get a bunch of people together, talk about all this stuff, use some of the materials we've been kind of building over the last few months, and kind of see where it goes.

And the inspiration was Hamel and Dan's like LLM Woodstock course last fall.

Swyx1:17

Mm-hmm.

Kwindla Hultman Kramer1:18

Which, you know, I know you were part of. I got a lot out of it. It was super fun, and lots of companies came and jumped in and helped everybody out, like with seeing all of the tools that you should know about if you're building stuff.

So I was like, "Well, let's try to do that for voice AI."

Swyx1:35

Yeah, totally. So I think Hamel's course honestly took-- was the biggest thing for Maven 'cause now Maven's like the default-

Kwindla Hultman Kramer1:42

Yeah

Swyx1:42

... voice platform. Even though there are many, many like it, Maven seems to be popular. Um, and, I mean, okay, so you have a lot of partners and maybe let's just go over like the syllabus, right? Like what do you think that people need to know in voice?

Syllabus1:59

Kwindla Hultman Kramer1:59

I think of it as three big things. The first is like, what's the landscape? What models do you use? Why? Like what are the best practices for low latency, real-time conversational multi-modality? And it turns out that if you're building like agents that are multi-modal, multi-turn, real-time, it's just a totally different shape of programming prob- problems and best practices than most other kinds of AI development even.

There's overlap in like how do you prompt stuff, how do you eval stuff, but the, the code just looks different. And so kind of best practices and models landscape, that's the first big thing. The second thing is like how do you run this stuff in production?

Again, because the shapes of the code are different, deploying a real-time agent to production and then scaling it and making sure you can run evals and you have monitoring and observability, like a whole new set of stuff there.

And to some extent, we're all figuring out this stuff as we go along, but there are some emerging best practices. There are people doing this stuff in production, so I think we can, we can jumpstart people's path from prototype to production.

And then the third thing for me is what's next. And this is a little bit selfish because like this was-- first I-- the first thing that I was like, I, I wanna do in this course that I don't know enough about is I wanna do voice-based programming.

Swyx3:17

Ha.

Kwindla Hultman Kramer3:18

And I know you have thought a ton about that, so I was like, "I have to pull swyx into this." But the other things in my like what's next bucket are what are the models that are coming out, you know, later this year from the big labs, from the open source community.

We'll have people from OpenAI and from DeepMind and from NVIDIA and from AWS doing office hours. What's happening with speech-to-speech models and APIs? Super hot topic, not widely used yet in production. Uh, we'll talk about that a bunch during the course.

We'll also talk about that sort of in, in what's next. Running some models locally. I'm really excited about like models getting good enough at medium size that you can run them locally, like on device or on robots or on local laptops.

And then real-time video. Like I really think real-time video is gonna hit the same inflection point that voice did by the end of the year or early next year. So we've got a couple of, uh, folks who are really focused on real-time video participating in the course and doing some credits and some office hours.

Swyx4:19

What do you think-- Sorry, w- w- there's, there's a lot there. Just double-clicking in reverse order, like what is real-time video really mean? Are, are you talking to like an avatar?

Real-time Video4:19

Kwindla Hultman Kramer4:30

Yeah. Yeah, yeah. It's like, just like you can now have a voice-based conversation with an LLM-

Swyx4:35

Like Tavis

Kwindla Hultman Kramer4:35

... you know, from, from folks like Tavis, you can have a video conversation with a transformer-based model or a collection of transformer-based models behind an API. And it's, you know, it changes the dynamic if you have that whole video channel.

Swyx4:52

Yeah. I, my, I, I tried this out and like it's... Who is this? I guess it's Charlie Witty. Okay. Yeah, yeah. I, I tried this out and I can't believe that there's a ton of traction behind this, 'cause it's kinda very robotic.

Kwindla Hultman Kramer5:06

We ser- we are the net- Daily is the network layer for Tavis and some others, and I can tell you there is a ton of traction behind it.

Swyx5:14

Yeah. Um, my sense is like it's, it's, uh, for coaching and...

What are the verticals that kinda use this stuff?

Kwindla Hultman Kramer5:24

Yeah. There's a bunch of coaching and kind of enterprise teaching and learning education stuff in-

Swyx5:29

Yeah

Kwindla Hultman Kramer5:29

... inside big companies.

Swyx5:30

Yeah.

Kwindla Hultman Kramer5:31

I also think we are starting to see the beginnings of some, some social stuff. I mean, a little bit to everybody's surprise, the voice AI monetization pull actually came from enterprise and B2B use cases. So a little bit to everybody's surprise, the inflection point with voice AI in terms of monetizable use cases came on the business side.

Uh, it's like telephony and customer support and, uh, you know, a bunch of like vertical SaaS with voice on top. I think real-time video may actually hit on the consumer side first. When it, when it's right, when it's on the right side of the uncanny valley- It's really, really compelling, and I think we're just starting to see some of that.

Swyx6:14

Yeah. Um, I'm starting to see, like, I think there's another platform, not Tavis, but, um, someone else doing more consumery video conversations. Are, are we talking about AI companions or is there some other product consumer that is interesting?

Kwindla Hultman Kramer6:27

Well, I think we're gonna have friends that are video in all our group chats.

Swyx6:31

I see.

Kwindla Hultman Kramer6:31

And the, the next TikTok is gonna be not just hyper-personalized recorded content, but hyper-personalized interactive content.

Swyx6:40

Yeah. Yeah, yeah. And, and, you know, for what it's worth, this, this is more, much more constrained domain for video than generalized video generation, which is a whole different thing. Like, here you're, maybe you're taking an avatar, maybe you're taking a clone.

You have, like, some kind of bone structure thing, and then you're animating that with, um, some expressions and, and the mouth movements, and that's about it.

Kwindla Hultman Kramer6:59

Yeah. There's a whole bunch of different approaches. You can go, like, single image to generate an avatar. You can take source video and you can kind of fine-tune a model. You can do rigged characters and have an LLM drive the rigged characters.

So people are experimenting in sort of every way with video. It feels early days, but-

Swyx7:15

Yeah, yeah

Kwindla Hultman Kramer7:16

... it's happening fast.

Swyx7:17

It's happening fast. Okay, so then recursing up, uh, there's just sections about-- You mentioned some open models that are just, you know, voice stuff. What are the main things and trends you're eyeing here? There's, there's sort of the end-to-end speech models like the KyuTime Mushis.

Open Models7:17

Swyx7:37

There are, like, the couple guys doing-- Never, never mind. I, I, I don't, I don't really know what the Parakeet and Dia type models are optimized for except for, for like, I guess podcast generation.

Kwindla Hultman Kramer7:50

Well, there's kind of the opposite ends of the spectrum there. Like, Dia is this amazing project by, like, one and a half Korean students. And it took, you know, it took at least my Twitter feed by storm like two weeks ago, and it's all the way on the, like, dynamic end of the spectrum.

Like, it's totally unhinged. And then Parakeet is NVIDIA's new speech model, speech transcription model that's number one on the leaderboards, and it's just, like, very enterprise tuned, like really, really rock solid, reliable at fairly small number of weights, like 600 million parameters or something.

So you've got, like, people building on the, "I need to use this in enterprise production" side of things, and you've got people building open source models that are like, "Let's just see how, let's just see how it goes.

Let's see what we can do."

Swyx8:33

Yeah, sorry. I-

Kwindla Hultman Kramer8:33

And they're both really valuable

Swyx8:34

... I made the link because I think the Dia guys said that they borrowed from Soundstorm and Parakeet in some ways. I, I don't know. I, I don't even know if there's code published, but they just took the papers and implemented them, which is the hard part.

Kwindla Hultman Kramer8:49

Yeah. NVIDIA does a great job publishing papers. It's super impressive if you follow NVIDIA research. They have a lot of GPUs, those researchers. They write good papers.

Swyx8:58

Definitely do. Um, yeah. What, what else is, um, interesting among all this? Like, is this, is this sort of a Pipecat centric? Are we gonna spend a lot of time on, I guess, this like guardrail scripting? I think, I think really, like, building smart voice is the-- A really quick demo is super easy, but then adding capabilities to it, making it reliable is really hard.

Kwindla Hultman Kramer9:20

That's a lot of the genesis of the course is like how do you get from a demo to something that you could actually use in production?

Swyx9:26

Mm-hmm.

Kwindla Hultman Kramer9:26

And there'll be a lot of Pipecat just 'cause I'm doing a lot of the work to kind of glue the course together, but it's not Pipecat centric or Pipecat only. There are a bunch of folks who are participating in the course who have their own stuff going on, like Freddy from Hugging Face is gonna talk about Fast RTC and hang out in the Discord and kind of be around.

I really love that work. Uh, Sean from OpenAI is gonna talk about their WebRTC and 1-800-CHATGPT layer, and so is Dominic, gonna talk about the OpenAI Agents framework, which they did a really elegant like voice layer on top of that original Agents framework in the most recent release.

Swyx10:00

Mm-hmm.

Kwindla Hultman Kramer10:01

We have like VAPI, which is a, a kind of higher up the stack, like everything together platform that I recommend often to a lot of people who start with Pipecat but kind of want more batteries included. There's a brand new competitor VAPI, I think it's fair to say, called LayerCode, uh, that has $1,000 of credits for students in the course that I got an early demo of.

I think you know those guys, and we were like, "Oh, well you should, you should hang out in the course too. You've got good ideas."

Swyx10:27

Do I know those guys? Have we-

Kwindla Hultman Kramer10:29

They told me that you know them.

Swyx10:30

Uh, I, I may have. I, I know a lot of people, but

Kwindla Hultman Kramer10:32

Yeah, I know.

Swyx10:33

Maybe, maybe I met them individually without knowing the name of the-

Kwindla Hultman Kramer10:36

LayerCode is very new, so actually that, that may be the case.

Swyx10:39

Yeah.

Kwindla Hultman Kramer10:39

It's like Damian Tanner.

Swyx10:41

Yeah, okay. I do know him. Yeah. Uh, interesting, 'cause when Damian was at Deepgram... Was it, was it Deepgram? Well, he, he went over to LayerCode and I was like, "Oh, this like seems like an established company." He did not say that or there was no indication that this is a new company.

Kwindla Hultman Kramer10:57

I think it's still in private beta. I think it's that new, which is fun.

Swyx11:00

Okay.

Kwindla Hultman Kramer11:00

There is so much stuff launching in voice and I, I've really tried to pull in like lots of people that I know because, like, you can't cover everything yourself. I mean, you do an incredible job covering like everything across all of AI, but we, we can't all be you.

Swyx11:15

Well, there, there should not be. You know, I think the world needs specialists, and then we need to be able to lean on their expertise and, and then, you know, build products. Like I, I-- There, there should not be that many of me who waste their time being journalists and, and wide on everything and shallow on, on a lot of them.

Voice & LLMs11:15

Swyx11:29

So, okay. I'll, I'll just dive in, double-click a little bit into how voice currently interacts with LLMs because you mentioned this thing about how the Agents framework from OpenAI adds a voice layer, and I think that's something that Pipecat does as well.

Everyone wants to stream. Everyone wants to, uh, do semantic detection. It's really hard, and I don't know if we have the right interfaces to do this without adding a ton of latency that makes it very unnatural.

Kwindla Hultman Kramer12:03

We basically spent all of 2024 trying to get to the 80/20 point on, you know, the half a dozen hard problems here that you were outlining. It's like how do you do the low latency networking of the audio to somewhere in the cloud where you can do the orchestration?

How do you do turn detection so you know when the user's done talking and the LLM should start talking? How do you do context management that's kind of optimized for the real-time multi-turn- You know, your LLMs are stateless, right?

So you have to maintain state, and that is a whole interesting rabbit hole to go down. How do you make it so that you can swap out different parts of the system because models are evolving so fast in, in every category, the models are leapfrogging each other, so you don't wanna only be tied to a single model?

And then how do you do things like function calling, which were never designed to even work in a multi-turn context, much less happen asynchronously or happen at low latency or, you know, whatever, how- whatever hole you have, you have a round peg or vice versa with things like function calling and structured data output and all that stuff.

So I think a lot of us who were working on voice last year were like, "Well, let's just get it so it works pretty well. And once we have all the pieces working pretty well, there's gonna be something really interesting here."

And it felt like that sort of happened at the end of last year, and now there are these kind of 2025 problems, which is how do you write really great evals? How do we give really great feedback to the people at the labs who are training these models, who, you know, a lot of what we're doing, it's become clear is pretty out of distribution.

The multi-turn, multimodal native audio stuff is all like it's frontier of the frontier models. And how do you like level up from the 80/20 point for things like turn detection? Like we have a-- we'll, we'll talk a lot in the class, I hope, about this open source, open data, open training code, uh, native audio turn detection model that a lot of folks, including me and the Pipecat community, have been working on.

So there's like the 2024 problems, and now there's a new set of higher, you know, in the Maslow's hierarchy, 2025 problems.

Swyx13:55

Yeah. Yeah, I, I remember last year talk, um, asking you publicly about the semantic turn detection. Now, like assuming that is solved, which is, you know, it's not, not really, but it's, it's getting better. Do people have... Do, does this Pipecat or any other framework support thinking while the other person's speaking?

Kwindla Hultman Kramer14:12

So there's a question about what the model can do, and there's a question about what you can put together from the orchestration pieces. And none of the, none of the SOTA models can kind of think while also taking input.

The, that Kyutai Mochi model you mentioned, which was my, my, my favorite academic paper last year, is a step towards a truly bidirectional streaming in both directions, thinking all the time LLM. That work has not been kind of scaled up to, to, to, you know, large LLM size.

So what we try to do in the Pipecat world is we do a bunch of things in parallel in separate Python tasks and have abstractions for kind of putting those things together or having those things interact in useful ways.

And you can do that with like state machine abstractions, you can do it with like multi-agent abstractions, all of which I think are just kind of abstractions on top of the idea that you often actually want multiple LLMs talking to each other.

One of my favorite Pipecat demos is this, uh, two connections to Gemini multimodal live-

Swyx15:14

Yeah, talking to each other

Kwindla Hultman Kramer15:14

... playing a word guessing game with you.

Swyx15:16

Yeah.

Kwindla Hultman Kramer15:16

One, one LLM is the judge and the moderator of the game, the other LLM is the player you're cooperating with. Increasingly, people are building stuff like that, and increasingly you want to have great abstractions in your agent framework to do that kind of thing.

Swyx15:30

Yeah. Amazing. And then, uh, one- always want to have one thing about telephony and other integrations. What are the trends in people accessing these things? So you have got the, you got the 1-800 ChatGPT, you've got, um, the little microphone in every, next to every text box.

Telephony15:30

Swyx15:47

What else are people doing? What's hard?

Kwindla Hultman Kramer15:48

I mean, 99% of the monetizable voice AI use cases today are telephony.

Swyx15:53

So like Twilio?

Kwindla Hultman Kramer15:54

Like Twilio. A lot more people are using Pipecat with Twilio than with anything else-

Swyx15:58

What's, what's the number two with Twilio?

Kwindla Hultman Kramer15:59

... whatsoever to see. I'm sorry. Say that again.

Swyx16:00

What's, what's number two after Twilio? I'm just-

Kwindla Hultman Kramer16:03

After Twilio, it gets more regional.

Swyx16:05

Okay.

Kwindla Hultman Kramer16:05

So like Plivo in India, where Twilio doesn't have great coverage, like can't, for regulatory reasons, can't do phone numbers and stuff. So there's a bunch of places all over the world where there are regional winners.

Swyx16:15

Yeah.

Kwindla Hultman Kramer16:15

Twilio has great infrastructure. It's hard not to, you know, buy with Twilio if you, if, if you're in a place Twilio serves.

Swyx16:21

Right. So like the web voice API connections, like the retails of the world, they're much smaller than compared to the Twilio ones, where you just have the phone line.

Kwindla Hultman Kramer16:33

Yeah. And all of the, you know, the, the VAPI Retail, Bland, Sierra only, like lo- I, I sort of mixed a bunch of categories there, but, but most of those, even if you've tried them from their website, most of their revenue, almost all of their revenue comes from inbound or outbound telephony-

Swyx16:52

Yeah. Yeah. Amazing

Kwindla Hultman Kramer16:52

... today.

Swyx16:53

Yeah, yeah. Um, I'm always curious, like just because, um, I would like to see more integration, more voice everywhere, you know, like I, I, I'm a supporter of, of like the B and like Limitless and like sort of just ambient listening devices.

I briefly looked into trying to have my, make my own home speaker thing that just listens at all times and, you know, takes notes. It's, it's surprisingly hard. Like even if you take the hackable like Raspberry Pi based, um, what is it called?

Hardware17:06

Swyx17:20

Home Assistant thing, it's very elementary. Like you can bas- you can just about tell it to turn on your lights.

Kwindla Hultman Kramer17:27

Yeah. It feels like, uh, compiling a Linux kernel in like 1999. It's like that, that's the level where we are for the home automation hardware do it yourself stuff. But I do think hardware is super interesting and hardware with voice is super interesting.

I, I wanna do-- I, we actually haven't scheduled it yet, but you know, you and I should like recruit a bunch of people to do a hardware session for the class.

Swyx17:49

I mean, I, I can get the, the B folks anytime. Um-

Kwindla Hultman Kramer17:53

I've been playing with little robots. For the people who are watching on YouTube, this is my voice controlled, latest voice controlled robot with a camera on it.

Swyx18:01

Okay. What does it do?

Kwindla Hultman Kramer18:03

It drives around and it recognizes you, and if you've had a conversation with it before, it remembers what you talked to it about.

Swyx18:09

No way. Okay, you gotta do a little demo of that. That's awesome.

Kwindla Hultman Kramer18:13

It, it started out as a hackathon project at one of the, one of the Antarctic hackathons.

Swyx18:16

Didn't you get the, the dogs too? You got the dog?

Kwindla Hultman Kramer18:20

Yeah, I got the dog, but the dog, we need to get Sean, uh, Sean Dubois' help on the dog because that's a sort of ESP32 and I just haven't managed to like compile a WebRTC stack for it. But Sean has one.

The robot I just held up is a Raspberry Pi, so you know, it's, it's easy mode. It runs Python.

Swyx18:36

Yeah, yeah. Excellent. Excellent. Um, cool. Um, that's a, you know, that's, that's a brief survey of the course. There's, there's a lot of other stuff. Anything else that you think we should highlight but haven't?

Kwindla Hultman Kramer18:47

I just wanna highlight how much of a community effort we want it to be. So like we want people to hang out in the Discord, trade ideas like the-- It's a course, but like in my mind, the course is just a Trojan horse for, for a community, like a festival.

Community18:47

Swyx19:01

It is. It is. And, uh, you know, you'll be at Google IO and the AI Engineer Summit at World's Fair. Got that right?

Kwindla Hultman Kramer19:07

Yeah.

Swyx19:08

And, uh, doing, doing a lot more voice, uh, vo-

Kwindla Hultman Kramer19:11

You know, we should actually credit AI Engineer World's Fair as the other thread that led to this. Like I used AI Engineer World's Fair as a forcing function for-

Swyx19:17

Oh, yeah

Kwindla Hultman Kramer19:17

... writing this like 16,000-word guide on voice AI.

Swyx19:19

I have the book too. It's upstairs, but yeah.

Kwindla Hultman Kramer19:21

And the, like the feedback on that was like, "Let's do reading groups about this. Let's talk about this in the Discord." And I was like, "That's interesting. Let's, let's do that." And kind of that pinged around and-

Swyx19:31

Mm-hmm

Kwindla Hultman Kramer19:31

... became this course.

Swyx19:32

Yeah. Excellent. What is the Discord? Is it, is it the Pipecat Discord?

Kwindla Hultman Kramer19:36

I think we'll just do it in the Pipecat Discord 'cause like that way I don't have to set up another Discord, but-

Swyx19:40

Yeah

Kwindla Hultman Kramer19:41

... we'll figure it out in the next 24 hours.

Swyx19:42

Too many Discords. Too many Discords. All right. Excellent. Um, well that's it. Uh, you know, everyone go sign up for the course and I'll put the link in the show notes. Any other, uh, any last calls to action or words of wisdom?

Kwindla Hultman Kramer19:53

Ping me if you have an idea for a session you wanna run.

Wrap-up19:53

Swyx19:56

Wow. Okay. You're still taking sessions. You already have, you already have too many, man.

Kwindla Hultman Kramer19:59

We have 28, so yeah, we do have too many but, you know.

Swyx20:02

What are you going for here?

Um, no, good, good, good. Uh, I mean, you know, I think it's a special time. I think voice is definitely one of these. The timing is right to work on these things because voice is here but it's not evenly distributed and I think more people can, can, can enhance their existing AI applications with voice just 'cause it actually makes it more, much more magical than-

Kwindla Hultman Kramer20:23

And it's the future

Swyx20:24

... just typing it in the chat window.

Kwindla Hultman Kramer20:25

I mean, uh, U- UX is gonna be, you know, 50%, 60%, 75% voice in the future. I 100% believe that and I did not believe that, you know, two years ago. Uh, but the trend line is just really, I think-

Swyx20:36

Yeah

Kwindla Hultman Kramer20:36

... clear.

Swyx20:37

I mean like this is our first year of having a full voice track, you know, and, uh, and if I were to subdivide it, I don't know how I would subdivide it, but clearly, you know, this is growing.

Excellent. All right. Thanks, Kwindla.

Kwindla Hultman Kramer20:49

Thank you. Bye-bye.