# [AIEWF Preview] Gemini in 2025 and Realtime Voice AI

Latent Space · 2025-06-02

<https://addtry.com/4f22fd3b-0e5f-4b86-a99b-e12712fb1f33>

At Google I/O 2025, Latent Space host Swyx and TWiML's Sam Charrington speak with Logan Kilpatrick, Shrestha Basu Mallick, and Kwindla Hultman Kramer about Gemini's reasoning upgrades, real-time voice AI, and the infrastructure challenges of building production voice agents. Kilpatrick announces thinking budgets and thought summaries for 2.5 Pro, now live in Flash and coming to Pro in June, while Mallick highlights native audio output with multilingual code-switching and proactive voice activity detection. Kramer discusses the need for sub-500ms latency and WebRTC packet routing, and the team debates single-model vs. componentized architectures, with Mallick noting that eventual convergence will bring capabilities like diffusion-based generative UI into the main Gemini model. The episode also covers implicit context caching, URL context for research agents, and the Live API's session length and function calling improvements.

## Questions this episode answers

### How might Google's diffusion language model enable generative UI?

Logan Kilpatrick shares that the Gemini Diffusion model could enable generative UI experiences where websites are built on-the-fly based on user interactions, generating code in response to clicks. This is possible because diffusion models generate tokens faster than autoregressive models, potentially unlocking real-time UI generation not feasible with slower models. However, the model still needs hardening before broader release.

[6:02](https://addtry.com/4f22fd3b-0e5f-4b86-a99b-e12712fb1f33?t=362000)

### How does Google's Gemini Live API handle real-time voice AI challenges like latency and long sessions?

Shrestha Basu Mallick explains the Gemini Live API's native audio-to-audio architecture targets 500–700 ms latency. It initially allowed only 15–20 minutes of audio and 5 minutes of video, but now offers session length controls, tool chaining, and system instruction updates. A 'proactive audio' feature ignores irrelevant sounds, and it unofficially recognizes different speakers, addressing key real-time voice challenges.

[7:37](https://addtry.com/4f22fd3b-0e5f-4b86-a99b-e12712fb1f33?t=457000)

### How can I control the thinking process in Gemini 2.5 Pro?

Logan Kilpatrick says Gemini 2.5 Pro will soon receive a thinking budget feature, likely in early June, letting developers set a token budget for reasoning or disable thinking entirely for a non-reasoning mode. Thought summaries are already live, providing condensed reasoning output instead of full thought traces. This gives developers more control over the model's reasoning behavior and output.

[1:31](https://addtry.com/4f22fd3b-0e5f-4b86-a99b-e12712fb1f33?t=91000)

## Key moments

- **[0:00] Intro**
- **[1:14] I/O Highlights**
  - [1:21] Gemini 2.5 Pro will get thinking budgets in early June, and thought summaries are live now, says Logan Kilpatrick.
  - [2:39] Native audio output enables seamless language switching, including Klingon, says Shrestha Basu Mallick.
  - [3:52] Implicit context caching now available, saving developers money without any action, says Logan Kilpatrick.
- **[6:01] Gemini Diffusion**
  - [6:09] Gemini Diffusion will enable on-the-fly generative UI that builds from user actions, predicts Logan Kilpatrick.
- **[7:37] Live API Hurdles**
  - [7:41] Shrestha Basu Mallick lists awareness, session length, and tool calls as key challenges for Live API adoption.
  - [10:16] Live API enables gaming agents, multi-hour customer support, and screen-assisted workflows like Shopify's DNS setup.
- **[11:04] Voice Architectures**
  - [11:04] Q: Is speech-to-text plus LLM a precursor to native audio-to-audio, or a distinct path?
  - [11:24] Most voice AI use cases will eventually transition to native audio-to-audio, predicts Shrestha Basu Mallick.
  - [12:04] Logan Kilpatrick explains Google's one-model strategy: merge separate capabilities like reasoning back into Gemini.
- **[14:35] Kwind Joins**
  - [16:05] Live API originally used separate TTS output from NotebookLM to hit latency and quality bars, says Shrestha.
- **[17:03] Voice Infra**
  - [17:26] Live API offers tunable voice activity detection and bring-your-own-VAD mode, says Shrestha Basu Mallick.
- **[18:49] Dev Framework**
- **[20:19] Proactive Audio**
  - [20:19] Proactive audio feature in native audio-to-audio enables Gemini to ignore irrelevant background speech.
  - [21:05] Native audio dialogue model can distinguish two speakers by voice, though not officially supported, says Shrestha.
- **[21:56] Async Calling**
  - [22:07] Asynchronous function calling now available in the Live API's cascaded architecture, with plans for native audio.
- **[22:45] Wrap**
  - [22:48] Kwindla Hultman Kramer predicts Gemini 5.0 will be announced at the next Google I/O.
  - [23:14] Gemini officially supports 24 languages and can already speak Klingon, hinting at fast expansion, says Shrestha.

## Speakers

- **Swyx** (host)
- **Kwindla Hultman Kramer** (guest)
- **Logan Kilpatrick** (guest)
- **Sam Charrington** (guest)
- **Shrestha Basu Mallick** (guest)

## Topics

Reasoning, Agent Infrastructure

## Mentioned

Daily (company), Google (company), LiveKit (company), AI Studio (product), Gemini (product), Imagen (product), Latent Space (product), Live API (product), NotebookLM (product), Pipecat (product), TWiML AI Podcast (product)

## Transcript

### Intro

**Sam Charrington** [0:05]
Hey, I'm Sam Charrington, welcome to another episode of the TWiML AI Podcast.

**Swyx** [0:09]
And I'm Swyx. Uh, this is a special episode of the Latency Space Pod with, uh, TWiML at Google I/O. Welcome .

**Sam Charrington** [0:15]
TW-

**Logan Kilpatrick** [0:15]
Thanks for, thanks for being here, and thanks for hanging out with us. I'm excited.

**Swyx** [0:20]
Logan, you are, you were our first guest. You came back remotely a few months ago, and, uh, now you're back. You're sort of the... A lot of the face of the AI studio, basically. Like, that, that a lot of people are using.

I'm using it, and, uh, you know, I think it's a really welcome change for people, like, being more accessible with, like, the rest of the Google pr- suite. And Shrestha, you've been... I actually don't super know your role.

Like, I, I just generally have you pegged as, like, PM of the API team with a s- particular focus on live.

**Logan Kilpatrick** [0:47]
Shrestha runs the show behind the scenes.

**Sam Charrington** [0:49]
Yeah.

**Logan Kilpatrick** [0:49]
Behind the sc- Is the least public face running the show behind the scenes.

**Sam Charrington** [0:53]
Yeah.

**Logan Kilpatrick** [0:53]
Model launches, the live API. Generally, all the stuff that's happening in the API is, uh, Shrestha's hard work, so.

**Shrestha Basu Mallick** [1:00]
Thank you for that, Logan, but I think everyone knows who really la- runs the show. There's public evidence there. But yeah, I, uh, I work with Logan and a few other excellent PMs, but I lead, uh, the API side of the house.

### I/O Highlights

**Swyx** [1:14]
There's a lot of announcements, and I think a lot of people have sort of done their recaps. What are you guys' personal highlights over I/O?

**Logan Kilpatrick** [1:21]
I'll break the rule, and I'll give two that are not the sort of big, big flashy ones. I think the two that I think developers are gonna be super excited about, one, thinking budgets coming to 2.5 Pro. So, and you can also...

You'll be able to disable thinking as well, so if you just want 2.5 Pro as, like, a raw, non-reasoning model, we'll have that, uh, hopefully in early June. And then thought summaries. So I'm, we, we've had this debate internally about, like, do we need to show full thoughts?

Do developers want full thoughts? I think s- developers say they want full thoughts. We have thought summaries right now as a sort of step in that direction. It'll be really interesting to find out and get the feedback around, like, what are the things that work, that work with thought summaries?

What are the things that don't work with thought summaries? I was reading some threads last night about, like, thought summaries are now live in Cursor as well, and people were sort of reacting to, you know, having summaries versus not full thoughts.

So it'll be interesting to see, but I'm, I'm excited for both of those things. Thought summaries are live now. Thinking budget for 2.5 Pro will land with the GA model in a couple of weeks.

**Shrestha Basu Mallick** [2:19]
Yeah, and I should say we already do have thinking budgets in 2.5 Flash.

**Logan Kilpatrick** [2:23]
In Flash. Yeah, yeah.

**Shrestha Basu Mallick** [2:24]
I, I do think, you know, with all of the, our features that we're releasing on top of our thinking models, summaries, budgets, I think this is our way of, you know, you have the models, but then we want to give developers as much control as they can on top of models.

But coming back to your question about my favorite feature, it, it's really hard to pick 'cause, like, all of these features we've been trying to push out for weeks. Uh, but I think, uh, native audio output is a personal-

**Swyx** [2:50]
I was just saying that with Quinn. Yeah .

**Shrestha Basu Mallick** [2:53]
Yeah, yeah, is a personal highlight. I actually, uh, Quinn and I have been playing with it together for a bit as well. I think, uh, especially with all the, obviously, the voices sound great. Um, the fact that it can switch in and out of languages, so Matt Veloso, our boss, actually has a, has a demo on Twitter where it actually speaks Klingon, even though that's not an officially supported language.

Uh, but, you know, for, I speak Bengali. Just being able to s- for it to switch into and out of Bengali and English, that's been special. And then if I get to pick another one, it's, uh, I'd say we released a new tool called URL context, and the idea is that you can use it by yourself or pair it with search to retrieve more in-depth information from th- uh, web pages in a way that's respectful of our publisher ecosystem, of course.

And I think that'll, this'll unlock new use cases, like if people want to build their own version of a research agent, which is something developers ask us for a lot.

**Swyx** [3:52]
Yeah.

**Shrestha Basu Mallick** [3:52]
Yeah.

**Sam Charrington** [3:52]
It's worth mentioning that just prior to I/O, there was a ton of new, interesting new capability-

**Shrestha Basu Mallick** [3:57]
Yeah

**Sam Charrington** [3:57]
... including the update to Gemini 2.5 Pro.

**Logan Kilpatrick** [4:00]
Yeah.

**Sam Charrington** [4:00]
As well as the implicit context caching, which I know a lot of folks were waiting for.

**Shrestha Basu Mallick** [4:04]
That's right.

**Logan Kilpatrick** [4:04]
We made implicit caching happen. I think there was lots of feedback that people are like, "Explicit caching is nice." Like, there's definitely use cases where it makes sense, but people want implicit caching, so I'm, I'm happy, um, passing the cost saving on to developers.

You don't have to do anything. It just works right now, and you're saving money. It's a, it's a great outcome.

**Swyx** [4:21]
Yeah. I don't want to manage that myself.

**Logan Kilpatrick** [4:23]
Yeah.

**Swyx** [4:23]
It's-

**Logan Kilpatrick** [4:23]
And there, there are, there are people, like, I think if you... There's so many use cases where, like, you're just doing chat on the same stuff over and over again. Um, and for those use cases, you know, you want to be able to explicitly cache the thing-

**Swyx** [4:35]
It has its place

**Logan Kilpatrick** [4:35]
... and make sure and guarantee you're cached so that you save money. So I'm, I'm happy we have that.

**Swyx** [4:39]
Is there any behind the scenes of, like, what makes caching hard, or anything that people don't appreciate about caching as a general concept?

**Logan Kilpatrick** [4:47]
That's a-

**Swyx** [4:47]
'Cause I think this is a very important pricing paradigm that people-

**Shrestha Basu Mallick** [4:49]
Yeah

**Swyx** [4:49]
... need to really get behind.

**Logan Kilpatrick** [4:51]
Yeah, that's a good question. I think there's a trade-off between, like, all of the dimensions of caching, which is around, like, the sort of latency, 'cause in some cases you're getting latency gains. In other cases, it's like, how much, you know, what's the cost for Google?

How much stuff do you wanna cache altogether? So we could have an entire episode and get a bunch of the caching people.

**Shrestha Basu Mallick** [5:08]
Yeah.

**Logan Kilpatrick** [5:09]
Um, it's like a good example of, like, an infrastructure problem to be solved, and a bunch of-

**Swyx** [5:13]
Yeah

**Logan Kilpatrick** [5:13]
... the folks who we work with love working on this problem, so we should do a deep dive episode.

**Swyx** [5:18]
Yeah, I wanna shout out that you've been doing more, uh, video stuff. You have your own podcast as part of your Gemini work. You've also been doing, uh, uh, video with, like, people on the team who have, like, done the work.

**Logan Kilpatrick** [5:28]
Yeah, yeah. It's been fun. We had-

**Swyx** [5:29]
As the long context video

**Logan Kilpatrick** [5:30]
... so we should do, like, a caching episode. That's a-

**Swyx** [5:30]
Exactly. Like, you did the long context one. People loved it.

**Logan Kilpatrick** [5:32]
Your reception was very, was very positive about the long context one, so thank you. That was the first time-

**Swyx** [5:36]
Yeah

**Logan Kilpatrick** [5:36]
... that we did, like, a more deep technical discussion with folks on the team, and, and Nikolay's awesome, and we actually just did one with, we did one with Shrestha about the live API, which I'm excited about. We did one with folks on the team about the multimodal capabilities in Gemini.

We're gonna do a pre-training one hopefully-

**Swyx** [5:53]
Ooh

**Logan Kilpatrick** [5:53]
... uh, which will be really cool. We've got a bunch of people who are excited to talk about that, so there's a bunch of them in the works, and it's, um, it's fun to make them happen and, and have those conversations.

### Gemini Diffusion

**Swyx** [6:01]
Yeah, and my underrated pick is, um, Gemini Diffusion

**Shrestha Basu Mallick** [6:04]
Yes

**Swyx** [6:04]
Yeah, yeah, yeah.

**Shrestha Basu Mallick** [6:05]
Uh-

**Logan Kilpatrick** [6:05]
It's not underrated

**Shrestha Basu Mallick** [6:06]
... oddly, it is underrated- ... for all the love it's getting

**Logan Kilpatrick** [6:09]
Which is the coolest thing ever

**Shrestha Basu Mallick** [6:09]
Yeah

**Swyx** [6:09]
So like apart from speed, I, I wonder like what the potential results of a diffusion language model could be-

**Logan Kilpatrick** [6:16]
Generative UI. Generative UI. This is the way the generative UIs happen is through, through this-

**Swyx** [6:20]
Language?

**Logan Kilpatrick** [6:21]
... experience. The UI bit, just like being able to like say, I want... You know, build the UI on the fly using code based on what a user does. So like you have no pre-compiled notion of what your website is, and as a user goes through, as they click buttons, thousand tokens generate, and it just like makes that UI for you.

**Swyx** [6:39]
Interesting.

**Logan Kilpatrick** [6:40]
I think that's gonna be possible. I mean, I think there's a lot of work to-

**Shrestha Basu Mallick** [6:43]
Yeah

**Logan Kilpatrick** [6:44]
... productionize, make Gemini diffusion like actually a high-quality model that meets the bar for us to bring to the world more generally, but-

**Swyx** [6:51]
Yeah

**Logan Kilpatrick** [6:51]
... I do think that's gonna be the... The killer use case will be like this generative UI experience that doesn't exist today because the models just take too long to generate tokens.

**Swyx** [7:00]
Yeah.

**Logan Kilpatrick** [7:01]
For me, it's really the role that audio and video are taking throughout a bunch of independent product releases, from the generative models to the live API to the on-the-fly, uh, transcription and, um-

**Shrestha Basu Mallick** [7:14]
Absolutely

**Logan Kilpatrick** [7:14]
... translation.

**Shrestha Basu Mallick** [7:15]
Yeah

**Logan Kilpatrick** [7:16]
Um, it's, uh, I think kind of sha- foreshadowing the role that that's gonna play in a lot of developer applications.

**Shrestha Basu Mallick** [7:25]
Yeah, transcription actually even before we released native audio, now of course you get, uh, text and audio interleaved in the output, but transcription used to be one of the biggest use cases we had on the live API.

### Live API Hurdles

**Swyx** [7:37]
Yeah

**Logan Kilpatrick** [7:37]
What are you seeing as the challenges for folks getting started with live?

**Shrestha Basu Mallick** [7:41]
Yeah, that's a great question. I think firstly, uh, awareness, right? Like people knowing that we have a live API-

**Logan Kilpatrick** [7:49]
You can do this.

**Shrestha Basu Mallick** [7:50]
That's why we're doing this, talking to you folks. I think some of the areas where... So we were actually the first to market with also video input, but one of the areas where we've been getting a lot of feedback is in session length.

Anybody who's been trying to put this in production, like when we started, you could do like fifteen to twenty minutes of audio, I'm sorry, and about five minutes of video. And so we've been putting in a lot of knobs for developers, and we can talk about that more if you guys want, to...

for people to have a sliding window or decide what resolution they want to send video in, but, you know, to basically increase the session length. And then tool calls, that was another area where we used to get a lot of feedback.

Again, we were very proud because we introduced tool chaining first, so you could chain search and code execution, do all kinds of analysis, but then we've had to do a lot of work in improving function calling-

**Logan Kilpatrick** [8:45]
Yeah

**Shrestha Basu Mallick** [8:45]
... improving the performance of search, and, and we, we continue to push on that.

**Logan Kilpatrick** [8:49]
I've got a quick one on this too, which is, I think the level of commitment you need to make to the model provider in the world of the live API, like I, I do think for developers is a higher bar.

**Shrestha Basu Mallick** [8:58]
Yeah

**Logan Kilpatrick** [8:58]
If you look at like what is chat completions or like what does, for us, generate content provide from just like a, a text modality perspective, it's like it's a pretty lightweight thing. There's a lot of model providers that have that option.

Like I could switch-

**Shrestha Basu Mallick** [9:11]
Yeah

**Logan Kilpatrick** [9:11]
... to a different provider if I end up not liking some model provider, which I think is good for the ecosystem. I think if you look at a lot of the live API infrastructure right now, like you really do need to commit that you're like gonna...

You know, there's... It's not easily interoperable between different model providers. Like everyone's infrastructure is all bespoke and different, so like it is a, it's a different level of commitment that you need to have to like really bet your company or your business or your product on the live API, which I, I do think is a challenge for developers to sort of make that level-

**Shrestha Basu Mallick** [9:40]
Yeah

**Logan Kilpatrick** [9:40]
... of commitment in this like fast-moving AI world. But I think hopefully there'll be like some level of like similarity, and you'll get some model-agnostic infrastructure to help make that, you know, make developers feel a little bit, uh, a little bit easier about being able to move between models potentially.

**Shrestha Basu Mallick** [9:56]
I could go on and on, but if you have, say, more complex workflows, then one of the things is being able to change the system instructions at every step of your workflow. And so yeah, so onboarding some of the more complex use cases with the live API has been a work in progress as we've released like more features.

**Swyx** [10:13]
So what kind of complex workflows are we talking about?

**Shrestha Basu Mallick** [10:16]
You know, like, like we have people who are building, say, gaming agents, but like which have multi states, for example, in them. Uh, we have a lot... I mean, this was a famous demo at Next, but we have, uh, folks who want to, uh, you know, uh, customer support agents, of course.

You know, they can- the sessions can last for hours, right? Then there's a lot of use cases around people showing a certain screen-

**Logan Kilpatrick** [10:41]
This is the coolest use case, honestly

**Shrestha Basu Mallick** [10:41]
... and getting... Yeah. Yeah, a- and I was, I was referring to like the famous demo at Next where Shopify showed how to set up a DNS using Cloudflare, right? So in certain cases, especially the longer your workflow runs, like you might have to go from one state to another state, and you might want to change their side, or if you handle it fro- hand it from one agent to another agent, you might have to change the system instruction.

### Voice Architectures

**Logan Kilpatrick** [11:04]
When you're thinking about building voice-based applications, is speech-to-text and then processing with a standard LLM, would you say that's like a precursor to the live era, or are these two distinct paths that are still viable and that you still see being viable going forward?

**Shrestha Basu Mallick** [11:24]
That's a tough question, and I'm still- Uh, right now we have both out. I do think perhaps eventually, for most use cases, as these audio-to-audio architecture models get better, a lot of use cases will probably transition to that.

But you know, when we talk to our developers, they still very much like those componentized, um, componentized components. So that's why we also put out, uh, two new text-to-speech models at I/O. Not available through the live API yet, but really high-performing, controllable, promptable text-to-speech models.

**Logan Kilpatrick** [12:04]
I have a, an angle of an answer to this question, which is I, I talked to Koray this morning, who's, who's our boss's boss, the, the CTO at DeepMind, um, and Koray had a, had a really interesting take, which is just around like what makes- One of the main things that makes what we're doing at Google with Gemini different than what a lot of the other labs are doing is, like we're here to make one model, and like that model is Gemini.

And like I think, I think you, you do need to, to Shrestha's point, like to make the capabilities work in some cases, like you do need to have these forks that like go off and, and make that capability and harden it, and then find a way to bring it back into the mainline model.

But like we want to make one model, and it's the Gemini model, and like not have the sort of splintering of all these different capabilities. And we've done a good job of, I think, thinking the reasoning stuff was like the best example of this.

We had those, they were separate from the mainline Gemini models so that those teams, the research teams could go and hill climb and make progress, and not need to be constrained about like, "How do we do this without having there be collateral damage on other capabilities like multimodal or something like that?"

But the teams went and did that, and then they find a way to sort of bring the capabilities together, and oftentimes what you see is there's tension in bringing them together, but it's the really exciting thing is what happens when you bring the capabilities together.

And like 2.5 Pro with reasoning is a great example of this, where like multimodal with video understanding ended up like having this huge, like it's having this beautiful moment. The model is like SOTA out of the box- ...

because of all the reasoning capabilities that were baked in. It wasn't because they like did a bunch of stuff to make video understanding really good, it was just like an artifact of bringing and merging those capabilities together. So I think that as like a north star for Gemini models makes, makes a ton of sense.

**Shrestha Basu Mallick** [13:45]
I agree with you, and, and that's what I said, right? Like I think eventually a lot of use cases will end up on Gemini, will end up on Natural Voice.

**Logan Kilpatrick** [13:51]
Yeah.

**Shrestha Basu Mallick** [13:52]
But I think in order to foster development, like we have these offshoots for d-

**Logan Kilpatrick** [13:56]
Yeah

**Shrestha Basu Mallick** [13:56]
... from time to time. We have our Imagine models for image generation, even though now another I/O, well, slightly pre-I/O announcement, you can do interleave text and image within Gemini also, right? And it un- unlocks different-

**Logan Kilpatrick** [14:10]
Yeah, but those are different models, right? The-

**Shrestha Basu Mallick** [14:11]
Those are different models

**Logan Kilpatrick** [14:12]
... the one is autoregressive, the other is diffusion.

**Shrestha Basu Mallick** [14:14]
Uh, the other is, that's what I'm saying, right? But for a lot of image generation, image editing, high-quality photorealistic use cases, developers are still using Imagine. Um-

**Logan Kilpatrick** [14:24]
Yeah

**Shrestha Basu Mallick** [14:24]
... but then, you know, slowly but surely, we're bringing those capabilities into Gemini as well.

**Logan Kilpatrick** [14:29]
Whoever's, whoever's watching this, we had a, like a mid- ... uh, I/O switch because obviously, uh, you know, there's a lot going on here. Um-

### Kwind Joins

**Shrestha Basu Mallick** [14:35]
This is not AI shapeshifting.

**Logan Kilpatrick** [14:36]
I know, I know. Uh, but we also have Kwind, actually, who, who made this podcast happen. Uh, but, uh, you're a founder and CEO of Daily. Welcome.

**Kwindla Hultman Kramer** [14:44]
I'm a big fan of all things voice and audio.

**Logan Kilpatrick** [14:47]
All things voice and audio.

**Kwindla Hultman Kramer** [14:47]
So it's good to be here with you and with Shrestha.

**Logan Kilpatrick** [14:50]
Kwind actually runs the Voice AI meetup in San Francisco. Like you are basically consistently the leading sort of community builder, uh, and you're very generous of your time and knowledge, so I really appreciate that. And obviously, also recently you started Pipecat, which is this open source framework for voice orchestration.

Um-

**Kwindla Hultman Kramer** [15:06]
Which has really great-

**Logan Kilpatrick** [15:07]
Yeah

**Kwindla Hultman Kramer** [15:08]
... uh, support for all the Gemini models.

**Logan Kilpatrick** [15:10]
Uh, you wanted to say something about the relationship with Gemini and Daily?

**Shrestha Basu Mallick** [15:13]
I just wanted to say that it's been a very, very, uh, fruitful partnership with Daily. They've been our partners since the launch of the Live API, and, uh, you know, a lot of their feedback that they continuously give has been, uh, uh, you know, instrumental to the success of the Live API.

So both Daily and LiveKit are, uh, we're partnered with them.

**Logan Kilpatrick** [15:35]
Kwind, I think you, I think, you know, we had, we had a little bit of a prep for this. You also wanted to dive into a little bit on like the cascade of, uh, models in Gemini Live.

**Kwindla Hultman Kramer** [15:43]
I mean, I think Shrestha's taken a really interesting approach designing these APIs. So you talked about components a little bit, you talked about how you want to be able to do things both in the Live API and in the more sort of traditional chat API.

And you've got, uh, originally you designed the Live API to have audio in, but then it's a separate text model, the NotebookLM model's audio out. What was the sort of driver for that originally?

**Shrestha Basu Mallick** [16:05]
I mean, that at the time was, um, we wanted to hit a certain quality bar, a certain latency bar, and, uh, you know, NotebookLM was, uh, already out and the TTS models that were powering NotebookLM are, were very, very good.

Uh, but we wanted an aspect of native, so it was native audio in but TTS out, and we still have that architecture available through the Live API, but then now we just released audio-to-audio architecture.

**Kwindla Hultman Kramer** [16:30]
I mean, the infrastructure for this stuff is so interesting-

**Shrestha Basu Mallick** [16:32]
Yeah

**Kwindla Hultman Kramer** [16:32]
... because you're always balancing latency, cost, output quality. There, there's no free lunch.

**Shrestha Basu Mallick** [16:39]
Yeah. Um, and, uh, other things like, uh, multilinguality. Coming back to your question earlier, Sam, a lot, we had a lot of users asking us for, say, better German language support-

**Logan Kilpatrick** [16:50]
Mm-hmm

**Shrestha Basu Mallick** [16:50]
... or something, which hopefully now we've delivered on with these models, so.

**Logan Kilpatrick** [16:55]
Yeah. Yeah.

**Kwindla Hultman Kramer** [16:55]
But now you have audio-to-audio in the, in the Live API as well.

**Shrestha Basu Mallick** [16:59]
In the Live API only is where we have the na- native audio output models.

### Voice Infra

**Logan Kilpatrick** [17:03]
Yeah.

**Shrestha Basu Mallick** [17:04]
Interesting.

**Sam Charrington** [17:04]
Now, continuing to pull on the component versus single model thread a little bit, when I think about voice, I think about it as being an area where to deliver solutions, you need to surround that strong model with a lot of voice-specific infrastructure-

**Shrestha Basu Mallick** [17:21]
Yeah

**Sam Charrington** [17:21]
... uh, that is, you know, I'm imagining challenging the scale.

**Shrestha Basu Mallick** [17:25]
Yeah, yeah.

**Sam Charrington** [17:26]
So Shrestha, can you talk a little bit about that? And maybe we can have Kwind talk about that from his perspective.

**Shrestha Basu Mallick** [17:30]
So the first thing that comes to mind is, of course, the voice activity detection models that we have, and we've done a lot of work like finessing that model server side. But we've also learned that we need to provide some knobs to developers, so now developers can actually tune the sensitivity on our voice activity detection model, as well as, you know, how much of the prefix pad, like how much of a time duration at the beginning, at the start or stop of saying things.

And, uh, we also have a mode where now you can, where you can disable our voice activity detection and bring your own.

**Logan Kilpatrick** [18:07]
Mm-hmm.

**Shrestha Basu Mallick** [18:07]
But I think the larger point that you're touching on, Sam, that I do want to mention is it is really, really hard to bring all these components together-

**Sam Charrington** [18:16]
Yeah

**Shrestha Basu Mallick** [18:16]
... and still get latency down to where it needs to be, you know, in the five hundred to seven hundred millisecond range. Like it's, it's one of the hardest things we've had to do with the Live API.

**Kwindla Hultman Kramer** [18:26]
What we see is that the shape of building these real-time voice agents is a different d- set of developer problems than the shape of, you know, non-real time or text mode things. One of the fun things about partnering with Shrestha and DeepMind is we work on this open source framework that people use to build these kind of production voice systems, and so we try to solve problems at the framework level like turn detection, like context management.

**Shrestha Basu Mallick** [18:49]
Yeah.

**Kwindla Hultman Kramer** [18:49]
As the models get better, as the use cases get more clear, some of those features migrate from the framework into the APIs, which makes life easier for developers. The use cases at the same time continue to broaden out, and so there's more things for the framework to do.

### Dev Framework

**Kwindla Hultman Kramer** [19:03]
So we're sort of filling the top of the use cases, building blocks, developer experience funnel, and pushing down as we all get better and we all figure out what this new world looks like.

**Shrestha Basu Mallick** [19:13]
And maybe this is also a good segue into WebSockets versus WebRTC, Quinn.

**Kwindla Hultman Kramer** [19:17]
Yeah.

**Shrestha Basu Mallick** [19:18]
Yeah.

**Kwindla Hultman Kramer** [19:18]
You know, there, there's so much infrastructure, like one-- I, I, for my whole career, I've been building, you know, large scale, low latency network stuff. What we saw from my perspective when we started to see the possibilities of voice AI was you need this packet routing, like, down underneath the inference layer even.

**Swyx** [19:33]
Oh my God.

**Kwindla Hultman Kramer** [19:33]
There's, like, the AI inference stuff, but then there's the just how do you move the audio and increasingly video around the internet. And so there's a whole new generation of developers who are interested in these networking protocols because voice AI and now real-time video are so interesting.

Which is super fun for me, 'cause, like, I've always thought moving packets around is one of the most fun things you can do on the internet.

**Swyx** [19:53]
Yeah. Seven layers of the OSI stack.

**Kwindla Hultman Kramer** [19:55]
Exactly. Exactly. Yeah, at, at pretty, pretty demanding real-time latencies, as Shrestha is saying. Like, human beings expect you to respond in a conversation in 500 milliseconds or so, and if we're talking to an AI, we don't relax that assumption.

We bring our assumptions about human conversation into that experience of interacting with an AI.

**Shrestha Basu Mallick** [20:16]
Yeah, or not respond. So one-

**Kwindla Hultman Kramer** [20:18]
That's, that's a great point.

### Proactive Audio

**Shrestha Basu Mallick** [20:19]
Yeah. So, like, one of the features that we've pushed out, a little more experimental but would love for people to test it, is what we're calling proactive audio, and it's available only in the native audio-- uh, in the audio-to-audio architecture right now.

And what this feature does is it's trained not to respond to irrelevant audio cues.

**Swyx** [20:40]
Okay, so it's like a refuse or kind of?

**Shrestha Basu Mallick** [20:42]
Yeah, or you could call it directionally like semantic voice activity detection, right? Um, so basically, uh, yeah, like, uh, let's say I'm talking to the AI and then Quinn comes and asks me a question, and I respond to Quinn.

It'll know when not to respond. Um, so yeah.

**Swyx** [20:58]
I saw that in one of the demos. The AI seemed to ignore a background question from someone else-

**Shrestha Basu Mallick** [21:04]
Yeah, yeah, yeah

**Swyx** [21:04]
... in the, the video.

**Kwindla Hultman Kramer** [21:05]
I think there's two threads to pull on there. One is that's another great example of things that we had to work really hard at the framework level to implement. It's much, much better if it actually migrates down into the model or the API.

The, the other is part of the magic there is this semi-separate feature, but I think they're multiplicative, of now your models can actually recognize two different people just based on their voices. You and I were playing with that.

**Swyx** [21:29]
They have to. They-- Yeah.

**Shrestha Basu Mallick** [21:30]
They-- This is not officially supported yet.

**Swyx** [21:33]
Oh, okay.

**Shrestha Basu Mallick** [21:33]
The model just does it.

**Swyx** [21:35]
But just try it, right?

**Shrestha Basu Mallick** [21:36]
Just, I mean-

**Swyx** [21:37]
Just try it and get feedback

**Shrestha Basu Mallick** [21:38]
... you try it, you'll give us feedback as in-

**Kwindla Hultman Kramer** [21:38]
But is it okay to talk about it? 'Cause it's-- it might be my single favorite thing you can do with these models that you previously have not been able to do.

**Shrestha Basu Mallick** [21:46]
I mean, you can talk about what you've observed, Quinn. I'm just saying it's not officially-

**Swyx** [21:49]
Not officially supported.

**Kwindla Hultman Kramer** [21:50]
And what specific models are we talking about? Because speaker identification and diarization has always been really hard-

**Shrestha Basu Mallick** [21:55]
Yeah

**Kwindla Hultman Kramer** [21:56]
... for these models.

### Async Calling

**Shrestha Basu Mallick** [21:56]
No, this is all the-- It's called, uh, gosh, like, model naming now has become so weird. It's called the native audio dialogue. You'll see it in the live API, uh, but that's the model.

**Swyx** [22:07]
Excellent.

**Shrestha Basu Mallick** [22:07]
And then, uh, you know, and to your point again, Sam, about architectures, one thing that we've launched on the cascaded architecture that we hope to eventually bring, uh, to the native audio as well is asynchronous function calling. Uh, so earlier the way it used to work is y- if you wanted the model to do a function call, you'd have to wait for the response, and now you can set a non-blocking parameter, and the model can go off and execute the function in the background, and-

**Swyx** [22:34]
Love you. I love you so much

**Shrestha Basu Mallick** [22:35]
... and then it's like, yeah.

**Swyx** [22:35]
Yeah. That's great. We do have to wrap up, um, so I think one fun thing that we can do to wrap up would be a wish list for, like, next year's I/O. What would be one thing that you would wish?

It doesn't have to come true, but, you know, wish happens, with Gemini.

### Wrap

**Kwindla Hultman Kramer** [22:48]
Well, I was hoping for Gemini 3.0 at this I/O, so maybe Gemini 5.0 at the next, the next I/O.

**Shrestha Basu Mallick** [22:55]
You gotta, you gotta tell us what you mean by Gemini 5. What do you want in Gemini 5.0 then? Uh, I'll let Quinn go.

**Kwindla Hultman Kramer** [23:02]
I'll just put on my hat as representative of a big community of people-

**Shrestha Basu Mallick** [23:05]
Yeah

**Kwindla Hultman Kramer** [23:05]
... building this stuff, more and more languages, because AI is global.

**Swyx** [23:08]
Oh, yeah.

**Kwindla Hultman Kramer** [23:09]
And there are so many communities all over the world that are starting to do this stuff.

**Swyx** [23:12]
Can we do language LoRAs? You know-

**Shrestha Basu Mallick** [23:14]
Yeah

**Swyx** [23:14]
... it's, it's hard to stuff everything in one language, and it's, uh, in one model. Yeah, yeah. Okay, but yeah, and then, Shrestha, you have to-

**Kwindla Hultman Kramer** [23:19]
But they're, they're building one model, as they said. They're building the one universal model.

**Shrestha Basu Mallick** [23:23]
Yeah. I, I think that, that would be bor- boring answer, but I think really-

**Swyx** [23:26]
Languages?

**Shrestha Basu Mallick** [23:27]
... more and more-- No, languages, I mean, wasn't I telling you earlier? Like, we, we officially support 24 languages, but you can try talking to the model in Klingon, and it'll respond to you.

**Swyx** [23:37]
Yeah.

**Shrestha Basu Mallick** [23:37]
So I think we'll get there way before next I/O. But I just think more and more capabilities into the main model is what I would say. I'll have to think about this.

**Swyx** [23:47]
Yeah, yeah. It's, it's a fun parlor game.

**Shrestha Basu Mallick** [23:49]
The real answer is-

**Swyx** [23:49]
Uh, but it also helps people align as to what is possible and what's coming up. Thanks for your time, everyone. This is, uh, very hastily organized, but I'm glad that we could make this happen. And it's nice to actually see Sam in person.

**Kwindla Hultman Kramer** [24:01]
Same.

**Shrestha Basu Mallick** [24:01]
Yeah.

**Swyx** [24:02]
Yeah.

**Shrestha Basu Mallick** [24:02]
This makes you think I am only the PM for the live API, but we did not get to talk about some of all of the other releases as well that-

**Swyx** [24:09]
We'll save that for your, your talk at World's Fair.

**Shrestha Basu Mallick** [24:12]
Yeah.

**Swyx** [24:12]
Uh, and you, you guys are all speaking, and, uh, and we'll, we'll be podcasting as well, so.

**Shrestha Basu Mallick** [24:15]
Sounds good.

**Swyx** [24:16]
Yeah.

**Shrestha Basu Mallick** [24:16]
Yeah.

**Swyx** [24:17]
Yeah, yeah. All right. That's it. Thank you so much.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
