# Gemini 2.0 Flash and Flash Thinking: the new SOTA models for the agentic era

Latent Space · 2025-02-28

<https://addtry.com/18edae3d-44ec-4eaf-9733-8db09d5a2bf5>

Logan Kilpatrick, Google AI Studio product lead, returns to discuss Gemini 2.0 Flash and Flash Thinking, the new models that balance frontier capabilities with cost efficiency for the agentic era. He explains the pricing strategy: Flash at 10 cents per million tokens (simplified from tiered pricing), Flash Lite preserving the 7.5-cent narrative for low cost, and Pro pushing the frontier. Flash Thinking, co-led by Noam Shazeer and Jack Rae, scales inference-time compute and already shows rapid improvements, with base model advances coupling with RL-based reasoning. Swyx reports that Gemini Flash outperforms o3-mini on long-context summarization, calling it a 'reporting model.' The multimodal live API enables real-time voice and vision interactions, and a memory layer mimicking Astra is in development. Grounding with Search as a Tool is also highlighted for agentic search use cases.

## Questions this episode answers

### What is Google's Gemini Flash Thinking and how is it different from other reasoning models?

Logan Kilpatrick explains that Flash Thinking is the Gemini 2.0 Flash model with built-in reasoning, not a separate series, co-led by Noam Shazeer and Jack Rae. It scales inference-time compute using reinforcement learning, showing rapid improvements. The model benefits from base capabilities plus reasoning, and multimodal and long context features further enhance it.

[8:25](https://addtry.com/18edae3d-44ec-4eaf-9733-8db09d5a2bf5?t=505000)

### How much does Gemini 2.0 Flash cost compared to the previous version?

Logan says Flash pricing increased from 7.5 cents to 10 cents per million tokens, but they eliminated the tiered pricing that doubled cost for inputs over 128K tokens. Now it's a flat 10 cents regardless of token volume. They also introduced Flash Lite to keep a low-cost model at the original price point.

[1:56](https://addtry.com/18edae3d-44ec-4eaf-9733-8db09d5a2bf5?t=116000)

### What does 'experimental' mean for Google's Gemini models?

Logan explains that experimental models are a fast release valve to get models into developers' hands quickly, not intended for production. They may change without notice; sometimes new versions hot-swap the old one with no ability to stay on it. It signals active improvement and validation of internal gains with real-world use.

[5:59](https://addtry.com/18edae3d-44ec-4eaf-9733-8db09d5a2bf5?t=359000)

## Key moments

- **[0:00] Intro**
  - [0:39] Logan Kilpatrick's first in-person podcast was Latent Space with Swyx and Alessio.
  - [1:05] "I lead product for Google AI Studio." — Logan Kilpatrick
- **[1:38] Model Suite**
  - [1:56] Gemini 2.0 Flash input price rises to 10 cents per million tokens, but Flash Lite preserves 7.5 cents for cost-sensitive builders.
  - [3:37] Gemini 2.0 Flash outperforms the previous Pro model on every dimension while being cheaper.
  - [4:56] Q: Where is Gemini Ultra? A: It faces practical bounds of size and cost, making Flash the right trade-off for production.
- **[5:10] Reasoning Models**
  - [7:16] Q: Will Cursor Composer support Gemini reasoning models? A: Logan Kilpatrick will talk to the Cursor team to advocate for it.
- **[8:23] Noam's Return**
  - [9:05] "It feels like the GPT-2 era" — Logan Kilpatrick on reasoning model progress.
- **[11:38] Long Context**
  - [11:38] Long context with reasoning will unlock finding 100 things in a two-hour video, predicts Logan Kilpatrick.
- **[15:01] DeepSeek Takeaways**
  - [15:45] DeepSeek's paper revealed that MCTS and process reward models are unnecessary for strong reasoning, notes Swyx.
- **[16:38] Multimodal Live**
  - [16:42] Google's multimodal live API enables real-time voice and screen sharing for AI co-presence at aistudio.google.com/live.
  - [18:58] Logan Kilpatrick predicts AI chat will shift to messaging apps like text and email to onboard the next 3-4 billion users.
- **[20:25] Nano & Local**
- **[21:56] Astra Memory**
  - [22:53] Google is building a memory service for Gemini that goes beyond embedding-based RAG and supports deletion.
- **[24:36] Search Grounding**
  - [24:47] Google's Search as a Tool feature lets developers build deep research-like experiences with Gemini.
- **[25:50] Wrap-up**
  - [26:01] Logan Kilpatrick runs two AI podcasts: "Around the Prompt" and "Google AI Release Notes."

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **Logan Kilpatrick** (guest)

## Topics

Reasoning, Language Models, Search

## Mentioned

Anthropic (company), DeepMind (company), Google (company), OpenAI (company), Astra (product), Claude (product), Claude Sonnet (product), Cursor (product), Gemini Flash (product), Gemini Flash Lite (product), Gemini Flash Thinking (product), Gemini Nano (product), Gemini Pro (product), Gemma (product), Google AI Studio (product), Multimodal Live API (product), NotebookLM (product), Search Grounding (product), Search as a Tool (product), o3-mini (product)

## Transcript

### Intro

**Alessio** [0:04]
Hey everyone, welcome to another Lightning Pod. This is Alessio, partner and CTO at Decibel, and I'm joined by Swyx, founder of Small AI.

**Swyx** [0:11]
Hey, and today is a, is a long time coming. We have the epic return of Logan Kilpatrick, now from Google. Welcome.

**Logan Kilpatrick** [0:18]
Thank you. I'm excited. It's been-- it feels like almost two years now or something like that since we, since we last did an episode, so it's great.

**Swyx** [0:25]
Yeah. Yeah. Um, thanks for helping us kick off Latent Space, to be honest. Like, I don't think people know or, like, have looked back that far in the, in the archives, but, you know, Latent Space has been one of the most successful things that Alessio or I have done, and, you know, you helped to kick that, kick that off for us.

So thank you so much.

**Logan Kilpatrick** [0:39]
It was my first in-person podcast, so it was y'all's first podcast-

**Swyx** [0:42]
Yes

**Logan Kilpatrick** [0:42]
...it was the first-- I'd done a bunch of virtual podcast recordings, but it was the first time I ever sat down and, like, did a legitimate podcast. So, uh, a special place in my heart that we got to, we got to do that together.

**Swyx** [0:52]
Yeah. Uh, since then, it's gotten a lot harder to, to do any podcast with OpenAI. They've locked down their PR a lot. But, you know, thank-thanks for, thanks for jumping on. Uh, so you're, you're now, uh, uh, AI studio lead, or what is-- what-- how do you describe your role?

**Logan Kilpatrick** [1:05]
Yeah, I, I do a lot of things, and I think people still think that my job is DevRel. My job's not DevRel. I, I sort of made the official transition as I joined Google to a product role, and today I lead product for Google AI Studio.

So really focused on, like, how do we build the right product for developers? How do we build the best product for developers, specifically, like, a AI developer platform? I still do a bunch of DevRel-y stuff just because it comes naturally to me, and I think it's, like, honestly the way to be the best is go and talk to developers, which is a lot like DevRel.

But yeah, working on, working on that and, and bringing Gemini models to the world.

### Model Suite

**Swyx** [1:38]
Yeah. Uh, I mean, let's, uh-- we're here to talk about it, Gemini 2, the, the, the, the, the sort of hottest, latest news. I think one of the interesting things that people are trying to figure out is the pricing strategy, where Pro is, and then what Flash Thinking is, and why there, there's, like, Flash Lite now.

How do you think about the product suite?

**Logan Kilpatrick** [1:56]
Yeah. So, you know, s-some of these are, are sort of our-- maybe we're in our, our own head too much with some of the mo- the decisioning on pricing and stuff like that. The sort of general principle is, you know, Flash, we really wanted to give developers, like, the best model without sort of a, a significant price increase.

But just, like, because of a bunch of trade-offs for that model, like, the price had to go up, so it went from seven and a half cents per million tokens to 10 cents. The sort of way that this was offset was we used to distinguish—I don't know if folks are familiar with this—but we used to distinguish based on input-

**Swyx** [2:27]
Per token

**Logan Kilpatrick** [2:27]
...or input token volume. So it was, like, over a hundred and twenty-eight K tokens would be twice as much, so it was, like, 15 cents. Um, now it's just, like, 10 cents straight up, whether it's one token or a million tokens, which should hopefully make people's lives simpler.

Um, but as far as why Flash Lite exists, we sort of spent the last year leaning into this Flash narrative of, um, you know, having the lowest, the lowest cost model with the highest performance in the ecosystem. And it kind of just felt weird to have, like, strung people along for this narrative and then, like, change the price point.

I think, like, a lot of people baked in this, like, seven-and-a-half cents into their mental model and, and we wanted to-- like, Flash Lite was really like, we wanna preserve this, this story that we've been telling developers, which is, like, basically e-eliminate the economic burden to build with AI and, and be able to get a really great model for, for cheap.

So that was the Flash Lite story. And then now with Pro, it's how do we push the frontier? And the, the exciting thing in the Pro models always end up being more expensive, but the exciting thing is we keep seeing this, this, this jump where the capabilities of the Pro models get superseded by the Flash models as the next generation comes.

So if you look at our 2.0 Flash model is actually o-on, like, every dimension a stronger and better model than the previous one, and it's X cheaper or whatever it is than the previous Pro model. So, like, I think that trend is going to continue, but we need to do the frontier research to continue to make the Pro models great and have the sort of cost benefit trickle down to the, the smaller size models.

**Swyx** [3:58]
Yeah, I think, like, the-- right now the meta game is really that the model labs have, like, the big model. We don't know, we don't know much about it except that I think now we know it's an MoE, um, and it's, it's gonna be as big as it can possibly be, and they're always the teacher models for the distillation artifact, which is the actual thing that you, you are supposed to use for production inference.

Um, I feel, I feel like there's maybe other variations of this where there's this distillation of the distillation, which is the Flash Lite. But um, you know, I, I think, like, it seems like there will always be this teacher model.

Like, you know, the rumor is that Anthropic has Claude 3.5 Opus or-- and is distilling for, for Sonnet, right? And Sonnet's the one that actually people use. I think it's an interesting pattern that, you know, we no longer encourage the big model in production because it's, it's too slow or it's too expensive, whatever, whatever that is.

Yeah.

**Logan Kilpatrick** [4:47]
Yeah, this, this has been interesting, and I think, like, we've had this conversation about, like, an Ultra series model internally where it's like what would the-

**Swyx** [4:55]
Well, yeah, where's Ultra?

**Logan Kilpatrick** [4:56]
Yeah. Where, where is Ultra is, is the comment that developers have asked, and I think it's like, it's a, it's a good question. But the, the real-- there's just, like, a lot of practical bounds around, like, would, would something of that size and cost, like, actually make any sense in production?

And I think the takeaway is, like, Flash is doing a great-- like, developers want better models, so we'll give developers better models, but Flash is really this, like, right trade-off and balance of all these things. And now we actually have the reasoning model, so it's like we have Flash Thinking, uh, which is a sort of skew sort of related to the 2.0 Flash model that can think and, and uses inference time compute in order to get better performance across a bunch of stuff like coding, math, science, et cetera, et cetera.

### Reasoning Models

**Logan Kilpatrick** [5:36]
And it feels like that is more likely in the, in the medium term going to be the thing that, like, continues to just, like, rapidly scale up relative to, to the other capabilities.

**Alessio** [5:47]
One question I had, what does the experimental keyword mean when it comes to Google models? Does it mean it might not be supported in the future? Uh, does it mean it's just early? Like, how should developers think about it?

**Logan Kilpatrick** [5:59]
Yeah, we're, we're trying to give ourselves the two sort of pieces to balance is like, on one hand, we want to ship models as fast as possible. Like the experimental model release train started as like with the idea of how do we eliminate the delta between like the model training process and the model that developers actually have.

So it's like a s- it's the super fast release valve to like get a model out the door into the hands of developers. Framing it as experimental is, is really to try to signal that like, one, you shouldn't use this thing in production.

Um, like it's, you know, rate limits are usually the limitation to stop people from doing this. But two, like the model might change and in certain cases like we come out with a new rev of the model, we literally just like hot swap the old one with the new version of it.

Like there's no opportunity to stay on the old one. Um, and this usually comes when like we're-- it's not, not really like an AB test, it's like we are clearly hill climbing on something, like there's still a lot of room to go.

Let's do some experimental model launches to like validate that, you know, the gains that we're seeing internally like actually correlate to the gains that developers see on, on a bunch of use cases that they care about.

**Alessio** [7:03]
Yep. And then the other thing people in our Discord were saying the model is cracked, uh, for code, uh, which sounds great. Uh, I try to use it with Cursor Composer, but Cursor doesn't support, uh, the reasoning models yet.

Is that something that you work with Cursor on? Like do you know if that's coming, kind of like what the limitations are? It feels like to me that's, for me, it's one of the main barriers to adopting these reasoning models for like most of my use case is coding and like if it doesn't work in Composers, like it needs to be so much better to get out of it that it's usually not worth it.

So any thoughts there?

**Logan Kilpatrick** [7:35]
Yeah. I'll, I'll text Michael and Aman after this and we'll see if we can convince them to turn it on for Composer. I think we, we've been collaborating closely with them and making sure they get access to the latest models and all that stuff.

I don't have, I don't have the specific context on, on the situation with Composer, but I've played around with it a bunch and it's, it's sick, so hopefully we can get the Gemini models in there.

**Swyx** [7:53]
It's funny. If it, if it doesn't exist in Cursor, it doesn't exist to us.

**Logan Kilpatrick** [7:57]
I think that's fair. It's, it's interesting like how quickly that's become the case. Like Cursor has been-

**Swyx** [8:02]
Yeah

**Logan Kilpatrick** [8:02]
... just continues to, to absolutely crush it, which is awesome to see.

**Swyx** [8:06]
Yeah. Uh, they're the fastest growing SaaS in the history of SaaS. One to 1 million ARR to 100 million in one year. And you know, we were talking about how y- you know, we were the first in-person podcast that you did.

We were also the first with Cursor, and just to see them grow so quickly beyond that, obviously we wish we invested and all that. But it's nice to at least have a long-term relationship with them. Yeah. Let's talk about reasoning.

### Noam's Return

**Swyx** [8:25]
A lot of people are very interested. Noam Shazeer is back. Uh, what's he cooking?

**Logan Kilpatrick** [8:29]
Yeah. Noam and, Noam and Jack are, uh, Jack Rae-

**Swyx** [8:32]
Jack, Jack who? Okay. Who, who's that?

**Logan Kilpatrick** [8:34]
Jack Rae. Yeah. He's, he's been a long time, uh, DeepMind research scientist, was, was previously a pre-training person. We, we actually overlapped at OpenAI together a little bit and then is now back at DeepMind with, with Noam co-leading the, the reasoning effort.

And really the intent is like, you know, it's now obviously clear to everyone that like you can scale up these reasoning models and like this is the new scaling frontier. I, I tweeted the other day that it's the-- it feels like the GPT-2 era, the like order in which the timeframe in which model progress can actually happen.

If you look at our first reasoning model launch in December, less than 20 days or less than 30 days later, we had the second version of the reasoning model publicly available and, and other versions already internally available. Um, and like the s- the scaling is like directly up and to the right, um, which is exactly what you want to see.

And it's, it's incredible how, how quickly this is happening and also like how quickly, you know, all of those threads about just like how quickly people adapt to this new model capab-- like it's just like the obvious outcome now is that these models are going to keep scaling up, which is really funny to see at the same time of this narrative around, you know, the models hitting a wall.

Literally it was like two months ago. Two months ago, it was like the entire media narrative was like, "There is no more value to be created with AI. It's hit the wall. This is the end of it." And in the last two months, it's like back to the pre-training era days of like, you know, you can just infinitely scale up RL and then all of a sudden the models just keep getting better infinitely, which is obviously not, is an exacerbated version of the narrative now, but it is what it sort of feels like to a certain extent.

**Swyx** [10:08]
Yeah. I, I kind of put it as like it's a different dimension of scaling. We, we did hit a wall, which is, you know, the way I put it is no model larger than 2 trillion parameters has, has been shipped, uh, 'cause we, we are no longer scaling parameter size really, like we're just scaling in inference time compute, and that's, that's okay.

But it's a different, different dimension.

**Logan Kilpatrick** [10:25]
I'm hopeful pr-- I think, um, my, my instinct is that pre-training will keep scaling. Like I think we're, we're seeing a bunch of progress on pre-training stuff, which is super exciting. Like I think the, the best outcome for the world is pre-training keep scaling.

And like actually if you look at the tangible examples of this, and, and I was having a conversation with Jack a f- a few weeks ago and talking about sort of the interplay between scaling the core model capabilities in the base model itself and scaling this RL thinking stuff that's happening, um, and how important the interplay is between those.

It's like you get this like ultimately the base model, and this is why I think like not, at least today, like we don't have a separate series of thinking models. Like it's not Gemini, you know, whatever, some other name.

It's like truly the Gemini 2.0 Flash models with thinking built into them. Um, so you benefit from all of the like base capabilities of the model and that ability for the model to reason over the tokens. You get like both ends of the scaling curve, which is like as the base model capabilities improve, you get that value, and then you also get the added sort of RL thinking chain of thought stuff that's happening, which is just super cool.

And then the capabilities like multimodal and long context start to really matter a lot in those, in those examples as well.

**Swyx** [11:38]
I mean, Google's long context, like a million, 2 million crazy. Like it's, it's, uh, it's wild.

### Long Context

**Logan Kilpatrick** [11:43]
Yeah. And like it feels like, a- again, like back to this, the thread around these capabilities, like it feels like long context with reasoning is like finally going to be that thing where like it actually just like blows the lid off of it and like it, it makes the use case really come to life.

Because like the challenge historically has been long context. If you read the original paper that we put out when long context landed, it was, you know, the models are pretty good at like needle in a haystack like tasks of like attending to, you know, one or two things or a small number of things in the million to 2 million token context window.

But as you scale that up to like, you know- Find me 100 things in this two-hour video, it, it becomes really hard for the model to do that. And I think reasoning is the moment where this starts to change, where the model's actually able to go and, and look over that number of, of things in a, in a single context window, which is really cool.

**Swyx** [12:31]
Yeah. Uh, I'll, I'll, I'll offer an anecdote which isn't science, but, you know, it's what, it's what we have. Which is I'm pretty public about AI News and how I run, um, basically regression tests against every single frontier model, uh, every day.

And the moment Gemini Flash came out, I ran it against 3 Mini as well, and Gemini Flash is winning. It is so much better at context utilization than o3-mini. It's actually, like, a bug in o3. Like, I I've told them about it, and they're looking into it.

But I'll endorse that Gemini Flash is, like, currently the best summarization and reporting model. I call this a reporting model because, like, I think people have this mental image of summarization as, like, we'll take a long document and, and trim it down to a f-few bullet points.

But what reporting is, is a, is a separate problem, is very, very long context, um, you know, hundreds of thousands of, of context, all different news stories and all different ways of ranking and clustering them that the Gemini is just way better at.

And so, so, you know, my-- on my sort of small benchmark, um, like, it-it's Gemini Flash is just incredible at context utilization. I, I don't have a, a better way of expressing this yet, but yeah, that's my feedback.

**Logan Kilpatrick** [13:33]
I feel like you should make this a public benchmark. I feel like everyone now has their own, like, external leaderboard/benchmark, and I actually feel like this would be, like, a really cool version of this where, like, I can go to your newsletter and, like, see the versions generated by all the different models and, like, whoever the, like...

Let people vote on that. You have your own LMSYS arena leaderboard of, like, people voting on the best version of the, of the newsletter on a given day would be super cool.

**Swyx** [13:57]
I have-- So I've gotten back and forth about this. Uh, obviously, I have the raw data every day. So one thing-- So we did a podcast with, with NotebookLM Risa, and what's the other person's name? I... Oh, God, I can't remember his name right now, sorry.

But anyway, so, like, when people ship AI products, actually, it helps to hide the process as much as possible, like not expose the guts because, um, it really helps you focus everyone on the final output and, and for you to just iterate on the final output without the constraints of exposing all the bells and whistles.

Like, one of the successes of NotebookLM, like imagine if they, they like re-released like a hundred different little knobs and tweaks to, to this from the start instead of the magic like one-click experience, right? I, I realized that less is more a lot-- for a lot of AI products.

**Logan Kilpatrick** [14:41]
Yeah, I think that's fair. My push for you on this is I think your, your audience is like the-

**Swyx** [14:47]
Yeah, yeah

**Logan Kilpatrick** [14:47]
... AI tinkerer super fan builder who like, who wants those. And like, I think that persona wouldn't mind seeing the sort of guts behind the scenes because like ultimately like we're all learning like what are the different capabilities of the model as much as that is worth it.

### DeepSeek Takeaways

**Logan Kilpatrick** [15:01]
Um-

**Swyx** [15:01]
Yeah. Yeah. Um, any other things on reasoning? Oh, any takeaways from like DeepSeek? Like did they get it right? Is Google pursuing a different path? Um, I, you know, I, I know, I know there's like not a ton that's public about Flash Thinking, but like, you know, i-is DeepSeek the right path?

Or they're like-- Are you seeing multiple paths going on?

**Logan Kilpatrick** [15:21]
Yeah. I think my sort of takeaway from the DeepSeek moment is a few things that we, we sort of knew, which is, you know, countries that have large investments in the technology sector like can and should build AI models.

Like I think it's like not surprising to me that given the, the position-

**Swyx** [15:37]
Oh, software AI? Yeah.

**Logan Kilpatrick** [15:39]
Yeah, that, that China's in, like they're gonna have companies that build great models. Like that seems like a, like a super obvious outcome to me. Um-

**Swyx** [15:45]
So I'll, I'll interrupt there. I, I was just talking more about like the technical insights from the paper, right? For example, a big, a big aha moment was no MCTS. Like they, they tried it, didn't, didn't work. No process reward models.

Tried it, didn't work. And for like most of the past year, we were talking about those things as like the future of reasoning, and it turns out they're not needed at all. I feel like Noam being able to land at Gemini and ship Flash Thinking in like two months is also a sign of like, yeah, like this is-- this, this was a, a head fake originated by Q*, you know, this magic secret code, code name that everyone extrapolated based on.

**Logan Kilpatrick** [16:20]
Yeah, I'll, I'll let Jack... I'll, I'll connect you with Jack and Noam. You should have them on the podcast to talk about when we, when we do the GA launch, hopefully very soon.

**Swyx** [16:27]
Mm-hmm.

**Logan Kilpatrick** [16:28]
I think doing a deep-

**Swyx** [16:28]
Oh, Flash Thinking?

**Logan Kilpatrick** [16:29]
... diving wall would be... Yeah, yeah, for-

**Swyx** [16:30]
Okay

**Logan Kilpatrick** [16:31]
... Flash Thinking, um, would be, would be a ton of fun. I don't wanna, I don't wanna be the one to, to spill any of the details before then.

**Swyx** [16:38]
Got it. Uh, shall we touch on the real-time multimodal stuff?

### Multimodal Live

**Logan Kilpatrick** [16:42]
Yeah, yeah. I think that'd be, that'd be super fun.

**Swyx** [16:45]
Something that of-people often overlook. So I'll, I'll give a huge shout-out to Paige Bailey, who it also seems like you're forming a proper DevRel team with Paige and with Christina Warren and, uh, I forget who else. They all announced on the same day that they were joining Google.

Um, that Paige, like on the day of Gemini 2.0 Flash's launch, came over to our New Rips event and demoed it, and it was, uh, really incredible.

**Logan Kilpatrick** [17:05]
Yeah, this-- If folks haven't tried this before, aistudio.google.com/live, it's the real-time live experience, and I think it is a front-end experience that's powered by the multimodal live API, but really hits on this sort of a couple of broader themes that I think matter a lot.

One of them is just like this AI co-presence, which I think we're getting... Like one of the big limitations of AI systems being useful is like they don't have the context. Like they can't see what you see. They don't have access to the stuff that you, that you do.

And I think the models actually being able to see and you being able to interact, interact with them using your voice, I think bridges us closer to this world where we're actually not limited by the context 'cause the models can see and do all the same stuff that I'm able to do.

But also, I think it's just like- It's one of those experiences that's so fundamentally visual. Like, you can show the model what you see on the screen. Uh, you can, you know, have it connect to your camera, you can send text to it, you can talk to it, all this stuff, which I think feels, feels like the AI experience, uh, which is really interesting to me.

Like, it's not just chat, it's, it's truly this, like, different experience than you get today, and I don't think we've seen a lot... Like, there's a couple of products. It's interesting, like, how much this is in parallel to, like, the chat app phenomenon.

Like, everyone, you know, it started off with a couple of chat apps that, you know, choose your favorite chat app that was, like, really good and super popular and widely known, and then that technology sort of broadly assimilated into society and into products and such.

And I think we're gonna see the same thing with this, like, showing that... Like, if every browser in the world doesn't have this experience in the future, I'd be surprised. If every IDE in the future doesn't have the ability to, like, "Here, let me just show the model what's happening here, share my screen, talk to it live," like, I think that's gonna be a huge miss.

So we're gonna see more and more of the tools that people use, um, enable AI to show up as a thought partner in, like, a very multimodal way, which, uh, will be, will be super exciting.

**Alessio** [18:58]
Yeah. Do you think chat is just, like, technical debt in a way? Like, we started there, and now people are just chatting, but, like, yeah, when I look at, like, the Gemini real-time stuff, it's like, why would you chat?

**Logan Kilpatrick** [19:10]
Yeah.

**Alessio** [19:10]
It doesn't really make sense to me, but I feel like people are kind of stuck in the UX that they started with, and I'm curious what it's gonna take to get out of it.

**Logan Kilpatrick** [19:18]
I think there's still a lot of things where chat probably makes a lot of sense. Uh, for what it's worth, like, I personally find, and, and maybe this is just my super personal bias, but, like, I feel like I'm in meetings, working, talking all day, and, like, it's really nice to just, like, chat with AI sometimes and not have to talk to it.

It's, like, kind of exhausting to always have to talk to the models. And, like, I, I think, like, through that lens, there's gonna be a lot of, of chat experiences. I think the chat experience that I'm most excited about, and, like, I've been pushing on this for, for years, and I feel like we still haven't even gotten there, and it's early sort of nuggets of this product success area working, which is, like, bringing the chat to the places that you're already chatting, like text, email, et cetera, et cetera.

Like, that's the place in which I wanna use. Like, the models are already good enough to do that. The interfaces, like text message has multimodal image audio. There's, like, app experiences built into a bunch of the messaging apps.

So, like, you can fully do that experience today, and, like, I feel like that's the chat experience that is going to make the most sense in the future, and especially critical to onboarding, like, the next three to four billion users.

### Nano & Local

**Logan Kilpatrick** [20:25]
Like, I don't think those people are going to come in through, you know, some front-end website somewhere. Like, those people are gonna come in through audio from a telephone to texting to email. Like, that's just the most obvious outcome.

**Alessio** [20:39]
Why are people not using the Gemini Nano thing more? That was one of the most exciting thing for me from Google I/O last year, and thanks again for, for having us there.

**Logan Kilpatrick** [20:48]
Yeah, yeah.

**Alessio** [20:49]
But it seems like, I mean, it shipped, right? But I don't really see a lot of adoption yet.

**Swyx** [20:53]
I think it's feature flags, right? Or I don't know. You can correct us.

**Logan Kilpatrick** [20:57]
Yeah, I'm not actually sure what the, what the deployment is for, for Gemini Nano.

**Swyx** [21:01]
So, uh, yeah, I'm really-- I'm pretty close to the browser guys, and, uh, it seems like it's still feature flag for some reason. It's like-- So it's in Chrome Canary.

**Logan Kilpatrick** [21:08]
Mm-hmm.

**Swyx** [21:08]
You have to, you have to turn it on with a diff- diff- developer flag. You can't just, like, do, like, you know, window.ai and then something. Um, it's not live yet for everybody.

**Logan Kilpatrick** [21:16]
Yeah.

**Alessio** [21:16]
Okay, that makes more sense.

**Logan Kilpatrick** [21:17]
It's a good push. I'll, I'll ping some folks and find out what the status is. I don't know if it's like, uh... Yeah, hopefully it launches soon. I think the-- in the JavaScript Council, the demos were super incredible that I saw, like, a few months ago, so I'm, I'll be intrigued to see when it lands.

**Swyx** [21:31]
Yeah. To, to me, like, this is the promise of sort of the local LLMs. It should be owned by the operating system, which is Apple, and we know the issues with Apple, or the browser, which is Google. And so it's one of the two really.

Like, y- you know, you probably shouldn't have to download a separate model for every single app that you own. Like, every model should just share the same model. Sorry, every app should share the same model. Yeah, I mean, uh, one, one thing I, I'm curious about, um, you know, you're talking about, like, the multimodality.

Have you tried Astra? Like, what, you know, what's-

### Astra Memory

**Logan Kilpatrick** [21:58]
Yeah, yeah.

**Swyx** [21:59]
Yeah.

**Logan Kilpatrick** [21:59]
I r- right after I/O, when Astra got announced, I got to spend a bunch of time just, like, testing out a bunch of use cases. Um, and, and to me, what's, what's actually astonishing is the multimodal live API, like, can do pretty much most of the Astra experience out of the box.

I think there's a couple of-

**Swyx** [22:15]
Yeah. Is that the, basically the API for Astra, or is this, like, a different-

**Logan Kilpatrick** [22:19]
It's, it's technically different. I think Astra was really intended as, like, a research experiment of, like, how do we push, like, truly the frontier of, of what's possible in this, in this user/product experience, um, and, like, didn't take into account, like, any of the thus far, like, practical realities of, like, could we ship this thing to millions of people and, like, have them actually deploy it in their product?

So different sort of constraint situation, but I think most probably, like, 80% of that experience you could get out of the box with the multimodal live API. The couple of things that I think we're still working on, and actually we'll hopefully bring those to the API as well, is, like, one of the big ones is memory.

Like, the cool thing about Astra is you can sort of show it something or talk to it and then as you go from session to session, it will actually remember the different contexts between the two sessions. And there's, like, a bunch of, you know, a bunch of very real constraints to make that happen, and I think a memory service is, is something that I've been talking to the team a bunch about as, as far as, like, building that layer for developers and doing it in a really elegant way that's not just, like, rag with embeddings.

**Swyx** [23:20]
How? Like, I, I've... Okay, you, you have some secret sauce or, or you wanna share?

**Logan Kilpatrick** [23:25]
Yeah, no. I, I, I don't wanna give the secret sauce because, like, we're-- we, we have to actually build this and, and make it work and scale.

**Swyx** [23:31]
I dig it.

**Logan Kilpatrick** [23:31]
But, like, I think the, I think the solution has to be something different, and part of the challenge is the scale, but also, like, the speed at which you get some of this context back. I don't think that, like, embeddings are going to be able to scale to, to...

I think they, they work well for some of this as, like, kind of the MVP version of the experience, but I think you're gonna need a different experience, and it's probably something like really smart caching with, like, a bunch of, you know, sparkles mixed in there that, uh, is, like, some of the infrastructure sweet sauce.

Like, it really feels like this core infrastructure problem that is not solved by the, by the embedding example.

**Swyx** [24:08]
Hmm. Okay. The one thing I, I'm fine with that, with caching. I mean, it basically means you're using full attention on your memory, which, which is always gonna be better than RAG and retrieval and, and all that. The one problem with that is you can't delete, um, things, and sometimes it's very helpful to delete things.

**Logan Kilpatrick** [24:26]
I, I think doing it in a way that lets you delete things is the secret sauce.

**Swyx** [24:30]
Okay.

**Logan Kilpatrick** [24:31]
That's the-- uh, that's the hope is that you, you can sort of have that level of control.

**Swyx** [24:36]
That's super exciting. Yeah, I'm looking forward to that. Yeah, I, I mean, I, you know, I hope that you publish a paper someday about it, you know, but launch it first and we'll see. Cool. Anything else? Uh, call to action, uh, shameless plugs, anything coming up?

### Search Grounding

**Logan Kilpatrick** [24:47]
Yeah, I think the only other thing is we're, we're also pushing on a bunch of these, like, search-powered use cases with Gemini. It's one of the things that I think very naturally makes a lot of sense for Google, and we shipped Search Grounding a few months ago, and then with 2.0 we updated a second generation of this called Search as a Tool.

Um, and I, I think, you know, we're sort of seeing the early versions of this with both deep researches, the Google one and the OpenAI one. But I think there's a huge amount of value to be created with that product experience and enabling developers actually to be the ones to sort of keep pushing in that frontier and, and building the right versions of that product experience I'm excited about and is, is something that we're, that we're pushing on.

And you can build, like, a lot of that experience today with Search as a Tool, but I think there's, there's some more work on our side to, to make it a lot less friction-full to, to bring that to life.

**Swyx** [25:35]
Yeah. The-- I, I'm strongly-- I like, uh... I think this is a category called, called, uh, online LLMs, like Perplexity kind of plays in this category. Other people have other search tools, but obviously Google is the best at search.

There's, there's no, no measurable question there. Cool. Alessio, anything else?

**Alessio** [25:50]
No, this was great, man. You were episode one. I think this is episode one-eighteen or something, so- ... uh, great, great to have you back.

### Wrap-up

**Swyx** [25:57]
And you have your own podcast. Are you still, are you still running that? You know, can we, can we plug it?

**Logan Kilpatrick** [26:01]
I have two, I have two podcasts. That's the downside-

**Swyx** [26:03]
What?

**Logan Kilpatrick** [26:03]
... is I have, I have a personal podcast and then I have a Google podcast, so-

**Alessio** [26:06]
He doubled us.

**Logan Kilpatrick** [26:07]
Okay. Okay.

**Alessio** [26:07]
You got two, we got one.

**Swyx** [26:09]
Plug, plug it, plug it, plug it.

**Logan Kilpatrick** [26:10]
I'm sorry.

**Swyx** [26:10]
What are they?

**Logan Kilpatrick** [26:11]
Yeah, so the personal podcast, Around the Prompt. Uh, me and my co-host, Nolan, uh, sit down and, and really we're trying to frame this as like very lightweight, unfiltered conversations. Uh, we j- we actually have an episode with, uh, Chamath Palihapitiya going out, I think today, if I put it out on time.

**Swyx** [26:28]
Wow. Nice.

**Logan Kilpatrick** [26:29]
CEO of, of GitHub, Thomas Dworaki tomorrow. Uh, so like lots of great folks and just chatting them with about AI and having very casual conversations. And then at Google, I run the Google AI Release Notes podcast, and this is really about sort of how do we peel back the layers of the stuff that's happening inside of Google.

And I've had amazing folks like Emmanuel Taropa talking about long context and Flash AB, Tulsee Doshi, who's the product director for Gemini, and then Jack Rae, actually the co-lead for reasoning models, and I just recorded an episode, so hopefully that will be out in the next few days at the time of this recording.

**Swyx** [27:01]
Dang. Okay, well, we'll have to catch up on those. Uh, I also know that DeepMind started a podcast with Hannah Fry. Uh, we have a small little Hannah Fry fan club 'cause her voice is so great.

**Logan Kilpatrick** [27:10]
She's great. The challenge for me is she makes my stuff look so bad that I'm like, "I'm glad Hannah's doing this because, like, she's the real deal. She's a professional."

**Swyx** [27:17]
You, you just need a British accent.

**Logan Kilpatrick** [27:20]
Truly. I'll-- There's a lot of people at DeepMind with British accents, so I, I feel like- ... I, I need to differentiate in a, in a different way other than a British accent.

**Swyx** [27:28]
Yeah. Awesome. Well, thanks for coming back. Uh, thanks for all the awesome work you're doing. Honestly, I'd just love to have you back for, for Worlds Fair for, for, uh, for this year, but, uh, we'll talk, we'll talk about that offline.

**Logan Kilpatrick** [27:36]
Yeah, yeah. We'll talk offline. I'm excited. I'll, I'll be there. Hopefully, the scheduling will not be, uh, will not be horrible like it was for me, for me last year- ... because of... It was-- I for- we were launching like five things or something like that-

**Swyx** [27:46]
Yeah

**Logan Kilpatrick** [27:46]
... at the same time. It was the-

**Swyx** [27:48]
Yeah

**Logan Kilpatrick** [27:48]
... the GA release or something. I don't even remember what it was, but.

**Swyx** [27:51]
Yeah, yeah. But also, I mean, Google stepped-- came through with like, uh, some other presenters and, uh, you, you also had Kathleen from the, the Gemma team, and I think people are very excited about open models still. Yeah.

There's lots of stuff from-- coming from Google.

**Logan Kilpatrick** [28:03]
Love it. Wonderful to see you both.

**Swyx** [28:05]
All right.

**Alessio** [28:05]
Bye, Logan.

**Swyx** [28:06]
Bye.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
