# Why is everyone cloning Deep Research?

Latent Space · 2025-02-18

<https://addtry.com/c4171a90-998f-4407-92f6-240496444633>

In this episode, hosts Alessio and Swyx interview Aarush Selvan and Mukund Sridhar, the lead PM and tech lead for Gemini Deep Research, about how Google created the deep research agent category and the technical decisions behind it. They finetuned a special version of Gemini 1.5 Pro for the feature, which autonomously browses the web for about five minutes to produce a research report. The team built an asynchronous platform to handle long-running tasks and developed an ontology of research behaviors—from broad shallow exploration to deep focus—to evaluate performance. They found that users value the visible browsing process even when it takes minutes, contrary to typical latency norms. The discussion covers iterative planning, trade-offs between model memory and web grounding, and future directions including thinking models, multimodality, and personalization.

## Questions this episode answers

### Why does Gemini Deep Research use a specially fine-tuned model rather than the standard Gemini API?

Aarush and Mukund explain that while the underlying model is similar, they performed extensive post-training to handle multi-step planning, iterative web browsing, and research plan generation reliably. They needed the model to propose and refine plans, decide when to search deeper, and produce coherent reports, requiring careful fine-tuning without losing pre-trained knowledge.

[9:15](https://addtry.com/c4171a90-998f-4407-92f6-240496444633?t=555000)

### How does the Gemini Deep Research team evaluate its agent's performance?

According to Mukund, they designed an ontology of research behaviour patterns, from broad/shallow to specific/deep, to create a test set. They use automatic metrics like the number of plan steps or browsing turn counts to spot distribution shifts, and they conduct human evaluations on comprehensiveness, completeness, and groundedness to ensure quality across diverse queries.

[22:41](https://addtry.com/c4171a90-998f-4407-92f6-240496444633?t=1361000)

### What unexpected user reaction did the team see regarding Deep Research's long processing time?

Aarush noted it was counterintuitive: they assumed users would dislike waiting 5–10 minutes, but many perceived the longer runtime as evidence of thorough work. Some even asked if the system deliberately delayed results. This reversed Google's typical latency-optimization mindset, leading the team to initially ship a 5-minute version with a 10-minute hard cap.

[31:30](https://addtry.com/c4171a90-998f-4407-92f6-240496444633?t=1890000)

## Key moments

- **[0:00] Intro**
- **[1:08] Deep Research 101**
  - [1:16] Q: What is Gemini Deep Research? — Aarush Selvan describes it as an AI assistant that browses the web for ~5 minutes and generates a research report.
  - [3:59] Deep Research's editable research plan lets users steer the AI before it spends 5–10 minutes searching, says Aarush Selvan.
  - [4:50] Mukund Sridhar: Deep Research shows a research plan to inform the user about the topic and elicit steering, rather than asking questions directly.
  - [5:17] Q: What is the top tip for using Deep Research? — Aarush Selvan: edit the research plan.
  - [6:14] Deep Research shows the websites it reads in real time as it browses, providing transparency during the search process.
  - [7:00] Mukund Sridhar: Deep Research iteratively searches and reads webpages, using two tools — general search and deep dive into specific pages.
- **[16:49] Under the Hood**
  - [16:49] Q: Does Deep Research use HTML-to-markdown conversion? — Mukund Sridhar says yes, but sometimes raw HTML is used for embedded snippets.
  - [18:39] Q: Does Deep Research ever exceed Gemini's 2M token context? — Mukund Sridhar: when it does, they use a RAG setup to retrieve older research material.
  - [20:03] Mukund Sridhar's rule of thumb for RAG vs long context: if the query involves multiple facets making cosine similarity unreliable, avoid RAG.
  - [21:44] Mukund Sridhar: if a follow-up is a natural progression of the same project, continue in the thread; otherwise, start new research.
  - [22:43] Q: How does the Deep Research team evaluate its model? — Mukund Sridhar: they use auto metrics on plan length, iteration steps, plus human evals on comprehensiveness and groundedness.
- **[23:41] Evals**
  - [24:49] Aarush Selvan: they built an ontology of research patterns — broad/shallow to deep/specific — to create a diverse eval set for Deep Research.
  - [26:25] Deep Research's eval set spans from 'options exploration' like shopping for summer camps to deep dives on specific topics, says Aarush Selvan.
  - [27:13] Q: What is the maximum conversation length Deep Research allows? — Aarush Selvan says there's no hard limit, but most users don't go very deep yet.
  - [28:42] Mukund Sridhar: dogfooding Deep Research revealed unexpected use cases, like wedding planning, that they hadn't initially designed for.
- **[29:12] Market & Latency**
  - [29:12] Aarush Selvan: shopping in Deep Research is more about options exploration than visual product comparison, complementing Google Shopping.
  - [31:07] Swyx notes a perverse incentive: users perceive longer research times as better, but many extra websites are irrelevant.
  - [32:08] Aarush Selvan: Jason Calacanis asked if they fake the 5-minute wait; users actually value seeing the work, contrary to typical latency optimization.
  - [33:20] Mukund Sridhar: the extra time in Deep Research is spent either exploring more topics for completeness or verifying uncertain information.
  - [34:52] Swyx suggests a toggle for verification vs search amount; Aarush Selvan worries users would always max it out, like a 'max power' button.
- **[36:21] Feedback & Learnings**
  - [36:21] NotebookLM team inspired Deep Research by focusing on a simple problem and being opinionated about the solution, says Aarush Selvan.
  - [37:02] Swyx suggests Deep Research should allow live chat during browsing, letting users edit the plan mid-execution like Devin does.
- **[40:31] Frontier Tech**
  - [40:31] Q: What has the Deep Research team learned from OpenAI's clone? — Aarush Selvan: they took bets on transparency, publisher-forward UX, and side-by-side artifacts, which others are adopting.
  - [42:26] Mukund Sridhar: Deep Research runs on 1.5 Pro; the new thinking models like Gemini 2.0 unlock better analytical reasoning but require careful balancing.
  - [44:02] Mukund Sridhar: thinking models know more and reason better, but they may over-rely on internal knowledge, making grounding in sources crucial.
  - [45:42] Mukund Sridhar: the hardest problem in Deep Research was teaching the model to iteratively plan across diverse topics without overfitting.
  - [46:38] Deep Research uses a custom async platform for durable execution of multi-minute jobs with state retention and retries, says Mukund Sridhar.
- **[49:02] Benchmarks & Discovery**
  - [49:02] Swyx says OpenAI beats Google on marketing for Deep Research because Google doesn't publish benchmarks; Aarush Selvan says they avoid benchmarks that don't reflect product experience.
  - [51:31] Q: Can Deep Research discover new ideas, not just summarize? — Mukund Sridhar: thinking models can generate second-order insights, but verifying hypotheses in domains like chemistry lacks synthetic environments.
- **[54:12] Agent Building**
  - [54:12] Aarush Selvan predicts Deep Research will personalize reports to user's knowledge level and add multimodal outputs like charts and maps.
  - [55:29] Aarush Selvan: a key next step for Deep Research is accessing private documents and subscriptions to go beyond the open web.
- **[57:35] Conference Talk**
  - [59:07] Swyx: Deep Research agents need to access private, auth-walled web content, requiring better credential delegation protocols.

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **Aarush Selvan** (guest)
- **Mukund Sridhar** (guest)

## Topics

Agent Infrastructure, Agent Platforms, Reasoning

## Mentioned

Google (company), Deep Research (product), Devin (product), Gemini (product), Gemini Extensions (product), Gemma (product), Google Assistant (product), Google Docs (product), Google Search (product), Google Shopping (product), NotebookLM (product), OpenAI Deep Research (product), o3 (product)

## Transcript

### Intro

**Alessio** [0:05]
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO at Decibel Partners, and I'm joined by my co-host Swyx, founder of Small AI.

**Swyx** [0:14]
Hey, and today we're very honored to have in our studio Aarush and Mukund from the Deep Research team, the OG Deep Research team. Welcome.

**Aarush Selvan** [0:21]
Hey, thanks.

**Mukund Sridhar** [0:22]
Thanks. Thanks for having us.

**Swyx** [0:22]
Yeah, thanks for making the trip up. I was, uh, fortunate enough to be one of the early beta testers of Deep Research when it came out and, uh, you know, I would say I was very keen on... Like, even I think even at the end of last year, people were already saying, like, it was one of the most exciting agents that was, you know, coming out of Google.

You know that previously we had on, uh, Raiza and Usama from the NotebookLM team, and I think, like, this is, this is, like, a increasing trend that, like, Gemini and Google are shipping interesting user-facing products that use AI.

So, uh, congrats on your success so far.

**Aarush Selvan** [0:56]
Yeah. It's been great. Thanks so much for having us here and yeah, excited.

**Swyx** [0:59]
Yeah, thanks for making the trip up, and I'm also excited for your talk that is, uh, happening next week. Obviously, we have to talk about what, what it, what exactly it is. But, you know, I'll ask you, I'll ask you towards the end.

But, uh, so basically okay, um, you know, we have the screen up. Maybe we just start at a high level for people who don't n- don't yet know. Like, what is Deep Research?

### Deep Research 101

**Aarush Selvan** [1:16]
Sure. So Deep Research is a feature where Gemini can act as your personal research assistant to help you learn about any topic, uh, that you want more deeply. It's really helpful for those queries where you wanna go from zero to 50 really fast on a new thing.

And the way it works is it, you know, it takes your query, browses the web for about five minutes, and then outputs a research report for you to review and, and ask follow-up questions.

**Mukund Sridhar** [1:43]
This is one of the first times, you know, that, uh, something takes about five, six minutes trying to-

**Aarush Selvan** [1:48]
Yeah

**Mukund Sridhar** [1:48]
... perform your, uh, research. So there's a, a few challenges that brings. Like, you wanna make sure you're spending that time and the compute doing what the user wants. Uh, so there's some ways of the UX design that we can talk about as we go through an example.

And then there's, um, also challenges in the browsers. Uh, the, the web is super fragmented, and, uh, being able to plan iteratively and as, as you parse through this noisy information is a challenge by itself.

**Swyx** [2:11]
Yeah. This is, like, the first time sort of Google automating yourself as searching. Like, you're, you know, you're supposed to be the, the experts at Search, but now you're, like, meta searching and, like, de- determining the search strategy.

**Mukund Sridhar** [2:23]
Yeah. I think, uh, at least we see it as two different use cases.

**Swyx** [2:27]
Okay.

**Mukund Sridhar** [2:27]
There are, uh, things that you know, you know exactly what you're looking for, and, uh, their search is still probably, you know, a very, you know, probably one of the best places to go. I think where Deep Research really shines is there are, like, multiple facets to your question, and you spend, like, a weekend, uh, you know, just opening, like, 50, 60 tabs, and many times I just give up.

And, uh, we wanted to solve that problem and, and give a great starting point, uh, for those kind of journeys.

**Alessio** [2:54]
Um, do we wanna start a query so that-

**Aarush Selvan** [2:56]
Yeah

**Alessio** [2:56]
... it runs in the meantime-

**Aarush Selvan** [2:57]
Let's do one. Let's do one

**Alessio** [2:57]
... and then we can chat over it.

**Aarush Selvan** [2:59]
Um, okay. Here's one query that, that we like. We love to test, like, super niche random things. Like, things where there's, like, no Wikipedia page already about this topic or something like that, right? Because that's where you'll see the most lift from, from a feature like this.

So for this one, I've come, I've come, come up with this query. Th- uh, this is actually Mukund's query that he's he, he loves to test, is, "Help me understand how milk and meat regulations differ between the US and Europe."

What's nice is the first step is actually where it puts together a research plan that you can review, and so this is sort of its guide for how it's gonna go about and carry out the research, right? And so this was, like, a pretty decently well-specified query, but, like, let's say you came to Gemini and were like, "Tell me about batteries," right?

That query, you could mean so many different things. You might wanna know about the, like, latest innovations in battery tech. You might wanna know about, like, a specific type of battery chemistry. And if we're gonna spend, like, five to even 10 minutes researching something, we wanna, one, understand what exactly are you trying to accomplish here, and two, give you an opportunity, like, to steer where the research goes, right?

Because, like, if you had an intern and you asked them this question, the first thing they'd do is ask you, like, a bunch of follow-up questions and be like, "Okay, so, like, help me figure out exactly what you want me to do."

And so the way we approached it is, um, we thought, like, why don't we just have the model produce its, its first stab at the, at the research query, at, at how it would break this down, and then invite the user to come and kind of engage with how they would wanna steer this.

**Mukund Sridhar** [4:31]
Yeah. And many times when you try to use a product like this, you often don't know what questions to look for or the things to look for. So, uh, we kind of made this decision very deliberately that instead of asking the users just follow-up questions directly, we kind of lay out, "Hey, this is what I would do.

Like, these are the different facets." For example, here it could be like what additives are allowed, uh, and how that differs, or labeling, uh, restrictions and so on in products. The aim of this is to kind of tell the user about the topic a little bit more and also get steer.

At the same time, we elicit for, like, uh, you know, follow-up questions and so on. So we-

**Aarush Selvan** [5:08]
Yeah

**Mukund Sridhar** [5:09]
... kind of did that in a joint function.

**Aarush Selvan** [5:10]
It's kind of like editable chain of thought.

**Mukund Sridhar** [5:11]
Right. Exactly.

**Aarush Selvan** [5:12]
Yeah.

**Mukund Sridhar** [5:12]
Exactly.

**Swyx** [5:13]
Yeah, I think that, you know, we were talking to you about, like, your top tips for using Deep Research.

**Aarush Selvan** [5:17]
Yeah.

**Swyx** [5:17]
And your number one tip is to edit the plan.

**Aarush Selvan** [5:18]
Just edit it, right? So, like, we actually... You can actually edit conversationally. We put in a button here just to, like, draw users' attention to the fact that you can edit this. Um-

**Swyx** [5:28]
Oh, actually you don't need to click the button.

**Aarush Selvan** [5:29]
You don't even-

**Mukund Sridhar** [5:29]
No, you can just check on that

**Aarush Selvan** [5:29]
... you can just click on this chat. Yeah. Actually, like, in early rounds of testing we saw no one was editing, and so we were just like, "If we just put a button here, maybe people will, will, like, engage more."

**Swyx** [5:37]
I, I confess, I just hit Start a lot.

**Aarush Selvan** [5:39]
I think-

**Swyx** [5:40]
You know

**Aarush Selvan** [5:40]
... like, we see that too. Like-

**Swyx** [5:41]
Yeah

**Aarush Selvan** [5:41]
... most people hit Start. Um-

**Swyx** [5:43]
Like, it's like the I'm feeling lucky.

**Aarush Selvan** [5:44]
Yeah.

**Mukund Sridhar** [5:46]
Yeah.

**Aarush Selvan** [5:46]
All right. So, like, I, I can just add a, add a step here, and what you'll see is it should, like, refine the plan-

**Swyx** [5:52]
Okay

**Aarush Selvan** [5:52]
... and, and show you a new thing to, uh, propose. Here we go. So it's added step seven, find information on milk and meat labeling requirements in the US and the EU. Or you can just go ahead and hit Start.

I think it's still, like, a nice transparency mechanism, even if users don't want to engage. Like, you still kind of know, okay, here's at least an understanding of why I'm getting the report I'm gonna get, um, which is kind of nice.

And then while it browses the web, and Mukund, you should maybe explain kind of how it-

**Swyx** [6:18]
Yeah

**Aarush Selvan** [6:18]
... how it browses. We show kind of the w- the websites it's reading in real time.

**Swyx** [6:23]
Yeah. I'll preface this with I haven't-- I forgot to explain the roles. You're a PM-

**Aarush Selvan** [6:28]
Yes

**Swyx** [6:28]
... and you're a tech lead.

**Mukund Sridhar** [6:28]
Yes.

**Swyx** [6:29]
Okay.

**Mukund Sridhar** [6:30]
Yeah.

**Swyx** [6:30]
Just for people who, who don't know.

**Aarush Selvan** [6:32]
Oh, okay. We maybe should have started with that-

**Swyx** [6:35]
Yeah, yeah

**Aarush Selvan** [6:36]
... I suppose.

**Mukund Sridhar** [6:36]
Yeah, yeah.

**Aarush Selvan** [6:36]
Yeah.

**Mukund Sridhar** [6:36]
We do each other's work sometimes as well, but, uh -

**Swyx** [6:39]
It's how it is

**Mukund Sridhar** [6:39]
... more or less that's the boundary. Yeah, yeah.

**Swyx** [6:41]
Yeah.

**Mukund Sridhar** [6:42]
Um, yeah, so, so what's happening behind the scenes actually is we kind of give this research plan that is a contract and that, uh, you know, has been accepted. But then if you look at the plan, there are things that are obviously parallelizable, so the model figures out which of the substeps that it can start exploring in parallel, and then it primarily uses, like, two tools.

It has the ability to do perform searches, and it has abilities to go deeper within, you know, a particular webpage of interest, right? And oftentimes it'll start exploring things in parallel, but that's not sufficient. Many times it, it has to reason based on information found.

So in this case, it c- one of the searches could have led the EU Commission has these additives banned. It wants to go and check if the FDA does the same thing, right? So, uh, this notion of being able to read outputs from the previous turn, uh, ground on that to decide what to do next, I think was, was key.

Otherwise, you have, like, incomplete information, and your report becomes, uh, a little bit of a, like, a high-level, uh, bullet points. So we wanted to go beyond that blueprint and actually figure out, you know, what are the key aspects here.

So yeah. So the-- this happens iteratively until the model thinks it's finished all its steps, and, uh, then we kind of enter this, uh, analysis mode, and here there can be inconsistencies across sources. You kind of come up with an outline for the report, start generating a draft.

The model tries to revise that by self-critiquing itself, uh, you know, to fi- to finalize the prompt, uh, finalize the report, and that's broadly what's happening behind the scenes.

**Swyx** [8:18]
What's the initial ranking of the websites? So, uh, when you first started it, there were 36. How do you decide where to start since it sounds like, you know, the initial websites kind of carry a lot of weight too because then they inform the following?

**Mukund Sridhar** [8:33]
Yes. So what happens in the initial turns, again, this is not like a-- it's not something we enforce. It's mostly the model making these choices. But typically, we see the model exploring all the different s- aspects in the, in the research plan that was presented.

So we kind of get, like, a breadth-first idea of what are the different topics to explore. And in terms of which ones to double-click on, I think it really comes down to every time you search, the model gets some idea of what the page is, and then depending on what pieces of it, sometimes there's inconsistency, sometimes there's just, like, partial information.

Those are the ones it double-clicks on. And, uh, yeah, it can continually, like, iteratively search and, and, and browse until it feels like it's done.

**Swyx** [9:15]
Yeah. I'm trying to think about how I would code this. Um, a simple question would be like, do you think that we could do this with the Gemini API, or do you have some special access that we cannot replicate?

You know, like, is-- if I model this with a tool call of, like-

**Mukund Sridhar** [9:30]
Yeah

**Swyx** [9:30]
... search, double-click, whatever.

**Mukund Sridhar** [9:32]
Yeah. I don't think we have special accsh-access per se. It's pretty much the same model. We of course have our own, uh, post-training work that we do, and y'all can also, like, you know, you will, you will-- can finetune from the base model and so on.

Uh-

**Swyx** [9:45]
I don't know that we can do that.

**Mukund Sridhar** [9:46]
I don't, I don't know.

**Swyx** [9:46]
Let's turn the head finetuning.

**Aarush Selvan** [9:48]
Well, if you use our Gemma open source models-

**Swyx** [9:50]
Yeah, yeah, yeah

**Aarush Selvan** [9:50]
... uh, you could finetune.

**Swyx** [9:51]
Yeah, yeah.

**Mukund Sridhar** [9:52]
Yeah. So I, I don't think there's a special access-

**Swyx** [9:54]
Okay

**Mukund Sridhar** [9:54]
... per se, but the, a lot of the work for us is first defining these, "Oh, there needs to be a research plan," and, and how do you go about presenting that, and then, uh, a bunch of post-training to make sure, you know, it's able to do this consistently well and, uh, with, with high reliability and all.

**Swyx** [10:10]
Okay. So, so 1.5 Pro with Deep Research is a special edition of 1.5 Pro.

**Mukund Sridhar** [10:14]
Yes.

**Swyx** [10:14]
Right?

**Mukund Sridhar** [10:14]
So it's a post-training version of-

**Swyx** [10:16]
It's not pure 1.5 Pro

**Mukund Sridhar** [10:16]
... it's, it's, it's, it's a post-training.

**Swyx** [10:17]
This also explains why you haven't just-- you can't just toggle on 2.0 Flash and just... Yeah.

**Aarush Selvan** [10:22]
Right.

**Swyx** [10:22]
Yeah. But I mean, I, I assume you have the data and, you know, it's should be doable.

**Mukund Sridhar** [10:27]
Yep.

**Swyx** [10:27]
There's still this, like, question of ranking.

**Mukund Sridhar** [10:29]
Yeah.

**Swyx** [10:30]
Right? And like-- Oh, it looks like you're, you're already done.

**Aarush Selvan** [10:32]
Yeah. Yeah, yeah, we're done.

**Swyx** [10:33]
Okay. We can look at it.

**Aarush Selvan** [10:34]
Yeah. So let's see. It's put together this report, and what it's done is it's sort of broken-- started with, like, milk regulation, and then it looks like it goes into meat probably further down, and then sort of covering how the US approaches this problem of, like, how to regulate milk, comparing and then, you know, covering the EU, and then, yeah, like I said, like, going into the meat production.

And then it'll also... What's nice is it kind of reasons over, like, why are there differences.

**Swyx** [11:03]
Mm-hmm.

**Aarush Selvan** [11:03]
And I think what's really cool here is, like, it's, it's showing that there's, like, a difference in philosophy between how the US and the EU regulate food. So the EU, like, adopts a precautionary approach. So even if there's inconclusive scientific evidence about something, it's still gonna prefer to, like, ban it, whereas the US takes sort of the reactive approach, where it's, like, allowing things until they can be proven to be harmful, right?

So, like, this is kind of nice is that you, you also un-- like, get the second-order insights from what it's being put, what it's putting together. So yeah, it's, it's kinda nice. It takes a few minutes to read and, like, understand everything, which makes for, like, a quiet period during a podcast, I suppose.

**Swyx** [11:43]
Oh, this is it.

**Aarush Selvan** [11:44]
But, but yeah, this is, this is kind of how it, how it looks right now.

**Swyx** [11:47]
Yeah. And then from here, you can kinda keep the usual chat and iterate thing. So this is more if you were to, like, uh, you know, compare it to other platforms, it's kinda like a, um, Anthropic artifact or like a ChatGPT Canvas where, like, you have the document on one side and, like, the chat on the other, and you're working on it.

**Aarush Selvan** [12:05]
Yeah. Yeah. This is something we thought a bit about and, um- One of the things we feel is like your learning journey shouldn't just stop after the first report. And so actually, what you probably wanna do is while reading, be able to ask follow-up questions without having to scroll back and forth.

And there's, like, broadly a few different kinds of follow-up questions. One type is, like, maybe there's like a factoid that you want that isn't in here, but it's probably been already captured as part of the web browsing that it did, right?

So we actually keep everything in context, like all the sites that it's read remain in context. So if there's a m- piece of missing information, it can just fetch that. Then another kind is like, okay, this is nice, but you actually wanna kick off more deep research.

You're like, "I also want to compare the EU and Asia," let's say, "in how they regulate milk and meat." For that, you'd actually w- want the model to be like, "Okay, this is sufficiently different that I want to go do more deep research to answer this question.

I won't find this information in what I've already browsed." And the third is actually maybe you just wanna, like, change the report. Like, maybe you wanna, like, condense it, remove sections, add sections, and actually, like, iterate on the report that you got.

So we broadly are basically a lot... have... try and teach the model to be able to do all three. And the kind of side-by-side format allows sort of for the user to do that more easily.

**Swyx** [13:25]
Yeah. So as a PM, there's a Open in Docs button-

**Aarush Selvan** [13:28]
Yeah

**Swyx** [13:28]
... there, right? How do you think about what you're supposed to build in here versus kinda it sounds like the condensing and things should be a Google Docs.

**Aarush Selvan** [13:36]
Yeah. Bard Extensions is different. It's just like an amazing editor. Like, sometimes you just wanna direct edit things, and now Google Docs also has Gemini in the side panel. So the more we can kind of help this be part of your workflow throughout the rest of the Google ecosystem, the better, right?

Like... And one thing that we've noticed is people really like that button and really like exporting it. It's also a nice way to just save it permanently. And when you do export all the citations, and in fact, I can just run it now, carry over, which is also really nice.

Gemini Extensions is a different feature. So that is really around Gemini being able to fetch content from other Google services in order to inform the answer. So that was actually the first feature that we both worked on on the team, is actually building, um, extensions in Gemini.

And so I think right now we have a bunch of different Google apps, as well as I think Spotify and, and a couple. I don't know if we have... and Samsung, uh, apps as well.

**Swyx** [14:29]
Who wants Spotify? I, I have this whole thing about like who wants Spotify-

**Aarush Selvan** [14:35]
I love Spotify

**Swyx** [14:35]
... for searches.

**Aarush Selvan** [14:35]
I think-

**Swyx** [14:35]
No, no, but who, who wants that in their deep research?

**Aarush Selvan** [14:37]
Sure. Sure. Sure.

**Swyx** [14:37]
Like what the

**Aarush Selvan** [14:39]
In deep research, I think less, but, like, the, the interesting thing is, like, we built extensions and we didn't-- we weren't really sure how people were gonna use it. And a ton of people are doing really creative things with them, and a ton of people are just doing things that they loved on the Google Assistant.

And Spotify is like a huge... Like, playing music on the go was like a huge va- a huge value use case.

**Swyx** [14:58]
Oh, it controls Spotify? It plays-

**Aarush Selvan** [15:00]
Yeah. So you can-

**Swyx** [15:01]
It's not, it's not deep research. For deep research-

**Aarush Selvan** [15:03]
Yeah, yeah

**Swyx** [15:04]
... you clearly use-

**Aarush Selvan** [15:04]
Yeah. Yeah, yeah.

**Swyx** [15:05]
This is search.

**Aarush Selvan** [15:06]
But otherwise, yeah. Like you can-

**Swyx** [15:07]
Okay. Okay

**Aarush Selvan** [15:07]
... you can have Gemini go.

**Swyx** [15:08]
Yeah, you have YouTube, Maps, and Search for Flash Thinking Experimental with Apps, the, the newest-

**Aarush Selvan** [15:13]
Yeah

**Swyx** [15:14]
... longest model name that has been launched.

**Aarush Selvan** [15:16]
Yeah.

**Swyx** [15:16]
Um, but like, yeah, I think Gmail is obvious one.

**Aarush Selvan** [15:19]
Yeah.

**Swyx** [15:19]
Calendar is obvious one.

**Aarush Selvan** [15:20]
Exactly.

**Swyx** [15:20]
You know, tho- tho- those I want.

**Aarush Selvan** [15:22]
Yeah.

**Swyx** [15:22]
And Spotify.

**Aarush Selvan** [15:24]
Yeah. Fair enough.

**Swyx** [15:25]
Yeah. And obviously feel free to dive in on your other work. I know you're, you're not just doing deep research, right? But, you know, we're just kind of focusing on, on, uh, deep research here. I actually have asked for modifications after this first run-

**Aarush Selvan** [15:37]
Mm-hmm.

**Mukund Sridhar** [15:38]
Yeah

**Swyx** [15:38]
... where I was like, "Oh, you, you stopped. Like, I actually want you to keep going. Like, what about these other things?"

**Aarush Selvan** [15:43]
Oh, yeah.

**Swyx** [15:43]
And then continue to modify it. So it really felt like a little bit of a co-pilot type experience, but more like an agent that would re- research. I thought it was pretty cool.

**Mukund Sridhar** [15:52]
Yeah. One of the challenges is currently we kind of let the model decide based on your query, like amongst the three categories. So some... There, there is, there is a boundary there. Like, some of these things, depending on how deep you wanna go, you might just want a quick answer versus, like, kick off another deep research.

And even from a UX perspective, I think the, the panel allows for this notion of, you know, not every follow-up is gonna take you like, uh-

**Aarush Selvan** [16:16]
Right

**Mukund Sridhar** [16:17]
... five minutes.

**Aarush Selvan** [16:17]
Right now it doesn't do any fo- Does it do follow-up search? It always does?

**Mukund Sridhar** [16:21]
It depends on your question. Since we have the liberty of like really long context, uh, models, we actually hold all these, uh, all the research material across dance. So if it's able to find the answer in, uh, things that it's found, you're gonna get a faster reply.

**Aarush Selvan** [16:36]
Yeah.

**Mukund Sridhar** [16:37]
Otherwise, it's just gonna go back to planning and-

**Aarush Selvan** [16:39]
Yeah, yeah. A, a bit of a follow-up on, on the... Since you brought up context, I had two questions. One, do you have a HTML to markdown transform step, or do you just consume raw HTML? You... There's no way-

### Under the Hood

**Mukund Sridhar** [16:49]
Yeah

**Aarush Selvan** [16:49]
... to consume raw HTML, right?

**Mukund Sridhar** [16:51]
We have both versions, right?

**Aarush Selvan** [16:52]
Oh.

**Mukund Sridhar** [16:53]
Uh, so there is... The models are getting... Like, every generation of models are getting much better at native, native understanding of these representations. I think the, the, the markdown step definitely helps in, in terms of, you know... There's a lot of noise, like as you can imagine with, with the pure HTML.

**Aarush Selvan** [17:08]
JavaScript.

**Mukund Sridhar** [17:09]
Right.

**Aarush Selvan** [17:09]
You deal with CSS.

**Mukund Sridhar** [17:09]
Exactly. Exactly. Uh, so yeah, when it makes sense to do it, there's... We don't artificially try to make it hard for the model . But sometimes it depends on the kind of access, uh, of, uh, what we get as well.

Like, uh, for example, if there's an embedded snippet that's HTML, we want the model to, you know, be able to work on that as well.

**Aarush Selvan** [17:28]
Yeah. And no Vision yet, but...

**Mukund Sridhar** [17:30]
Currently no Vision.

**Aarush Selvan** [17:31]
Yeah, yeah. It's-- The reason I ask all these things 'cause I've done the same.

**Mukund Sridhar** [17:34]
Got it.

**Aarush Selvan** [17:34]
Like, I haven't done Vision.

**Mukund Sridhar** [17:36]
Yeah. So the, the, the tricky thing about Vision is I, I think the models are getting significantly better, es- especially if you look at the last six months, natively being able to do like VQA stuff and so on.

But the challenge is the trade-off between having to, you know, actually render it and, and so on, the, the gap, the trade-off between the added latency versus the value add you get.

**Aarush Selvan** [17:58]
You have a latency budget of minutes.

**Mukund Sridhar** [18:00]
Yeah, yeah, yeah. Yeah, yeah, yeah.

**Aarush Selvan** [18:01]
Minutes.

**Mukund Sridhar** [18:02]
It's, it's true. In my opinion, uh, the places you'll see a real difference is like the- Like, I don't know, a small part of the tail, um, especially in, like, this kind of an open domain setting if you just look at what people ask.

**Swyx** [18:13]
Yeah.

**Mukund Sridhar** [18:14]
Uh, there's definitely some use cases where it makes a lot of sense to do it, but I still feel it's not in the, the, not in the head cases. Uh, and we'll, we'll-

**Swyx** [18:22]
I mean-

**Mukund Sridhar** [18:22]
We'll do it when we get there, I guess.

**Swyx** [18:23]
The classic is, like, it- it's a JPEG that has some important information-

**Mukund Sridhar** [18:27]
Right

**Swyx** [18:27]
... and you can't, you can't touch it. Yeah.

**Mukund Sridhar** [18:29]
Yeah.

**Swyx** [18:29]
Okay. And then the, the, the other technical follow-up was just, uh, you have one million to two million token context. Has it ever exceeded two million? Uh, and what do you do there?

**Mukund Sridhar** [18:39]
Yeah. So we had this challenge, uh, sometime last year where we said, uh, when we started, like, wiring up this multi-turn where we said, "Hey, let's see how long somebody in the team can take, uh, DR," you know?

**Swyx** [18:51]
Yeah. What's the most challenging question you can ask that-

**Mukund Sridhar** [18:53]
Yeah

**Swyx** [18:53]
... takes the longest?

**Mukund Sridhar** [18:54]
Yeah, no, we, we-

**Swyx** [18:54]
Second game. Keep doing it. Like, keep-

**Mukund Sridhar** [18:55]
We also keep asking follow-ups. Like, for, for example, here you could say, "Hey-

**Swyx** [18:58]
Oh, I see, I see

**Mukund Sridhar** [18:59]
... I also wanna compare it with, like, how it's done."

**Swyx** [19:00]
Okay. So you're, you're guaranteed to bust it. Yeah.

**Mukund Sridhar** [19:02]
Yeah, yeah, yeah. We also have, uh, we have retrieval mechanisms if required. So, uh, we natively try to use the context, uh, as much as it's available, uh, beyond which, uh, you know, we have, like, a RAG set up to figure out the-

**Swyx** [19:16]
Okay. This is all in-house, in-house tech?

**Mukund Sridhar** [19:18]
Yes.

**Swyx** [19:19]
Okay.

**Mukund Sridhar** [19:19]
Yes.

**Swyx** [19:20]
What are some of the differences between putting things in context versus RAG? And when I was in Singapore, I went to the Google Cloud-

**Mukund Sridhar** [19:26]
Long context versus RAG.

**Swyx** [19:27]
Well, when, when I was in Singapore, I went to the Google Cloud team, and they talked about Gemini plus grounding. Is Gemini plus search kinda like Gemini plus grounding? Or like how should people think about the different shades of, like, I'm doing retrieval and data versus I'm using Deep Research versus I'm using grounding?

Uh, sometimes the labels can be hard, too.

**Mukund Sridhar** [19:47]
Yeah. I can... Let me try to answer the first part of the question.

**Swyx** [19:51]
Yeah, yeah, yeah.

**Mukund Sridhar** [19:51]
Uh, the, the second part I'm not fully sure of, of the grounding offering, so, uh, uh, but I can at least, at least talk about the first part of the question. So I think, uh, you're asking, like, the difference between, like, being able to...

When do, when would you do RAG versus rely on the long context?

**Swyx** [20:07]
Well, I think we all, we all get that. I was more curious, like from a product perspective, when you decide to do RAG versus just... Like this, you didn't need to, you know. Uh, do you get better performance-

**Mukund Sridhar** [20:17]
Yeah

**Swyx** [20:17]
... just putting everything in context or?

**Mukund Sridhar** [20:19]
So the tricky thing for RAG, it really works well because a lot of these things are doing, like, cosine distance, like a dot product kind of a thing, and that kind of gets challenging when your query side has multiple different attributes.

Uh, the dot product doesn't really work as well. I would say, at least for me, that's, that's my guiding principle on, uh, when to avoid RAG. That's one. The second one is, I think every generation of these models are, uh, like the initial generations, even though they offered, like, long context, their performance as the context kept growing was you would see some kind of a decline.

But I think, uh, as the newer generation models came out, uh, they were really good even if you kept filling in the context and being able to piece out, uh, like these really fine grain information. So I think these two, at least for me, are like guiding principles on when to.

**Swyx** [21:11]
Just to add to that, I think, like, just like a simple rule of thumb that we use is, like, if it's the most recent set of research tasks where the user is likely to ask lots of follow-up questions, that should be in context.

But, like, as stuff gets 10 tasks ago, you know, it's fine if that stuff is in RAG because it's less likely that the user needs to do-- you need to do, like, very complex comparisons between what's currently being discussed and the stuff that you asked about, you know, 10 turns ago, right?

So that's just, like, a, a very, like, the rule of thumb-

**Mukund Sridhar** [21:44]
Yeah

**Swyx** [21:44]
... that we follow.

**Mukund Sridhar** [21:44]
And so from a user perspective, is it better to just start a new research instead of, like-

**Swyx** [21:49]
Ooh

**Mukund Sridhar** [21:49]
... extending the, the context?

**Swyx** [21:51]
Yeah.

**Mukund Sridhar** [21:51]
I think that's a good question. I think if it's a related topic, I think there's benefit to continue with the thread-

**Swyx** [21:57]
Mm-hmm

**Mukund Sridhar** [21:57]
... uh, because you could, the model, since it has this in memory, could figure out, "Oh, I've found this niche thing, uh, about, uh, I don't know, milk regulation in this case in the US. Let me check if your, you know, follow-up country or, or place also has something like that."

So these kind of things you might have not caught if you start a new thread.

**Swyx** [22:16]
Mm-hmm.

**Mukund Sridhar** [22:16]
So I think it really depends on, on the use case. If there's a natural progression, uh, and you feel like this is, like, part of one cohesive kind of a project, you should just continue using it. My follow-up turn's gonna be like, "Oh, I'm just gonna look for summer camps or something," then, yeah, I don't think it should make a difference, but we haven't really-

**Swyx** [22:33]
Mm-hmm

**Mukund Sridhar** [22:33]
... uh, you know, pushed that to see, uh, and, and, and tested that, that aspect of it for us. Most of our tests are, like, more natural trans-

**Swyx** [22:40]
Yeah

**Mukund Sridhar** [22:40]
... transitions. Yeah.

**Swyx** [22:41]
How do you eval Deep Research?

**Mukund Sridhar** [22:43]
Oh, boy. Uh, yeah, this is a hard one. I think the entropy of the output space is so high, like It's, um, like, people love auto raters, but it brings its own, own, own set of, uh, challenges. And so for us, we have some metrics that we can auto generate, right?

So for example, as we move, uh, when we do post-training and have multiple, uh, models, we kind of wanna make sure, uh, the distribution of, like, certain stats, like for example, how long is spent on planning, how many, how many iterative steps it does on, like, some dev set.

If you see large changes in distribution, that's, that's kind of like a early, uh, signal of, of something has changed. It could be for better or worse. Uh, so we have some metrics like that, that we can auto compute.

**Swyx** [23:28]
So every time you have a new version, you run it across a test suite of cases, and you see how long it takes?

**Mukund Sridhar** [23:33]
Yeah. So we have, like, a dev set, and we have, like, some kind of automatic metrics that we can detect in terms of, like, the behavior end-to-end. Like, for example, how long is the research plan? Do we like...

### Evals

**Mukund Sridhar** [23:44]
Does a new model produce really longer, uh, many more steps than-

**Swyx** [23:47]
Just number of characters. So it's-

**Mukund Sridhar** [23:49]
Uh, like number of steps in, in case of-

**Swyx** [23:51]
Oh, the research

**Mukund Sridhar** [23:51]
... research plan. In the plans, it could be like, like we spoke about how it iteratively plans based on, like, previous searches. How many steps does that go on an average over some, uh, dev set? Uh, so there are some things like this you can automate, but beyond that- There are auto raters, but we definitely do a lot of human evals, and there we have defined with product about certain things we care about and being super opinionated about is it comprehensive, is it complete, uh, like groundedness, and these kind of things.

So it's, it's a mix of these, uh, these two attributes. There's another challenge, but I'll, I'll, uh, if you, if you-

**Aarush Selvan** [24:27]
Is this where... Is the other challenge in that sometimes you just have to have your PM review examples?

**Mukund Sridhar** [24:33]
Yeah. Yeah. So, yeah. Yeah, exactly.

**Aarush Selvan** [24:35]
Yeah.

**Mukund Sridhar** [24:35]
Yeah. And, and, and for latency-

**Swyx** [24:37]
So you're the human, human rater.

**Aarush Selvan** [24:38]
I am the human rater. But broadly what we tried to do in-- is for the eval question is, like, we tried to think about, like, what are all the, the ways in which a person might use a feature like this?

And we came up with what we call an ontology of, of use cases.

**Swyx** [24:53]
Yes.

**Aarush Selvan** [24:53]
And really what we, what we tried to do is, like, stay away from, like, verticals like travel or shopping and things like that, but really try and go into, like, what is the underlying research behavior type that a person is doing.

So there's queries on one end that are just you're going very broad but shallow, right? Things like shopping queries are an example of that, where... Or like, I wanna find the perfect summer camp. My kids love soccer and tennis.

And really you just want to find as many different options and explore all the different options that are available and then synthesize, okay, what's the TLDR about each one? Kind of like those journeys where you open many, many Chrome tabs, but then, like, need to take notes somewhere of, of the stuff that's appealing.

On the other end of the spectrum, you know, you've got like a specific topic, and you just wanna go super deep on that and really, really understand that. And there's like all sorts of points in the middle, right?

Around like, okay, I have a few options, but I wanna compare them. Or like, yeah, I, I wanna go not super deep on a topic, but I wanna cover slightly, slightly more topics. And so we, we sort of developed this ontology of different research patterns, and then for each one came up with queries that would fall within that.

And then that's sort of the eval set by which we then run human evals on and make sure we're trying to doing well across the board on all, all, all of those.

**Swyx** [26:15]
Yeah, you mentioned three things. Is it literally three or, or, or is it three out of like 20 things? How wide is the ontology?

**Aarush Selvan** [26:20]
I, I basically just told the, told the, like-

**Swyx** [26:22]
Like the full set.

**Aarush Selvan** [26:22]
Yeah. I told-

**Swyx** [26:23]
Okay.

**Aarush Selvan** [26:23]
No, no, no. I told you the like extremes, right? So like-

**Swyx** [26:25]
The extremes. Okay.

**Aarush Selvan** [26:25]
Yeah. And then we-

**Swyx** [26:26]
I understand

**Aarush Selvan** [26:26]
... we, we had like several, several midpoints. So basically, yeah, going from like something super broad and shallow to something very specific and deep. We weren't actually sure which end of the spectrum users are gonna really resonate with.

And then on top of that you have compounds of those, right? So you can have things where you wanna make a plan, right? Like, uh, a great one is like, "I wanna plan a wedding in, you know, Lisbon, and I, you know, I need you to help with like these 10 things," right?

And so-

**Swyx** [26:51]
Oh, that becomes like a project with research enabled.

**Aarush Selvan** [26:54]
Right. And so then it needs to research planners and venues and catering, right? And so there's, there's sort of compounds of when you start combining these different underlying ont- ontology types. And so that we also thought about that when we, when we tried to put, put together our eval set.

**Swyx** [27:09]
What's the maximum conversation length that you allow or design for?

**Aarush Selvan** [27:13]
We don't have any hard limits on the how many turns you can do. One thing I will say is most users don't go very deep right now.

**Swyx** [27:21]
Yeah.

**Aarush Selvan** [27:22]
It might just be that it takes a while to get comfortable, and then over time you start pushing it further and further. But like right now we don't see a ton of users.

**Swyx** [27:30]
I, I think the, the way that you, uh, visually present it suggests that you stop when the doc is created.

**Aarush Selvan** [27:36]
Right.

**Swyx** [27:36]
So you don't actually really encourage... The, the, the UI doesn't encourage cont- ongoing chats as though it was like a project.

**Aarush Selvan** [27:43]
Right. I think, I think there's definitely some things we can do on the UX side to basically invite the user to be like, "Hey, this is the starting point. Now let's keep going together. Like, where else would you like to explore?"

**Swyx** [27:55]
Mm.

**Aarush Selvan** [27:56]
So I think there's definitely some, some explorations we, we could do there. I think the... In terms of sort of how deep, I, I don't know. We've seen people internally just really push this thing-

**Swyx** [28:05]
Yeah

**Aarush Selvan** [28:06]
... uh, to quite, quite s- a lot of ways.

**Mukund Sridhar** [28:08]
I think the other, other thing I think will change with, with time is people kind of uncovering different ways to use, uh, Deep Research as well. Like for, for the wedding planning thing, for example. It's, it's not one of the, you know, first thing that comes to mind when, when we tell people about this product.

So that's another, uh, thing I think as people explore and, and, and find that this can do these various different kinds of things, some of this can naturally lead to longer conversations. And even for us, right? When we dogfooded this, we saw people use it in like ways we hadn't really thought of before.

**Swyx** [28:42]
Yeah.

**Mukund Sridhar** [28:42]
So, uh, that was one. Because this was like a little new, like we didn't know, like, will users wait for five minutes? What kind of tasks will... Are they, you know, gonna try for something like that takes five minutes?

So our primary goal was not to specialize in, in a, in a particular vertical or, or target one type of user. We just wanted to put this in the hands of like, like we had like this busy parent persona and, and like various different user profiles and, and see like what people try to use it for and learn more from that.

**Swyx** [29:12]
And how does the ontology of the DR use case tie back to like the Google main product use cases? So you mentioned shopping as one ontology, right? There's also Google Shopping.

### Market & Latency

**Aarush Selvan** [29:23]
Yeah.

**Swyx** [29:23]
To me, this sounds like a much better way to do shopping than going on Google Shopping and looking at the, the wall of items. How do you collaborate internally to figure out where AI goes?

**Aarush Selvan** [29:33]
Yeah, that's a good question. So when, when I meant like shopping, I, I sort of tried to boil down underneath what exactly is the behavior, and that's really around like I called it like options exploration. Like you just wanna be able to see...

And whether you're shopping for summer camps or shopping for a product or shopping for like scholarship opportunities, it's sort of the same action of just like I need to curate from a large... Like I need to sift through a lot of information to curate a bunch of options for me.

So that's kind of what we tried to distill down rather than like thinking about it as a vertical. But yeah, Google Search is, is like awesome if you want to have really fast answers. You've got high intent for like I know exactly what I want, and you want like super up-to-date information, right?

And I still do kind of like Google Shop because it's like multimodal, you see the best prices and stuff like that. I think creating a good shopping experience is hard, especially like when you need to look at the thing.

If I'm shopping for shoes and like I, I don't wanna use Deep Research because I wanna look at how the shoes look. But if I'm shopping for like HVAC systems, great, like I don't care how it looks or I don't even know what it's supposed to look like, and I'm fine using Deep Research because I really wanna understand the specs and like how exactly does this work and the voltage rating and stuff like that, right?

So like and I need to also look at contractors who know how to install each HVAC system. So I'd say like where we really shine when it comes to shopping is those, that kind of end of the spectrum of like it's more complex and it, it matters less what it...

Like it's, it's maybe less on the consumery side of, of shopping.

**Swyx** [31:07]
One thing I, I, I've also observed, um, just about the, I guess, I guess the metrics or like the communication of what value you provide, and also this, this goes into the latency budget, is that I think there's a perverse incentives for research agents to take longer and it be perceived to be better, to people are like, "Oh, you're, you're searching like seventy, seventy websites for me."

**Aarush Selvan** [31:30]
Yeah.

**Swyx** [31:30]
You know, but like thirty of them are irrelevant. You know, like I feel like right now we're in kind of a honeymoon phase where you get a pass for all this. Be- being inefficient is actually good for you because, you know, people just care about quantity and not quality, right?

So they're like, "Oh, this thing took an hour for me, like it's doing so much work," like or it's slow.

**Aarush Selvan** [31:48]
That was super counterintuitive for us. So actually the first time I realized that, what you're saying, is when I was talking to Jason Calacanis and he's like, "Do you actually just make the answer in ten seconds and then just make me wait for the balance?"

**Swyx** [32:00]
Yeah.

**Aarush Selvan** [32:00]
Which we hadn't expected that people would actually value the, the like work that it's putting in because-

**Swyx** [32:07]
You were actually worried about it.

**Aarush Selvan** [32:08]
We were really worried about it. We were like... I remember we actually built two versions of Deep Research.

**Swyx** [32:12]
Ah.

**Aarush Selvan** [32:13]
We had like a hardcore mode that takes like fifteen minutes. And then what we actually shipped is a thing that takes five minutes, and I even went to Eng and I was like, "There has to be a hard stop, by the way.

It can never take more than ten minutes."

**Swyx** [32:25]
Yep.

**Aarush Selvan** [32:25]
Because I think at that point, like users will just drop off. Um-

**Swyx** [32:28]
Nope.

**Aarush Selvan** [32:29]
But, but what's, what's been surprising is like that's not the case at all, and it's been going the o- the other way. Because when we worked on Assistant at least and other Google products, the metric has always been if you improve latency, r- like all the other metrics go up, like satisfaction goes up, retention goes up, all of that, right?

And so when we pitch this, it's like, hold on. In contrast to like all Google orthodoxy, we're actually gonna slow everything right down and we're gonna hope that like users still stay.

**Swyx** [32:56]
Not on purpose.

**Aarush Selvan** [32:57]
Yeah.

**Swyx** [32:57]
Not on purpose.

**Aarush Selvan** [32:58]
Right.

**Mukund Sridhar** [32:58]
Yeah. I think it comes down to the trade-off, like what are you getting in return for-

**Swyx** [33:02]
Yes

**Mukund Sridhar** [33:02]
... for, for the wait. And from an engineering/modeling perspective, it's just trading off inference, compute, and time to s- to do two things, right? Either to explore more to be like more complete or to verify more on things that you probably know already.

And since it's like a spectrum and we don't claim to have found the perfect spot, uh, we had to start somewhere and we're trying to see where... Like there, there's probably some cases where you actually care about verifying more than the others.

In an ideal world based on the query and conversation history you know what that is. So I think, yeah, it basically boils down to these three things. From a user perspective, am I getting the right value add? From an engineering/modeling perspective, are we using the compute, uh, to either explore effectively and, uh, also, uh, verify and go, uh, in depth for things that are, uh, vague or uncertain in the initial steps?

The other point about the, the more number of websites, I think again it comes with a trade-off. Like sometimes you wanna explore more early on before you kind of narrow down on either the sources or the topics you wanna go deep.

So that's one of the... If you look at like the way at least for most queries, the, the way, uh, Deep Research works here is initially it'll go broad. It'll... If you look at the kinds of websites, it's trying to explore all the different topics that we measured in the research plan, and then you would see choices of websites getting a little bit narrower on, on a particular topic or a particular entity that it has come across and so on.

So that's roughly how the number kind of fluctuates. So we, we don't do anything deliberate to either keep it low or, or, you know, uh, try to, uh-

**Swyx** [34:45]
Would you be... Would it be interesting to have an explicit toggle for amount of verification versus amount of search?

**Aarush Selvan** [34:52]
I think so. I think like, um, users would always just hit that toggle. I think I, I worry that like-

**Swyx** [34:57]
Max everything.

**Aarush Selvan** [34:58]
Yeah. If you like give a max power button, users are always just gonna hit that button, right? So then the question becomes like why don't you just decide from the product POV where's the right, where's the right balance?

**Swyx** [35:07]
OpenAI has a, a preview of this like... I think it's e- either in Anthropic or OpenAI, and there's a preview of this, uh, model routing feature where you can choose, uh, intelligence, cheapness, and speed. Uh, but then they're all zero to one values, so then you just choose one for everything.

**Aarush Selvan** [35:24]
Right.

**Swyx** [35:25]
Obviously they're gonna like do a normalization num- thing, but user's always gonna want one, right?

**Mukund Sridhar** [35:30]
We've discussed this a bit. Like from, if I wear my pure user hat, I don't wanna set anything. Like I, I come with a query, you figure it out. Like-

**Swyx** [35:38]
Yeah

**Mukund Sridhar** [35:38]
... sometimes I feel like there will be based on the query... Like for example, right? If I'm, uh, asking about, hey, how does rising rates from the Fed household income for a middle class and, and how, how has it traditionally happened?

These kind of things you wanna be very accurate, uh, and you wanna be very precise on historical trends of this, and so on and so on. Whereas there is a little bit more leeway when you're saying, "Hey, I'm trying to find businesses near me to go celebrate my birthday-"

**Swyx** [36:06]
Yeah

**Mukund Sridhar** [36:06]
... or something like that. So in an ideal world, we kind of figure that trade-off based on, uh, the conversation history and the topic. I don't think we are there yet, uh, uh, as a, as a, as a research community, and it's, it's an interesting challenge by itself.

**Swyx** [36:21]
So this reminds me a little bit of the, the NotebookLM approach. Uh, Raisa, we also asked this thing to Raisa, and she was like, "Yeah, just people want to click a button and see magic."

### Feedback & Learnings

**Aarush Selvan** [36:31]
Yeah, like peop- like, like you said, you just hit start every time, right? You don't-- most people don't even-

**Swyx** [36:35]
Yeah.

**Aarush Selvan** [36:36]
... wanna, wanna add the plan.

**Swyx** [36:37]
No, so, so, okay, my feedback on this, if you, if you want feedback, is that, uh, I am still kind of a champion for Devin in a sense that Devin will show you the plan while it's working the plan, and you can say like, "Hey, the plan is wrong," and I can chat with it while it's still working.

**Aarush Selvan** [36:53]
Right.

**Swyx** [36:54]
And it will live update the plan and then, you know, pick off the next item on the plan. I think it's static, right? Like, while you're working on a plan, I cannot chat. It's just normal.

**Aarush Selvan** [37:01]
Yeah.

**Swyx** [37:01]
Bolt, Bolt also has this. Like, uh, you know, and that's the most default experience. But I think you should never lock the chat. You should always be able to chat with the plan and update the plan, and the plan, uh, scheduler, whatever orchestration system you have under the hood, should just pick off the next job on the list.

That would be my two cents.

**Aarush Selvan** [37:18]
Especially if we spend more time researching, right? 'Cause, like, right now if you watched that query we just did, it was done within a few minutes, so your chance, your opportunity to chime in was actually like-- or it left the research phase after a few minutes, so your opportunity to chime in and steer was less.

**Swyx** [37:32]
Mm.

**Aarush Selvan** [37:32]
But especially imagine, you could imagine a world where these things take an hour, right, and you're doing something really complicated, then yeah, like, your intern would totally come check in with you, be like, "Here's what I found. Here's, like, some hiccups I'm running into the plan.

Give me some steer on how to change that or how to change direction." And, and you would, you would do that with them. So I totally would see, especially as these tasks get longer, we actually want the user to come engage-

**Swyx** [37:57]
Yeah

**Aarush Selvan** [37:57]
... way more, um-

**Swyx** [37:58]
Yeah

**Aarush Selvan** [37:58]
... to, like, create a, create a good output.

**Swyx** [38:01]
I guess Devin had to do this because some of these jobs, like, take hours.

**Aarush Selvan** [38:04]
Right.

**Swyx** [38:04]
So yeah.

**Aarush Selvan** [38:05]
Yeah. I can totally imagine.

**Swyx** [38:06]
And, and it's pervasive since this where they charge by hour.

**Aarush Selvan** [38:09]
Oh.

**Swyx** [38:10]
So they make more money the slower they are.

**Aarush Selvan** [38:11]
Interesting. Have we thought about that?

Have we?

**Swyx** [38:15]
I, that's why I'm just-- I'm calling this out because, like, everyone is like, "Oh my God, it takes hours for-- it does hours of work autonomously for me."

**Aarush Selvan** [38:20]
Yeah.

**Swyx** [38:20]
And like, they are like, "Okay, it's good."

**Aarush Selvan** [38:22]
Yeah.

**Swyx** [38:22]
But, like, this is a honeymoon phase. Like, at some point-

**Aarush Selvan** [38:24]
Right

**Swyx** [38:25]
... we're gonna say like, "Okay, but you know, it's very slow."

**Aarush Selvan** [38:29]
Yeah.

**Swyx** [38:30]
Anything else that, like... I mean, obviously within Google you have a lot of other initiatives. You-- I'm sure you, like, sit close to the NotebookLM team. In any learnings that are coming from shipping AI products in general.

**Aarush Selvan** [38:42]
They're really awesome people. Like, they're really nice, friendly thought. Just like as people, I'm sure you met them, you like realized this with Razor and stuff. So, like, they've actually been really, really cool collaborators or just, like, people to bounce ideas off.

I think o-one thing I found really inspiring is they just picked a problem, and hindsight's 20/20, but, like, in advance just like, "Hey, we just wanna build, like, the perfect IDE for you to do work and, like, be able to upload documents and ask questions about it and just make that really, really good."

And I think we were definitely really inspired by their ability, their vision of just like, let's pick a, a simple problem, really go after it, do it really, really well, and have-- be opinionated about how it should work, and just hope that users also resonate with that, and that's definitely something that we tried to learn from.

Separately, they've also been really good at, you know, and maybe Mukund you wanna chime in here, just extracting the most out of Gemini 1.5 Pro, and they were really friendly about just, like, sharing their ideas about how to do that.

**Mukund Sridhar** [39:38]
Yeah. I think, I think you'll, you'll, you learn a bit, like, when you're trying to do the last, last mile of, of these products and, and, and pitfalls of, of any, any given model and so on. So yeah, we definitely have a healthy relationship and, and, and share notes, and we're doing the same for other, other products.

**Swyx** [39:54]
You'll never merge, right? It's just different teams.

**Aarush Selvan** [39:57]
They are different teams. So they're in like-

**Swyx** [39:58]
Yeah

**Aarush Selvan** [39:58]
... Labs as an organization that the mission of that is to really explore kind of-

**Mukund Sridhar** [40:02]
What's possible

**Aarush Selvan** [40:03]
... different bets and, and explore what's possible. Um-

**Swyx** [40:05]
Even though I think there's a paid plan for NotebookLM now.

**Aarush Selvan** [40:08]
Yeah. So I think-

**Swyx** [40:08]
So it's, it's a product

**Aarush Selvan** [40:09]
... and it's the same plan as us actually.

**Swyx** [40:11]
Yeah, yeah, yeah.

**Aarush Selvan** [40:11]
So it's like-

**Swyx** [40:12]
It's more than just the labs is what I'm saying.

**Aarush Selvan** [40:14]
It's more than just labs- ... 'cause I mean, yeah, ideally you want things to graduate and into-

**Swyx** [40:18]
Yeah, yeah

**Aarush Selvan** [40:19]
... and, and stick around. But hopefully one thing we've done is, uh, like not created different SKUs, but just being like, "Hey, if you pay the AI Premium School-"

**Swyx** [40:28]
Gemini 1.

**Aarush Selvan** [40:29]
Yeah.

**Swyx** [40:29]
Yeah, whatever.

**Aarush Selvan** [40:29]
You get, you get everything.

**Swyx** [40:30]
Good thing.

### Frontier Tech

**Alessio** [40:31]
What about learning from others? Obviously, I mean, OpenAI has Deep Research, literally-

**Aarush Selvan** [40:36]
What?

**Alessio** [40:36]
... has the same name.

**Aarush Selvan** [40:37]
What?

**Alessio** [40:37]
I'm sure-

**Swyx** [40:38]
Yeah, uh-

**Alessio** [40:38]
... I'm sure there's a lot of, you know, contention. Is there anything you've learned from other people trying to build similar tools? Like, do you have opinions on maybe what people are getting wrong that they should do differently?

It seems like from the outside, a lot of these products look the same. Ask for a research, get back a research, but obviously when you're building them, you understand the nuances a lot more.

**Aarush Selvan** [41:00]
When we built Deep Research, I think there was a few things that we took a few different bets, uh, around how this-- how it should work, and what's nice is some of that is actually where we feel like was the right way to go.

So we felt like agents should be transparent around telling you up front, especially if they're gonna take some time, what they're gonna do. So that's really where that research plan, we showed that in a card. We really wanted to be very publisher forward in this product.

So while it's browsing, we wanted to show you, like, all the websites it's reading in real time, make it super easy for you to, like, double-click into those while it's browsing. And the third thing is, you know, putting it into a side-by-side artifact so that you could ideally easy for you to read and ask at the same time.

And what's nice is you kind of, as other products come around, you see some of these ideas also appearing in, in other iterations of this product. So I definitely see this as a space where, like, everyone in the industry is learning from each other, good ideas get reproduced and built upon, and so yeah, we'll, we'll definitely keep iterating on and kind of following our users and seeing, seeing how we can make, make our feature better.

But yeah, I think, I think like it's, it's like this is the way the industry works, is like everyone's gonna kind of see good ideas and wanna replicate and, and build off of it.

**Alessio** [42:13]
And on the model side, OpenAI has the o3 model, which is not available through the API, the full one. Have you tried already with the 2 model? Like, is it a big jump, or is a lot of the work on the post-training?

**Mukund Sridhar** [42:26]
Yeah. I, I would say stay tuned. Uh, definitely it currently is running on, on 1.5. The, the new generation models, especially with these thinking models, they unlock a few things. So I think one is obviously the, the better capability in like analytical thinking, like in math, coding, and these type of things.

But also this notion of, you know, as they- Produce thoughts and think, uh, before taking actions. They kind of inherently have this notion of being able to critique the, the partial steps that they take and so on. So yeah, uh, we're definitely exploring multiple different options to make better value for the, for our users as we, as we iterate.

Yeah.

**Swyx** [43:04]
Yeah. I feel like there's a little bit of a conflation of inference time compute here in the sense of, like, one, you can inference time compute within the mod- the thinking model.

**Mukund Sridhar** [43:13]
Right.

**Swyx** [43:14]
And then two, you can inference time compute by searching and-

**Mukund Sridhar** [43:17]
Yeah

**Swyx** [43:17]
... reasoning.

**Mukund Sridhar** [43:17]
Doing more iterative.

**Swyx** [43:18]
I wonder if they, like, gets in the way. Like, when you... Presumably you've tested thinking plus Deep Research, if the thinking actually does a little bit of verification, so maybe saves you some time or it, like, tries to draw too much from its internal knowledge and then therefore it searches less.

You know, like, does it step on each other?

**Mukund Sridhar** [43:37]
Yeah. No, I think that's a, that's a really nice call-out. And, uh, this also goes back to, uh, to the kind of use case. Uh, the reason I bring that up is there are certain things that I can tell you from model memory last year's the, the Fed did X number of updates and so on, but unless I sourced it, it's, it's gonna be-

**Swyx** [43:57]
Hallucinated. Yeah

**Mukund Sridhar** [43:58]
... Yeah. Like, one is a hallucination or even if I got it right, as a user-

**Swyx** [44:01]
Yeah

**Mukund Sridhar** [44:01]
... I'd be very wary of, uh, that number unless I'm able to, like, uh, source the, uh, the, the, the .gov website for it and so on, right? So that's another challenge. Like, there are things that you might not optimally spend time verifying even though the model's like, like this is a very common fact the model already knows and it's able to, like, uh, reason over.

And balancing that out between, uh, trying to leverage the model memory versus being able to ground this in is in, uh, you know, some kind of a, a source is, is the challenging part. And I think as, as like you rightly called out with the thinking models, this is, this is even more pronounced because the models know more.

They're able to, like, draw second order insights more just by reasoning over.

**Swyx** [44:44]
Technically they don't know more, they just use their internal knowledge more, right?

**Mukund Sridhar** [44:49]
Yes, but also like for, for example, uh, things like math, uh-

**Swyx** [44:53]
I see. They've been-

**Mukund Sridhar** [44:54]
Mm.

**Swyx** [44:54]
They've been post-trained to do better math.

**Mukund Sridhar** [44:56]
Yeah, I think they just, they probably do way better job and in, uh, like in, in math-

**Swyx** [45:01]
Sure

**Mukund Sridhar** [45:01]
... than, than previous, so in that sense, yeah.

**Swyx** [45:03]
Yeah, I mean obviously reasoning is a topic of huge interest and people want to know what a engineering best practice is. Like we think we know like, you know, how to prompt them better, but engineering with them I think also very, very unknown.

Again, you guys are gonna be the first to figure it out.

**Mukund Sridhar** [45:20]
Yeah, definitely interesting times and yeah-

**Swyx** [45:22]
No pressure, Mukund. Yeah. If you have tips, let us know. While we're on the sort of technical elements and technical bends, I'm interested in like other parts of the Deep Research tech stack that might be worth calling out.

Any hard problems that you solved just more generally?

**Mukund Sridhar** [45:38]
Yeah, I think the iterative planning one to do it in a generalizable way.

**Swyx** [45:42]
Yeah.

**Mukund Sridhar** [45:42]
That was the thing I was ve- most wary about. Like, you don't wanna go down the route of being able to teach how to plan iteratively per domain or like per type of problem. Uh, like, like even in the going back to the ontology, if, if you had to teach the model for every single type of ontology how to come up with these traces of planning, that would've been nightmarish.

So, uh, trying to do that in a super data efficient way by, you know, leveraging a lot of like things model memory as well as like there's this very tricky balance when you work on like on the, the product side of, of the any of these models is knowing how to post-train it just enough without losing things that it knows in pre-training.

Basically not overfitting in, in, in the most trivial sense I guess. But yeah, so the techniques there, data augmentations there and, and multiple experiments to tune this trade-off, I think, uh, that's, that's one of the challenges.

**Swyx** [46:37]
Yeah.

**Mukund Sridhar** [46:38]
Yeah.

**Swyx** [46:38]
On the orchestration side, this is basically you're, you're spinning up a job. I'm an orchestration nerd, so how do you do that? Is, is like a self internal tool?

**Mukund Sridhar** [46:45]
Yeah, so we built this, uh, asynchronous platform for Deep Research-

**Swyx** [46:49]
Okay

**Mukund Sridhar** [46:50]
... which is basically to like most of our interactions before this were like sync in nature. They like, you know-

**Swyx** [46:56]
Yeah

**Mukund Sridhar** [46:57]
... uh-

**Swyx** [46:57]
All chat-

**Mukund Sridhar** [46:58]
Right

**Swyx** [46:58]
... things are sync, right?

**Mukund Sridhar** [46:59]
Exactly.

**Swyx** [46:59]
And now, now you can leave the chat and come back.

**Mukund Sridhar** [47:01]
Exactly. And-

**Swyx** [47:02]
Close your computer and-

**Mukund Sridhar** [47:04]
And, and now it's on Android and-

**Swyx** [47:05]
It's, yeah

**Mukund Sridhar** [47:06]
... rolling out on iOS.

**Swyx** [47:07]
We should make a plugin.

**Mukund Sridhar** [47:07]
So yeah.

**Swyx** [47:08]
Yeah. You know, I saw, I saw you say that on the-

**Mukund Sridhar** [47:10]
I told you we switch roles sometimes.

**Swyx** [47:13]
Okay. Like you're reminding him, right? Yeah, yeah. Uh, yeah, we wrapped on all Android phones and then iOS is this week. Um, but yeah what's, what's neat though is like you can close your computer-

**Mukund Sridhar** [47:25]
You get a notification on your phone

**Swyx** [47:26]
... you get a notification and so on. So it's, so it's some kind of sync engine that you made.

**Mukund Sridhar** [47:29]
Yes. Yes. So we-- the other, uh, one is this notion of synchronicity and the, the, the user able to leave, but also if you're, if you build like five, six minute jobs, there are bound to be like failures and you don't wanna like-

**Swyx** [47:41]
Retry

**Mukund Sridhar** [47:42]
... lose your progress-

**Swyx** [47:43]
Yeah

**Mukund Sridhar** [47:43]
... and so on. So this notion of like keeping state, uh, knowing, uh, what to retry and, and, and kind of keep the journey going.

**Swyx** [47:50]
Is there a public name for this or just some internal thing?

**Mukund Sridhar** [47:53]
Uh, no, I don't think there's a public name for this. Yeah.

**Swyx** [47:55]
Okay. Yeah, right. Data scientists would be like, "This is a Spark job," or, you know, it's like a Rave, you know, thing or whatever. In the old Google days might be like MapReduce or, you know, whatever. But like it's, it's a different scale and nature of work-

**Mukund Sridhar** [48:08]
Yeah. Yeah, yeah

**Swyx** [48:09]
... than those things. So we j- I'm trying to find a name for this and right now-

**Mukund Sridhar** [48:13]
We can name it now.

**Swyx** [48:14]
This is our opportunity to name it.

**Mukund Sridhar** [48:15]
Yeah, we can name it now.

**Swyx** [48:16]
Yeah. Yeah. Well, the, the, the classic name 'cause I used to work in this area. This is why I'm asking.

**Mukund Sridhar** [48:21]
I see. I see.

**Swyx** [48:21]
It's, it's workflows. Nice. Yeah. Yeah. Uh, sort of durable workflows.

**Mukund Sridhar** [48:24]
This was, uh, like back when you were involved with-

**Swyx** [48:26]
Temporal.

**Mukund Sridhar** [48:26]
Okay.

**Swyx** [48:27]
Uh, so Apache Airflow, Temporal. You guys were both at Amazon by the way. Yeah. AWS Step Functions would be one of those where you define a graph of execution but, uh, Step Functions are more static and would not be as able to ac-accommodate Deep Research style backends.

What's neat though is we built this to be like quite flexible so like you can imagine once you start doing hour or multi-day jobs like-

**Mukund Sridhar** [48:48]
Yeah. You have to model what the agent wants to do.

**Swyx** [48:49]
Exactly. And but also like ensure it re- like it's stable

**Aarush Selvan** [48:53]
You know, for, for like hundreds of LLM calls.

**Swyx** [48:56]
Yeah. It's boring but like, you know, this is the thing that makes it run autonomously, you know?

**Aarush Selvan** [49:01]
Right.

**Swyx** [49:02]
Yeah. So like it's... Yeah. Anyway, I'm excited about it. Just to clo- close out the OpenAI thing, I would say OpenAI easily beat you on marketing, and I think it's because you don't launch your benchmarks. And my question to you is, should you care about benchmarks?

### Benchmarks & Discovery

**Swyx** [49:17]
Should you care about humanity's last exam or not MMLU, but whatever.

**Aarush Selvan** [49:21]
The like-- I think benchmarks are great. The thing we wanted to avoid is like the day Kobe Bryant entered the league, who was the president's nephew, and like weird like benchmarks.

**Swyx** [49:31]
He's a big Kobe fan.

**Aarush Selvan** [49:32]
Okay, perfect.

**Swyx** [49:33]
I see.

**Aarush Selvan** [49:33]
Just like these like weird things that like nobody talks that way. So like why would we over solve for like some sort of a benchmark that doesn't necessarily represent the, the product experience we wanna build.

**Swyx** [49:44]
Okay.

**Aarush Selvan** [49:44]
Nevertheless, like benchmarks are great for the industry-

**Swyx** [49:47]
Yeah

**Aarush Selvan** [49:48]
... and like rally a community and, and help us like understand where we're at. I don't know. Do you have any...?

**Mukund Sridhar** [49:53]
No, I think you, you kind of hit the points. I think the-- for us, our primary goal is like solving the deep research user value for the user-

**Swyx** [50:02]
Yeah

**Mukund Sridhar** [50:02]
... uh, use case. The benchmarks, at least the ones that we are seeing, they don't directly translate to the product. There's definitely some technical challenges that you can benchmark against, but they don't really... Like if I do great on HLE, that doesn't really mean I'm a great deep researcher.

So, uh, we wanna avoid, uh, going into that rabbit hole a bit. But I-- we also feel like, yeah, benchmarks are great, especially in the whole GenAI space with like models coming every other day and everybody claiming to be like Soda.

**Swyx** [50:33]
Soda.

**Mukund Sridhar** [50:33]
So, uh, uh, it's, it's, uh, it's tricky. The other big challenge with benchmarks, especially when it comes to like the models these days, is the output space entropy is like everything is like, uh, text and, uh... So there's a notion of verifying even if you got the right answer, different labs do it in like different ways and but it-- we all compare numbers.

So there's a lot of, you know, art slash, uh, figuring out like how you verify this or how you run this in a, a level plane.

**Swyx** [51:01]
Mm.

**Mukund Sridhar** [51:01]
Uh, but yeah. So I think there's trade-offs. There's definitely value to doing benchmarks. Uh, but at the same time we also-

**Aarush Selvan** [51:06]
From like a selfish PM perspective-

**Mukund Sridhar** [51:08]
Yeah

**Aarush Selvan** [51:09]
... benchmarks are a really great way to motivate researchers. Like, um-

**Swyx** [51:12]
Yeah, make number go up.

**Aarush Selvan** [51:13]
Exactly. Or just like prove you're the best. Like-

**Swyx** [51:16]
Yeah

**Aarush Selvan** [51:16]
... it's like a really good way of like rallying the researchers within your company. Like, um, I used to work on the MLPerf benchmarks, and like that was like, yeah, you'd put like a bunch of engineers in a room, and in a few days they'd do like amazing performance improvements on, on our TPU stack and things like that, right?

So just like having a competitive nature and a pressure like really motivates people. There's one benchmark that is impossible to benchmark, but I just wanna leave you with it, which is that deep research, most people are chasing this idea of discovering new ideas, and deep research right now will summarize the web in a way that, you know, is, is, is much more readable, but it will ne- like, you know, what will it take to discover new things from the things that you've searched?

**Mukund Sridhar** [51:57]
First, I think the thinking style models definitely help here because, uh, they are significantly better on how they reason natively and being able to, you know, draw these second-order insights, which is like very premise. Uh, the-- like if you can't do that, you can't, uh, think of doing what you mentioned.

So that's, that's one step. And, uh, the other thing is I think it also depends on the domain. So sometimes you can riff with a model for like new hypothesis, but depending on the domain, you might not be able to verify that hypothesis, right?

So like coding math, there are reasonably good tools that the model already knows to interact with, and you can run a verifier, test the hypothesis, and so on. Like even if you think about it from a purely agent perspective, saying, "Hey, I have this hypothesis in this area.

Go figure out and come back to me," right? But let's say you're a chemist, right? So, uh, what are you gonna do there? We don't have like synthetic environments yet where the model's able to verify these hypothesis by playing in a playground and have this like, uh, very accurate verifier or a reward signal.

The computer uses another one, uh, where, where there are this both, both in the open source research and so on, there's like nice playgrounds coming up. Uh, so I think for f- if you're talking about truly being able to come up with...

My personal opinion is the model doesn't-- uh, has to do the second order thinking and so on that we are seeing now with these new models, but also be able to play-

**Swyx** [53:21]
Verify

**Mukund Sridhar** [53:21]
... and test that out in an environment where you can, you know, verify and, and, and give it feedback so that it can continue iterating.

**Swyx** [53:28]
Yeah. So basically, like code sandboxes for now.

**Mukund Sridhar** [53:33]
Yeah, yeah. So in those kind of cases, I think, yeah, it's, it's a little bit more easy to envision this like end to end, uh, but not for all, uh, domains or verticals.

**Swyx** [53:41]
Physics engines.

**Mukund Sridhar** [53:42]
Yeah, yeah.

**Swyx** [53:43]
So if you think about agents more broadly, there's like a lot of things that go into it. What do you think are like the most valuable pieces that people should be spending time on? Like things that come to mind that I'm seeing a lot of early stage companies is like memory.

You know, like we already touched on evals, we touched a little bit on, uh, tool call. There's kinda like the auth piece. Like should this agent be able to access this? If yes, how do you verify that? What are things that you want more people to work on that will be helpful to you?

### Agent Building

**Aarush Selvan** [54:12]
I can take a stab at this from the lens of like deep research, right? Like I think some of the things that we're really interested in, in how we can push this agent are one, like similar to memories, like personalization, right?

Like if I'm giving you a research report-

**Swyx** [54:26]
Oh

**Aarush Selvan** [54:26]
... the way I would give it to you if you're a fifteen-year-old in high school should be totally different to the way I give it to you if you're like a PhD or postdoc, right?

**Swyx** [54:33]
You can prompt it.

**Aarush Selvan** [54:34]
You can prompt it, right. But the second thing though is like it should like ideally know where you're at and like w- everything you know up to that point, right, and kind of further customize, right? Have, have this understanding of like where you are in your learning journeys.

I think modality will be also really interesting. Like right now we're, we're text in, text out. We should go multimodal in, right? But also multimodal out, right? Like I would love if my reports are not just

**Swyx** [55:03]
Text, but like charts, maps, images, like make it super interactive and multimodal, right? And optimize for the type of consumption, right? So the way in which I might put together an academic paper should be totally different to the way I'm trying to do like a learning program for a, for a kid, right?

And just the way it's structured. Ideally, like you wanna do things with generative UI and things like that to really customize reports. I think tho- those are definitely things that I'm personally interested when it comes to like a, a research agent.

I think the other part that's super important is just like we will reach the limits of the open web, and you want to be able to... Like, a lot of the things that people care about are things that are in their own documents, their own corpuses, things that are within subscriptions that they personally really care about, right?

Like, especially as you go more niche into specific industries. And ideally, you want ways for people to be able to complement their Deep Research experience with that content in order to further customize their answers.

**Mukund Sridhar** [55:57]
There's two answers to this. So one is, I feel, uh, in terms of like the approach for us at least, or for me rather, trying to figure out the core mission for like an agent building that, I feel like it's still early days for us like to try to platformatize or like try to build these, oh, there are these five horizontal pieces and you can plug and play and build your own agent.

My personal opinion is we are not there yet. In order to build a super engaging agent, I would-- if I were to start, uh, thinking of a new idea, I would, I would start from the idea and try to just, just do that one thing really well.

Yes, at some point there will be a time where like these common pieces can be pulled out and, and, and, and, you know, platformatized. I know there's a lot of work across companies and in the open source community about providing these tools to re- really build, uh, agents very easily.

I think those are super useful to start, uh, building agents. But at some point, once those tools enable you to build the basic layers, I think me as an individual would, would, you know, try to focus on really curating one experience before going super broad.

**Alessio** [57:05]
Yeah. We have Brett Taylor from Sierra and he said they mostly built everything in-house.

**Swyx** [57:09]
Built everything in-house-

**Alessio** [57:09]
Yeah

**Swyx** [57:09]
... which is very sad for VCs. They wanna find the next great framework and tooling and all that.

**Mukund Sridhar** [57:16]
Yeah. No, but, but the space is moving so fast, like I, like the problem I described might be obsolete six months from now and, and I, I don't know, like.

**Swyx** [57:23]
We, we'll fix it with one more LLM ops platform.

**Mukund Sridhar** [57:25]
Yes. Yes.

**Swyx** [57:27]
Okay. So just a, just a final, final, uh, point on just plugging your talk. People will be hearing this before your talk. Uh, what are you gonna talk about? What are you looking forward to in New York?

**Mukund Sridhar** [57:35]
I mean-

**Swyx** [57:35]
I would love to like actually learn from you guys. Like-

### Conference Talk

**Mukund Sridhar** [57:37]
Yes

**Swyx** [57:37]
... what would you like us to talk about now that we've had this conversation with you?

**Mukund Sridhar** [57:41]
Yeah.

**Swyx** [57:41]
Yeah. What would, what do you think people would find most interesting? I think a little bit of implementation and a little bit of vision, like kind of 50/50 and it, and I think both of you can, can sort of fill those roles very well.

Everyone, you know, looks at you, uh, you're very polished Google products and I think Google always, uh, does, does polish very well. But everyone will have to want, want like Deep Research for their industry. You know- Yeah ...

he's invested in-

**Mukund Sridhar** [58:03]
Yeah

**Swyx** [58:03]
... Deep Research for finance.

**Mukund Sridhar** [58:04]
Yeah.

**Swyx** [58:04]
And they, they focus on their, their thing. And there will be Deep Researches for everything, right? Like you have created a category here that OpenAI has cloned. And so like, okay, let's, let's talk about like what are the hard problems in this brand of agent that is like probably the first real product market fit agent, I would say more so than the computer use ones.

This is the one where like, yeah, people are like, yeah, easily pays for $200 worth, uh, a month worth of, of stuff. Probably 2,000 once you get it really good. So I'm like, okay, like let's talk about like how to do this right from the people who did it, and then where is this going?

So yeah.

**Mukund Sridhar** [58:37]
Yeah.

**Swyx** [58:37]
It's very simple.

**Mukund Sridhar** [58:38]
Happy to talk about that.

**Swyx** [58:40]
Sounds good.

**Mukund Sridhar** [58:40]
Yeah. Thank you.

**Swyx** [58:40]
Yeah.

**Mukund Sridhar** [58:41]
Yeah.

**Swyx** [58:41]
For me as well, uh, you know, I'm also curious to see you interact with the other speakers because then, you know, there will be other sort of agent problems that... You know, I'm very interested in personalization and very interested in memory.

I think those are related problems. Planning, orchestration, all those things. Auth and security is something that we haven't talked about. There's a lot of the web, the web that's behind auth walls. Uh, can I-- how do I delegate to you my credentials so that you can go and search the things that I have access to?

I don't think it's that hard, you know. It's just, you know, people have to get their protocols together.

**Mukund Sridhar** [59:13]
Mm.

**Swyx** [59:13]
Uh, and that's what conferences like that is hopefully meant to-

**Mukund Sridhar** [59:15]
Yeah

**Swyx** [59:15]
... meant to achieve.

**Mukund Sridhar** [59:17]
Yeah, no, I'm, I'm super excited. I think for us, like, um, it's we often like live and breathe within Google and which is like a really big place, but it's really nice to like take a step back, meet people like approaching this problem at other companies or totally different industries, right?

Like inevitably, at least where we work, we're a very consumer-focused space.

**Swyx** [59:35]
I see.

**Mukund Sridhar** [59:36]
Right?

**Swyx** [59:36]
Yeah. I'm more B2B. Yeah.

**Mukund Sridhar** [59:37]
It's also really great to understand like, okay, what's going on within the B2B space-

**Swyx** [59:43]
I see. I see. I see

**Mukund Sridhar** [59:43]
... and like within different verticals.

**Swyx** [59:45]
Yeah. The first thing they wanna do is Deep Research for my own docs, right? My, my company docs.

**Mukund Sridhar** [59:48]
Yeah.

**Swyx** [59:48]
Yeah.

**Mukund Sridhar** [59:48]
Yeah.

**Swyx** [59:48]
So obviously you're gonna get asked for that. Yeah, I mean, uh, there'll be, there'll be more, uh, to discuss. I'm really looking forward to your talk and, uh, yeah, thanks for joining us.

**Mukund Sridhar** [59:57]
Yeah.

**Swyx** [59:57]
Cool.

**Mukund Sridhar** [59:57]
Yeah. Thanks for having us. It was great to have-

**Swyx** [59:58]
Thanks so much, guys.

**Mukund Sridhar** [59:59]
Yeah.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
