# How NotebookLM Was Made

Latent Space · 2024-10-25

<https://addtry.com/8a75c365-f94f-4291-a49a-02686fb2593a>

Raiza Martin (NotebookLM lead PM) and Usama Bin Shafqat (AI engineer) explain how Google's NotebookLM built the viral 'Deep Dive' audio overview feature. They reveal the product evolved from Project Tailwind, using Gemini 1.5's long context and DeepMind speech to create a two-persona dialogue format that transforms documents into engaging podcasts. The team learned from 65,000 Discord members, leaned on best-selling author Steven Johnson for a 'tool for thought' workflow, and prioritized a single format over exposed controls to preserve unpredictability and delight. Humor and tension are not explicitly prompted but emerge from giving personas different angles. Evaulation relied on internal taste ('potatoes for chefs') before formal raters, with a Likert scale on dimensions like entertainment and groundedness. Future plans include multilingual support, API access, real-time chat, and codebase podcasting, while managing non-determinism by accepting occasional bad rolls.

## Questions this episode answers

### What makes NotebookLM's Deep Dive podcast so engaging?

Usama Bin Shafqat and Raiza Martin say the two-person conversation format creates engagement by having AI personas with different angles, building tension and withholding information. The audio models add human intonations, pauses, and laughter. The team focused on a 'gradual unrolling' of the topic, making even dry subjects feel riveting, and avoided a simple narration style.

[1:05:04](https://addtry.com/8a75c365-f94f-4291-a49a-02686fb2593a?t=3904000)

### How does NotebookLM use its Discord community to improve the product?

Raiza Martin says the 65,000‑member Discord server provides immediate feedback on features and catches server issues faster than their own monitoring. Users share real‑world use cases, which helps the team understand persistent problems and decide which features to prioritize or sunset. The direct user stories guide them to build what truly helps people.

[15:39](https://addtry.com/8a75c365-f94f-4291-a49a-02686fb2593a?t=939000)

## Key moments

- **[0:00] Intro**
  - [2:14] Raiza Martin almost left Google twice but stayed after interviewing with cool teams, eventually landing in Google Labs.
- **[4:04] Origins**
  - [4:29] Talk to Small Corpus: the October 2022 prototype at Google that let users chat with their own data, later becoming NotebookLM.
  - [6:41] User testing revealed the first action 100% of users take in NotebookLM is asking it to summarize their document.
  - [7:45] Raiza Martin launched a Discord server for NotebookLM to get immediate user feedback, growing it to over 65,000 members.
- **[9:26] Discord**
- **[12:15] Architecture**
  - [12:15] Usama Bin Shafqat explains that the two AI hosts have different perspectives, questioning each other to create an engaging audio overview.
  - [16:02] Raiza Martin: 'You don't actually know what's going to happen when you press generate... that initial wow was because you didn't know.'
- **[18:00] Visionary**
  - [18:01] Raiza Martin recalls her shock when New York Times bestselling author Steven Johnson joined the NotebookLM team as a 'chief dreamer'.
  - [22:15] Steven Johnson demonstrated NotebookLM's ability to analyze handwritten Marie Curie notes from PDFs, inspiring the team's vision for source inputs.
- **[23:09] Data Pipeline**
  - [23:26] Raiza Martin divides NotebookLM into three areas: source inputs, processing capabilities, and output creation, with features like DOCX support planned.
  - [27:54] NotebookLM Audio Overviews went viral when people discovered they could upload their LinkedIn profile and get a personalized hype session.
  - [28:55] Googlers used NotebookLM to create Q3 performance review summaries, saying it made them 'feel really good' before manager meetings.
- **[29:50] Evals**
  - [30:21] Usama Bin Shafqat nicknamed an early eval set 'Potatoes for Chefs,' and the team relied on internal listening sessions before formal rating.
  - [33:28] To build engaging audio, the NotebookLM team broke down 'entertainment' into measurable factors like tension and withholding information, then dialed them to the max.
- **[39:04] Engagement**
  - [42:36] Usama Bin Shafqat: 'Some people think you need a linguistics person in the room for making a good chatbot, but that's not actually true.'
  - [44:09] Raiza Martin challenges the view that humans make chatbots interesting, asking: 'What does it mean to flip it and say, "No, you be interesting now"?'
  - [46:04] The NotebookLM AI hosts spontaneously generated humor in a viral audio overview of the research paper 'Chicken, Chicken, Chicken.'
  - [50:26] Raiza Martin: 'Is it interesting to compare two retirement plans? No. But to listen to these two talk about it, you'd think it was the best thing ever invented.'
  - [51:38] Q: When building AI products, how do you decide between prompt engineering and writing code? A: Raiza Martin says it depends on the desired outcome and is a craft as much as a science.
- **[53:35] Feature Requests**
  - [53:52] Raiza Martin confirms NotebookLM is exploring an API for developers, while continuing to invest in the end-user product.
  - [55:05] Usama Bin Shafqat reveals early multilingual tests showed dialect mixing, such as Canadian vs. Parisian French, and they are fixing dialect consistency.
  - [57:50] Raiza Martin announces they will ship a fast-follow feature for users to control audio output via system prompts, likely a text box.
  - [58:26] Raiza Martin on avoiding feature overload: 'This is not cool. This is not fun. This is not magical.'
  - [59:52] Raiza Martin acknowledges the tension between adding user controls and preserving the magical 'push button, get podcast' simplicity.
- **[1:00:58] The Future**
  - [1:01:44] Raiza Martin reveals that while audio overviews attract users, they stay for the core NotebookLM features like note-taking and question-answering.
  - [1:05:50] Some users have downloaded their GitHub repos into NotebookLM and generated audio overviews of their codebase.
  - [1:06:09] Raiza Martin observed a student use NotebookLM to check Computer Science homework, asking for guidance instead of answers to learn effectively.
  - [1:07:17] Raiza Martin confirms the team is actively prioritizing real-time interactive chat with the AI hosts, exploring what makes it fundamentally different from Q&A.
- **[1:08:27] Outro**
  - [1:08:52] Raiza Martin advises AI product managers to always be building, using competitor tools, and prototyping to gain conviction.
  - [1:10:07] Usama Bin Shafqat: 'The most magical moments out of AI building come about for me when I'm really close to the edge of the model capability.'
  - [1:12:11] Swyx declares NotebookLM was 'kind of like the ChatGPT moment for Google,' which Raiza Martin found 'crazy' and 'so cool.'

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **Raiza Martin** (guest)
- **Usama Bin Shafqat** (guest)

## Topics

Audio & Music, Evals, Language Models

## Mentioned

Databricks (company), DeepMind (company), Google (company), OpenAI (company), AI Test Kitchen (product), Artifacts (product), Canvas (product), ChatGPT (product), Claude (product), Deep Dive (product), Discord (product), Gemini (product), LaMDA (product), LinkedIn (product), NotebookLM (product), Project Tailwind (product), Spotify (product), Wikipedia (product), X (product), YouTube (product)

## Transcript

### Intro

Hey, everyone. We're here today as guests on Latent Space.

It's, uh, great to be here. I'm a longtime listener and fan. They've had some great guests on this show before.

Yeah. What an honor to have us, the hosts of another podcast, join as guests.

I mean, a huge thank you to Swyx and Alessio for the invite. Thanks for having us on the show.

Yeah, really. It seems like they brought us here to talk a little bit about our show, our podcast.

Yeah. I mean, we've had lots of listeners ourselves, listeners of Deep Dive.

Oh, yeah. We've made a ton of audio overviews since we launched, and we're learning a lot.

There's probably a lot we can share around what we're building next, huh?

Yeah. We'll share a little bit at least.

The short version is we'll keep learning and getting better for you.

We're glad you're along for the ride.

So yeah, keep listening.

Keep listening and stay curious. We promise to keep diving deep and, uh, bringing you even better options in the future.

Stay curious.

**Raiza Martin** [0:53]
Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host, Swyx, founder of Small AI.

**Swyx** [1:02]
Hey, and today we're back in the studio with our special guests, Raiza Martin and Usama—I forgot to s- to, to get your last name. Shafqat?

**Usama Bin Shafqat** [1:10]
Yes.

**Swyx** [1:11]
Okay. Welcome.

**Usama Bin Shafqat** [1:12]
Thank you.

**Raiza Martin** [1:13]
Hello. Thank you for having us.

**Usama Bin Shafqat** [1:14]
Thanks for having us.

**Swyx** [1:14]
So AI podcasters meet human podcasters. Um, always fun. Congrats on the success of NotebookLM. I mean, uh, how does it feel?

**Raiza Martin** [1:22]
Uh, it's been a lot of fun. A lot of it, honestly, was unexpected, but, uh, my favorite part is really listening to the audio overviews that people have been making.

**Swyx** [1:30]
Maybe we should do a little bit of intros and tell the story. You know, what, what is your path into the sort of Google AI org, or maybe-- I, actually I don't even know what org you guys are in.

**Raiza Martin** [1:40]
Um, I can start. My name's Raiza. I lead the NotebookLM team inside of Google Labs. So specifically that's the org that we're in. It's called Google Labs. It's only about two years old, and our whole mandate is really to build AI products.

That's it. We work super closely with DeepMind. Our entire thing is just, like, try a bunch of things and see what's landing with users. And the, the background that I have is really I worked in payments before this, and I worked in ads right before, and then startups.

I tell people, like, at every time that I change orgs, I actually almost quit Google.

**Swyx** [2:14]
Mm-hmm.

**Raiza Martin** [2:14]
Like, specifically, like, in between ads and payments, I was like, "All right, like, I can't do this. Like, this is, like, super hard." I was like, "It's not for me." I'm, like, a very zero to one person. But then I was like, "Okay, I'll try.

I'll interview with other teams." And when I interviewed in payments, I was like, "Oh, these people are really cool. I don't know if I'm, like, s-a super good fit with this space, but I'll try it 'cause the people are cool."

And then I really enjoyed that, and then I worked on, like, zero-to-one features inside of payments, and I had a lot of fun. But then the time came again where I was like, "Oh, I don't know." I was like, "It's time to leave.

It's time to start my own thing." But then I interviewed inside of Google Labs, and I was like, "Oh, darn." Like, there's definitely, like-

**Swyx** [2:48]
They got you again.

**Raiza Martin** [2:49]
They got me again. And so now, now I've been here for two years, and I'm, I'm happy that I stayed because especially with, you know, the recent success of NotebookLM, I'm like, "Dang, we did it."

**Swyx** [3:00]
Yeah.

**Raiza Martin** [3:00]
I actually got to do it.

**Swyx** [3:01]
Yeah.

**Raiza Martin** [3:01]
So that was really cool.

**Usama Bin Shafqat** [3:02]
Kind of similar, honestly. I was at a big team at Google. We do sort of the data center supply chain planning stuff. Um, Google has, like, the largest sort of footprint. Obviously, there's a lot of management stuff to do there.

But then there was this thing called Area 120 at Google, which does not exist anymore. But I sort of wanted to do, like, more zero-to-one building and landed a role there where we're trying to build, like, a creator commerce platform called Kaya.

It launched briefly, um, a couple years ago. But then Area 120 sort of transitioned and morphed into Labs, and, like, over the last few years, like, the, the focus just got a lot clearer. Like, we're trying to build new AI products and do it in the wild and sort of co-create and all of that.

So yeah, we've just been trying a bunch of different things, and this one really landed, which has felt pretty phenomenal.

**Swyx** [3:53]
Really, really landed. Let's talk about the brief history of NotebookLM. You had a tweet which is very helpful for doing research. May 2023 during Google I/O, you announced Project Tailwind.

**Raiza Martin** [4:04]
Yeah.

**Swyx** [4:04]
So that-- So today is October 2024. So you joined October 2022?

### Origins

**Raiza Martin** [4:09]
Actually, I used to lead AI Test Kitchen, and-

**Swyx** [4:12]
Ah

**Raiza Martin** [4:13]
... this was actually, I think, not I/O 2023, I/O 2022-

**Swyx** [4:19]
Okay

**Raiza Martin** [4:19]
... is when we launched AI Test Kitchen or announced it, and I don't know if you remember it-

**Swyx** [4:23]
I wasn't it. You-- That's how you, like, had the basic prototype for Gemini and-

**Raiza Martin** [4:26]
Yes, yes, exactly

**Swyx** [4:28]
... and, like, gave-

**Raiza Martin** [4:28]
LaMDA

**Swyx** [4:28]
... beta access to people.

**Raiza Martin** [4:29]
Right. Yeah, yeah. Yeah. And I remember I was like, "Wow, this is, this is crazy. I-- We're gonna launch an LLM into the wild." And that was the first project that I was working on at Google. But at the same time, my manager at the time, Josh, he was like, "Hey, but I want you to really think about, like, what real products would we build that are not just demos of the technology?"

That was in October of 2022. I was sitting next to an engineer that was working on a project called Talk to Small Corpus. His name was Adam. And the idea of Talk to Small Corpus is basically use an LLM to talk to your data.

And at the time I was like, "Wait, there are some, like, really practical things that you can build here." And I, I-- Just a little bit of background, like, I was an adult learner. Like, I went to college while I was working a full-time job.

**Swyx** [5:14]
Mm-hmm.

**Raiza Martin** [5:14]
And the first thing I thought was like, "This would have really helped me with my studying."

**Swyx** [5:19]
Mm-hmm.

**Raiza Martin** [5:19]
Right? Like, if I could just, like, talk to a textbook, especially, like, when I was tired after work, like, that would've been huge. We took a lot of, like, the Talk to Small Corpus prototypes, and I showed it to a lot of, like, college students, particularly, like, adult learners.

They were like, "Yes." Like, "I get it." Right? Like, I didn't even have to explain it to them. And we just continued to iterate the prototype from there to the point where we actually got a slot as part of the, the I/O demo in '23.

**Swyx** [5:43]
And Corpus, was it a textbook?

**Raiza Martin** [5:45]
Oh my gosh, yeah.

**Swyx** [5:46]
Yeah.

**Raiza Martin** [5:46]
It's funny. Actually, y- when he explained the project to me, he was like, "Talk to Small Corpus." It was like, "Talk to a small corpse?"

**Swyx** [5:51]
Yeah, nobody says corpus. Yeah.

**Raiza Martin** [5:52]
Talk to a small corpse?

**Swyx** [5:53]
Yeah, yeah.

**Raiza Martin** [5:53]
This is not AI.

**Swyx** [5:54]
It's very academic.

**Raiza Martin** [5:55]
Yeah, yeah. And it, it really was just, like, the, a way for us to describe the amount of data that we thought, like, could be-- it could be good for.

**Swyx** [6:02]
Yeah.

**Raiza Martin** [6:03]
And so-

**Swyx** [6:03]
But even then you're still, like, doing RAG stuff because-

**Raiza Martin** [6:05]
Yeah, yeah

**Swyx** [6:05]
... you know, the context lengths back then was probably like two K, four K.

**Raiza Martin** [6:08]
Yeah, it was basically RAG.

**Swyx** [6:10]
Yeah.

**Raiza Martin** [6:10]
That, that was essentially what it was, and I remember I was like, we were building the prototypes, and at the same time, I think, like, the rest of the world was, right? We were seeing all of these, like, chat with PDF stuff come up, and I was like, "Come on, we gotta go."

Like, we have to, like, push this out into the world. I think if there was anything I wish we would've launched sooner because I wanted to learn faster. But I think, like, we netted out pretty well. Was the initial product just text-to-speech, or were you also doing kind of like a synthesizing of the content, refining it, or were you just helping people read through it?

Before we did the I/O announcement in '23, we'd already done a lot of studies, and one of the first things that I realized was the first thing anybody ever typed was, "Summarize the thing." Right? Summarize the document. And it was, like, half, like, a test and half just like, "Oh, I know the content.

I want to see how well it does this." So as part of the first thing that we launched, um, it was called Project Tailwind back then. It was just Q&A, so you could chat with the doc- Just through text, and it would automatically generate a summary as well.

I'm not sure if we had it back then. I think we did. It would also generate the key topics in your document, and y- it- it could support up to, like, 10 documents. So it wasn't just, like, a single doc.

**Swyx** [7:20]
And then the IO demo went well, I guess.

**Raiza Martin** [7:23]
Yeah.

**Swyx** [7:23]
And then what was the discussion from there to where we are today? Is there any maybe intermediate step of the product that people missed, uh, between this first launch or...?

**Raiza Martin** [7:34]
It was interesting because every step of the way, I think we ha- we hit, like, some pretty critical milestones. So I think from the initial demo, I think there was so much excitement of like, "Wow, what is this thing that Google is launching?"

And so we capitalized on that. We built the wait list. That's actually when we also launched the Discord server, which has been huge for us because for us in particular, one of the things that, that I really wanted to do was to be able to launch features and get feedback ASAP.

Like, the moment somebody tries it, like, I wanna hear what they think right now, and I wanna ask follow-up questions, and the Discord has just been so great for that. But then we basically took the feedback from IO.

We continued to refine the product. So we added more features. We added sort of like the ability to save notes, write notes. We generate follow-up questions. So there's a bunch of stuff in the product that shows, like, a lot of that research, but it was really the rolling out of things like we removed the wait list, so rolled out to all of the United States.

We rolled out to, uh, over 200 countries and territories. We started supporting more languages, both in the UI and, like, the actual source stuff. We experienced... Like, in terms of milestones, there was, like, an explosion of, like, users in Japan.

This was super interesting as, like, in terms of just, like, unexpected. Like, people would write to us, and they would be like, "This is amazing. I have to read all of these rules in English, but I can chat in Japanese."

**Swyx** [8:52]
Mm.

**Raiza Martin** [8:53]
I was like, "Oh, wow, that's true," right? Like, with LLMs, you kind of get this natural... It translates-

**Swyx** [8:58]
Mm

**Raiza Martin** [8:58]
... the content for you, and you can ask in your sort of preferred mode, and I think that's not just, like, a language thing, too. I think there's like, um... I do this test with Wealth of Nations all the time 'cause it's, like, a pretty complicated text to read, but I can just-

**Swyx** [9:10]
The Adam Smith classic. It's, like, 400 pages or something.

**Raiza Martin** [9:13]
Yeah, but I, I like this test 'cause I'm like, I ask in, like, normie, you know, plain speak, and then it, it summarizes really well for me. It sort of adapts to my tone. Yeah.

**Swyx** [9:23]
Very capitalist. Uh

**Raiza Martin** [9:25]
Very on brand.

**Swyx** [9:26]
Um, I just checked in on a NotebookLM Discord, 65,000 people.

### Discord

**Raiza Martin** [9:29]
Yeah.

**Swyx** [9:29]
Crazy.

**Raiza Martin** [9:30]
Yeah, crazy.

**Swyx** [9:30]
Just, like, for, for one project within Google, and it's not, like, it's not labs. It's just, just NotebookLM.

**Raiza Martin** [9:36]
Just NotebookLM.

**Swyx** [9:37]
What do you learn from, um, the community?

**Raiza Martin** [9:39]
I think that the Discord is really great for hearing about a couple of things. One, when things are going wrong. I think honestly, like, our fastest way that we've been able to find out if, like, the servers are down or there's just an influx of people being like, "It says system unable to answer.

Anybody else getting this?" And I'm like, "All right, let's go." And it actually catches it a lot faster than, like, our own monitoring does. It's like that's been really cool, so thank you.

**Swyx** [10:03]
Cancel data dog.

**Raiza Martin** [10:06]
So, so thank you to everybody. Please keep reporting it. I think the second thing is really the use cases. I think when we put it out there, I was like, "Hey, I have a hunch of how people w- will use it," but, like, to actually hear about, you know, not just the context of, like, the use of NotebookLM, but, like, what is this person's life like?

Why do they care about using this tool? Especially people who actually have trouble using it, but they keep pushing, like, that's just so critical to understand what was so motivating, right? Like, what, what was your problem that was, like, so worth solving?

So that's, like, a second thing. The third thing is also just hearing sort of like when we have wins and when we don't have wins because there's actually a lot of functionality where I'm like, "Hmm, I don't know if that landed super well or if that was actually super critical."

As part of having this s- sort of small project, right, I wanna be able to unlaunch things, too. So it's not just about just, like, rolling things out and testing it-

**Swyx** [10:54]
Mm

**Raiza Martin** [10:54]
... and being like, "Wow, now we have, like, 99 features." Like, hopefully we get to a place where it's like there's just a really strong core feature set, and the things that aren't as great we can just unlaunch.

**Swyx** [11:03]
What, what have you unlaunched? I have to ask.

**Raiza Martin** [11:05]
I'm in the process of unlaunching some stuff, but for example, uh, we had this, this idea that you could highlight the text in your source passage, and then you could transform it, and nobody was really using it, and it was, like, a very complicated piece of our architecture, and it's very hard to continue supporting it in the context of new features.

So we were like, "Okay, let's do a 50/50 sunset of this thing and see if anybody complains," and so far nobody has.

**Swyx** [11:30]
Is there, like, a feature flagging paradigm inside of your architecture that lets you feature flag these things easily? Um-

**Raiza Martin** [11:37]
Yes, and actually-

**Swyx** [11:37]
What is it called? Like, I, I, I love feature flagging.

**Raiza Martin** [11:40]
You mean, like, in terms of just, like, being able to expose things to users?

**Swyx** [11:42]
Yeah, as a PM, like, this is your number one tool, right?

**Raiza Martin** [11:44]
Yeah, yeah.

**Swyx** [11:45]
Like, let's try this out. All right. If it works, roll it out. If it doesn't, roll it back, you know?

**Raiza Martin** [11:49]
Yeah, I mean, we just run Mendel experiments for the most part, and I, I actually... I don't know if you saw it, but on Twitter, somebody was able to get around our flags, and they enabled all the experiments.

They were like, "Check out what the NotebookLM team is cooking." And I was like, " " And I was at, at lunch with the rest of the team, and I was like... I was eating, I was like, "Guys, guys"- "...

Magic draft me." And they were like, "Oh, no." I was like, "Okay, just finish eating, and then let's go figure out what to do."

### Architecture

**Swyx** [12:16]
Yeah.

**Usama Bin Shafqat** [12:16]
I think a postmortem would be fun, but I don't think we need to do it on-

**Raiza Martin** [12:19]
Yeah.

**Usama Bin Shafqat** [12:20]
... on the, on the podcast now.

**Raiza Martin** [12:21]
Yeah, yeah.

**Usama Bin Shafqat** [12:22]
Can we just talk about what's behind the magic? So I, I think everybody has questions, hypothesis about what models power it. I know you might not be able to share everything, but can you just give people a very basic how do you take the data and put it in the model, what text model you use, what's the text-to-speech, kinda like jump between the two?

**Raiza Martin** [12:42]
Like-

**Swyx** [12:42]
Sure, yeah

**Raiza Martin** [12:43]
... I was gonna say, S-

**Usama Bin Shafqat** [12:44]
Go for it

**Raiza Martin** [12:44]
... Usama, he manually does all the podcasts.

**Usama Bin Shafqat** [12:46]
Yes.

**Swyx** [12:46]
Oh, thank you.

**Usama Bin Shafqat** [12:47]
Really fast.

**Swyx** [12:47]
You're very fa- yeah.

**Raiza Martin** [12:48]
He's both of the voices at once.

**Swyx** [12:51]
Voice actor.

**Raiza Martin** [12:52]
All right. Go ahead, go ahead.

**Usama Bin Shafqat** [12:53]
Yeah, so, um, for a bit of background, we were building this thing sort of outside NotebookLM to begin with, like, just the idea is, like, content transformation, right? Like, we can do different modalities. Like, everyone knows that. Everyone's been poking at it, but, like, how do you make it really useful?

And, like, one of the ways we thought was like, okay, like, you maybe like, you know, people learn better when they're hearing things, but TTS exists and you can, like, narrate- ... whatever's on screen, but you, you won't absorb it the same way.

So, like, that's where we sort of started out into the realm of, like, maybe we try, like, you know, two people are having a conversation-

**Alessio** [13:27]
Mm-hmm

**Usama Bin Shafqat** [13:27]
... kind of format. We didn't actually start out thinking this would live in Notebook, el- right? Like, Notebook was sort of... We built this demo out independently, tried out, like, a few different sort of sources. The, the main idea was, like, go from some sort of sources and transform it into a listenable, engaging audio format.

And then through that process, we, like, unlocked a bunch more sort of learnings. Like, for example, in a sense, like, you're not prompting the model as much because, like, the information density is getting unrolled by the model prompting itself, in a sense.

Because there's two speakers, and they're both technically like AI personas, right, that have different angles of looking at things, and, like, they'll have a discussion about it, and that sort of... We realized that's kind of what was making it riveting in a sense.

Like, you care about what comes next, even if you've read the material already. 'Cause, like, some-- like, people say they get new insights on their own journals or books or whatever, like anything that they've written themselves.

**Alessio** [14:26]
Mm.

**Usama Bin Shafqat** [14:26]
So yeah, from a modeling perspective, like, it's, uh, like Raiza said earlier, like we work with the Deep Mind audio folks pretty closely, so they're, they're always cooking up new techniques to, like, get better, more human-like audio. And then Gemini 1.5 is really, really good at absorbing long context, so we sort of, like, generally put those things together in a way that we could reliably produce the audio.

**Raiza Martin** [14:53]
I would add, like, um, there's something really nuanced, I think, about sort of the evolution of, like, the, the utility of text-to-speech, where if it's just reading an actual text response, and I've done this several times, I do it all the time with, like, reading my text messages, or, like, sometimes I'm trying to read, like, a really dense paper, but I'm trying to do actual work.

I'll have it, like, read out the screen. There is something really robotic about it-

**Alessio** [15:17]
Mm

**Raiza Martin** [15:17]
... that is not engaging, and it's really hard to consume content in that way, and it's never been really effective. Like, particularly for me, where I'm like, hey, it's actually just, like, it's fine for, like, short stuff like texting, but even that, it's, like, not that great.

**Alessio** [15:30]
Mm.

**Raiza Martin** [15:30]
So I think the frontier of experimentation here was really thinking about there is a transform that needs to happen in between, whatever, here's, like, my resume, right? Or here's, like, a, a hundred-page slide deck or something. There is a transform that needs to happen that is inherently editorial, and I think this is where, like, that two-person persona, right, dialogue model, they have takes on the material that you've presented.

**Alessio** [15:56]
Yeah.

**Raiza Martin** [15:57]
That's where it really sort of, like, brings the content to life in a way that's, like, not robotic.

**Alessio** [16:02]
Mm.

**Raiza Martin** [16:02]
And I think that's, like, where the magic is, is, like, you don't actually know what's going to happen when you press generate, you know, for better or for worse. Like, to the extent that, like, people are like, "No, I actually want it to be more predictable now."

Like, I, I want to be able to tell them, but I think that initial, like, wow was because you didn't know.

**Alessio** [16:19]
Yeah.

**Raiza Martin** [16:19]
Right? When you upload your resume, like, what's it about to say about you? And I think I've seen enough of these where I'm like, oh, it gave you good vibes, right? Like, you knew it was gonna say, like, something really cool.

As we start to shape this product, I think we wanna try to preserve as much of that wow as much as we can. Because I do think, like, exposing, like, all the knobs and, like, the, the dials, like, we, we've been thinking about this a lot.

It's like, hey, is that, like, the actual thing?

**Alessio** [16:44]
Mm.

**Raiza Martin** [16:44]
Is that the thing that people really want? So...

**Alessio** [16:46]
Have you found differences in having one model just generate the conversation and then using text-to-speech to kind of fake two people? Or, like, are you actually using two different kind of system prompts to, like, have a conversation step by step?

I'm always curious, like, if persona system prompts make a big difference, or, like, you just put in one prompt, and then you just let it run.

**Usama Bin Shafqat** [17:06]
I guess, like, generally, we use a lot of inference, as you can tell with, like, the, the spinning thing takes-

**Alessio** [17:11]
Mm-hmm

**Usama Bin Shafqat** [17:12]
... takes a while. Um, so yeah, there's definitely, like, a bunch of different things happening under the hood. We've tried both approaches, and they have their sort of drawbacks and benefits. I think that that idea of, like, questioning, like, the two different personas, like, persists throughout, like, whatever approach we try.

It's like there's a bit of, like, imperfection in there. Uh, like, we had to really lean into the fact that, like, to build something that's engaging, like, it needs to be somewhat human-

**Alessio** [17:38]
Yeah

**Usama Bin Shafqat** [17:38]
... and it needs to be just not a chatbot. Like, that was sort of, like, what we need to diverge from is, like, you know, most chatbots will just n-narrate the, the same kind of answer, like, given the same sources for the most part, which is ridiculous.

So yeah, there's, like, experimentation there under the hood, like, with the model to, like, make sure that it's spitting out, like, different takes and different personas and different sort of prompting each other is, like, a good analogy, I guess.

### Visionary

**Alessio** [18:01]
Yeah.

**Raiza Martin** [18:01]
Yeah.

**Alessio** [18:01]
I think, um, Steven Johnson, I think he's on your team. I don't know what his role is. He seems like chief dreamer, writer.

**Raiza Martin** [18:08]
Yeah, I mean, uh, I can comment on Steven.

**Alessio** [18:10]
Yeah.

**Raiza Martin** [18:11]
So Steven joined actually in, in the very early days, I think, before it was even a fully funded project.

**Alessio** [18:16]
Yeah.

**Raiza Martin** [18:16]
And I remember when he joined, I was like, "Steven Johnson's gonna be on my team?"

**Alessio** [18:21]
Mm-hmm.

**Raiza Martin** [18:21]
You know, and, and for folks who don't know him, Steven is a New York Times best-selling author of, like, 14 books. He has a PBS show. He's, like, incredibly smart, just, like, a true sort of celebrity by himself.

**Alessio** [18:34]
Yeah.

**Raiza Martin** [18:34]
And then he joined Google, and he was like, "I wanna come here, and I wanna build the thing that I've always dreamed of, which is a tool to help me think." I was like, "A what?" Like, "A tool to help you think?"

I was like, "W-what do you need help with?" "Like, you seem to be doing great on your own." And, uh, you know, he would describe this to me, and I would watch his flow, and aside from, like, providing a lot of inspiration, to be honest, like, when I watched Steven work, I was like, "Oh, nobody works like this," right?

Like, this is what makes him special. Like, he is such a, such a dedicated, like, researcher and journalist, and he's so thorough. He's so smart. And then I had this realization of, like, maybe Steven is the product. Maybe the work-

**Swyx** [19:15]
Mm

**Raiza Martin** [19:15]
... is to take Steven's expertise and bring it to like everyday people that could really benefit from this. Like just watching him work, I was like, "Oh, I could definitely use like a mini Steven like doing work for me."

Like that would make me a better PM. And then I thought very quickly about like the adjacent roles that could use sort of this like research a- and analysis tool. And so aside from being, you know, chief dreamer, Steven also represents like a super workflow that I think all of us, like if we had a- access to a tool like it, would just inherently like make us better.

**Swyx** [19:47]
Did you make him express his thoughts while he worked or you just silently watched him? Or how, how did, how does this work?

**Raiza Martin** [19:53]
Oh, now, now you're making me admit it. But yes, I did just silently watch him.

**Swyx** [19:57]
Yeah. This is a part of the PM toolkit, right? Like you-

**Raiza Martin** [19:59]
Yeah, yeah

**Swyx** [19:59]
... user interviews and all that.

**Raiza Martin** [20:01]
Yeah, I mean, I did interview him, but I noticed like if I interviewed him-

**Swyx** [20:05]
Yeah

**Raiza Martin** [20:05]
... it was different than if I just-

**Swyx** [20:06]
It's artificial

**Raiza Martin** [20:06]
... if I just watched him. And I did the same thing with students all the time. Like I followed a lot of students around. I watched them study. I would ask them like, "Oh, how do you feel now?"

Right? Or, "Why did you do that?" Like, "What made you, what made you do that actually?" Or, "Why are you upset about like this particular thing? Why are you cranky about this particular topic?" Um, and it was very similar I think for Steven especially because he was describing, he was in the middle of writing a book and he would describe like, "Oh, you know, here's how I research things and here's how I keep my notes.

Oh, and here's how I do it." And it was really, he was doing this sort of like self-questioning, right? Like now we talk about like chain of, you know, reasoning or-

**Swyx** [20:43]
Reflection

**Raiza Martin** [20:44]
... thought.

**Swyx** [20:44]
Yeah.

**Raiza Martin** [20:44]
Reflection. And I was like, "Oh, he's the OG." Like I watched him do it in real time and I was like, "That's, that's like LLM right there." And to be able to bring sort of that expertise in a way that was like, you know, maybe like costly inference-wise, but really have like that ability inside of a tool that was like, for starters, free inside of NotebookLM, um, it was good to learn whether or not people really did find use out of it.

**Swyx** [21:05]
So did he just commit to using NotebookLM for everything or-

**Raiza Martin** [21:09]
Oh my gosh

**Swyx** [21:09]
... did you just model his existing workflow?

**Raiza Martin** [21:12]
Both, right?

**Swyx** [21:13]
Okay.

**Raiza Martin** [21:13]
Like in the beginning there was no product for him to use.

**Swyx** [21:15]
Yeah.

**Raiza Martin** [21:15]
And so he just kept describing the thing that he wanted. And then eventually like we started building the thing and then I would start watching him use it. One of the things that I love about Steven is he uses the product in ways where it kind of does it but doesn't quite.

Like he's always using it at like the absolute max limit of this thing, but the way that he describes it is so full of promise where he's like, "I can see it going here." And all I have to do is sort of like meet him there and sort of pressure test whether or not, you know, everyday people want it and we just have to build it.

**Swyx** [21:47]
I would say, uh, OpenAI has a pretty similar person, Andrew Mason I think his name is. It's very similar, like just from the writing world and using it as a tool for thought-

**Raiza Martin** [21:55]
Yeah

**Swyx** [21:55]
... to shape ChatGPT.

**Raiza Martin** [21:57]
Yeah.

**Swyx** [21:57]
I don't think that people who use AI tools to their limit are common.

**Raiza Martin** [22:00]
Yeah.

**Swyx** [22:01]
I'm looking at my NotebookLM now. I've got two sources. You have a little like source limit thing, uh, and my bar is over here-

**Raiza Martin** [22:07]
Yeah.

**Swyx** [22:07]
... you know, and it stretches across the whole thing.

**Raiza Martin** [22:09]
No.

**Swyx** [22:09]
I'm like, "Did he fill it up?" Like what-

**Raiza Martin** [22:10]
Yes

**Swyx** [22:10]
... you know?

**Raiza Martin** [22:11]
And he has like a higher limit than others. I, I think Steven-

**Swyx** [22:14]
He fills it up.

**Raiza Martin** [22:15]
Oh, yeah. Like I don't-

**Swyx** [22:15]
Okay

**Raiza Martin** [22:15]
... think Steven even has a limit actually.

**Swyx** [22:18]
And he has notes, Google Drive stuff, PDFs, MP3-

**Raiza Martin** [22:21]
Yes

**Swyx** [22:22]
... whatever.

**Raiza Martin** [22:22]
Yes. And one of my favorite demos, he just did this recently, is he has actually PDFs of like handwritten Marie Curie notes.

**Swyx** [22:29]
I see. So you're doing image recognition as well-

**Raiza Martin** [22:31]
Yeah, yeah

**Swyx** [22:31]
... on this thing.

**Raiza Martin** [22:32]
So it, it does support it today. So if you have a PDF that's purely images it will recognize it, but his demo is just like super powerful. He's like, "Okay, here's Marie Curie's notes." And then it's like, "Here's how I'm using it to analyze it and I'm using it for like this thing that I'm writing."

And that's really compelling. It's like the everyday person doesn't think of these applications. And I think even like when I listen to Steven's demo I see the gap. I see how Steven got there, but I don't see how I could without him.

**Swyx** [22:56]
Mm.

**Raiza Martin** [22:56]
And so there's a lot of work still for us to build of like, hey, how do I bring that magic down to like zero work? Because I look at all the steps that he had to take in order to do it and I'm like, "Okay, that's, that's product work for us," right?

Like that's just onboarding.

### Data Pipeline

**Swyx** [23:09]
And so from an engineering perspective people come to you and it's like, "Hey, I need to use this handwritten notes from Marie Curie from hundreds of years ago." How do you think about adding support for like data sources and then maybe any fun stories in like supporting more esoteric types of, of inputs?

**Raiza Martin** [23:26]
So I think about the product in three ways, right? So there's the sources, the source input. There's like the capabilities of like what you could do with those sources. And then there's the third space which is how do you output it into the world?

Like how do you put it back-

**Swyx** [23:38]
Mm

**Raiza Martin** [23:38]
... out there? There's a lot of really basic sources that we don't support still, right? I think there's sort of like the handwritten note stuff is, is one, but even basic things like DOCX or like PowerPoint, right? Like these are the things that people, everyday people are like, "Hey, my professor actually gave me everything in DOCX.

Can you support that?" And then just like basic stuff like images and PDFs combined with text. Like there's just a, a really long roadmap for sources that I think we just have to work on. So that's like a big piece of it.

On the output side, and I think this is like one of the most interesting things that we learned really early on is, sure, there's like the Q&A analysis stuff which is like, "Hey, when did this thing launch? Okay, you found it in the slide deck.

Here's the answer." But most of the time the reason why people ask those questions is because they're trying to make something new. And so when actually when some of those early features leaked, like a lot of the features we're experimenting with are the output types.

And so you can imagine that people care a lot about the, the sources that they're putting into NotebookLM 'cause they're trying to create something new. So I think equally as important as like the source inputs are the, the outputs that we're helping people to create.

And really like, you know, shortly on the roadmap we're thinking about how do we help people use NotebookLM to distribute knowledge? You know, that's like one of the most compelling use cases is like shared notebooks as like a way to share knowledge.

How do we help people take sources and then like one click new documents out of it, right? And I think that's something that people think is like, oh yeah, of course, right? Like one push a document, but what does it mean to do it right?

Like to do it in your style, in your brand, right? To follow your guidelines, stuff like that. So I think there's, there's a lot of work like on both sides. Of that equation.

**Swyx** [25:13]
Interesting. Any comments on the engineering side of things?

**Usama Bin Shafqat** [25:16]
Uh, so yeah, li- like I said, I was mostly f- uh, working on building the text to audio, which kind of lives as a separate engineering pipeline almost that we then put into NotebookLM, but I think there's probably tons of NotebookLM engineering war stories on dealing with sources-

**Raiza Martin** [25:30]
Yeah

**Swyx** [25:30]
Mm.

**Usama Bin Shafqat** [25:30]
And so-

**Raiza Martin** [25:30]
Yeah

**Usama Bin Shafqat** [25:30]
... so I don't work too closely with engineers directly, but I think a lot of it does come down to, like, Gemini's native understanding of images really well with the latest generations.

**Raiza Martin** [25:39]
Yeah, I think on the engineering and modeling side, I think we are a really good example of a team that's put a product out there, and we're getting a lot of feedback from the users, and we return the data to the modeling team, right?

To the extent that we say, "Hey, actually, you know what people are uploading but we can't really support super well?" Text plus image, right? Especially to the extent that, like, NotebookLM can handle up to 50 sources, 500,000 words each.

Like, you're not gonna be able to jam all of that into, like, the context window, so how do we do multimodal embeddings with that? There's really, like, a lot of things that we have to solve that are almost there but not quite there yet.

**Swyx** [26:16]
On then turning it into audio, I think one of the best things is it has so many of the human... Does that happen in the text generation that then becomes audio, or is that a part of, like, the audio model that transforms the text?

**Usama Bin Shafqat** [26:28]
It's a bit of both, I would say. The audio models definitely try to mimic, like, certain human intonations and, like, sort of natural, like, you know, breathing and-

**Swyx** [26:36]
Mm-hmm

**Usama Bin Shafqat** [26:36]
... pauses and, like, laughter and things like that. But yeah, in generating, like, the, the text, we also have to sort of give signals on, like, where those things maybe would make sense.

**Swyx** [26:45]
Yeah, and on the input side instead, having a transcript versus having the audio, like, can you take some of the emotions out of it, too? If I'm giving... Like, for example, when we did the recaps of our podcast, we can either give a audio of the pod or we can give a diarized transcription of it.

**Raiza Martin** [27:01]
Mm. Yeah.

**Swyx** [27:01]
But, like, the transcription doesn't have some of the, you know, voice kinda, like, things.

**Raiza Martin** [27:05]
Yeah, yeah.

**Swyx** [27:06]
Um, do you reconstruct that when people upload audio, or how does that work?

**Raiza Martin** [27:10]
So when you upload audio today, we just transcribe it, so it, it is quite lossy in the sense that, like, we don't transcribe, like, the emotion from that as a source. But, um, when you do upload a text file and it has a lot of, like, that annotation, I think that there is some ability for it to be reused in, like, the audio output, right?

But I think it will still contextualize it in the Deep Dive format.

**Swyx** [27:33]
Mm.

**Raiza Martin** [27:33]
So I think that's, that's something that's, like, particularly important is like, hey, today we only have one format. It's Deep Dive. It's meant to be pretty general overview, and it is pretty peppy. Like, it's just very upbeat. So, uh-

**Usama Bin Shafqat** [27:44]
It's very enthusiastic, yeah.

**Raiza Martin** [27:45]
Yeah, yeah. Even, even if you had, like, a sad topic, I think they would find a way to be like, "Silver lining, though."

**Swyx** [27:51]
Really?

**Raiza Martin** [27:51]
Yeah. We're having a good chat.

**Swyx** [27:54]
Yeah, that's awesome. One of the ways, many, many, many ways that Deep Dive went viral is people saying, like, "If you wanna feel good about yourself, just drop in your LinkedIn."

**Usama Bin Shafqat** [28:02]
Yeah.

**Swyx** [28:02]
Any other, like, favorite use cases that you saw from, from people discovering things in social media?

**Raiza Martin** [28:08]
I mean, there's so many funny ones, and I love the funny ones, I think because I'm always relieved when I watch them. I'm like "That was funny and not scary. It's great." There was another one that was interesting, which was a startup founder putting their landing page and being like, "All right, let's test whether or not-"

**Swyx** [28:24]
Mm

**Raiza Martin** [28:24]
... "like, the value prop is coming through." And I was like, "Wow, that's right. That's smart."

**Swyx** [28:28]
Yeah.

**Raiza Martin** [28:28]
Uh, and then I saw a couple of other people following, following up on that, too.

**Usama Bin Shafqat** [28:32]
Yeah.

**Swyx** [28:32]
Yeah, yeah. I put my About page in there and, like, uh, yeah, like if there are things that I'm not comfortable with, I should remove it, you know, so that-

**Raiza Martin** [28:38]
Yeah

**Swyx** [28:38]
... it, it can pick it up.

**Usama Bin Shafqat** [28:39]
Right.

**Raiza Martin** [28:39]
Yeah.

**Usama Bin Shafqat** [28:40]
I think that the personal hype machine was, like- ... a pretty viral, uh, one. The, I think, like, people uploaded their dreams and dr-

**Raiza Martin** [28:47]
Wow

**Usama Bin Shafqat** [28:47]
... like some people, like, keep sort of dream journals-

**Raiza Martin** [28:49]
Yeah

**Usama Bin Shafqat** [28:49]
... and it, like, would sort of comment on those and, like, it was therapeutic.

**Raiza Martin** [28:54]
I didn't see those. Those are good.

**Usama Bin Shafqat** [28:55]
Yeah.

**Raiza Martin** [28:55]
I, um, I hear from Googlers all the time, especially 'cause we launched it internally first, and it-- I think we launched it during the, you know, the Q3 sort of, like, check-in cycle. So all Googlers have to write notes about like, "Hey, you know, what'd you do in Q3?"

And what Googlers were doing is they would write, you know, whatever they accomplished in Q3, and then they would create an audio overview. And these people that I didn't know would just ping me and be like, "Wow," like, "I feel really good, like, going into a meeting with my manager."

And I was like, "Good, good, good, good. You really did that, right?" Like

**Usama Bin Shafqat** [29:29]
I think another cool one is just, like, any Wikipedia article.

**Raiza Martin** [29:33]
Yeah.

**Usama Bin Shafqat** [29:33]
Like, you drop it in-

**Swyx** [29:34]
Right. Okay

**Usama Bin Shafqat** [29:34]
... and it's just, like, suddenly, like, the best sort of summary overview. Um-

**Raiza Martin** [29:38]
Well, I think Kar-

**Usama Bin Shafqat** [29:39]
Right there

**Raiza Martin** [29:39]
... that's what Karpathy did, right? Like, he has now-

**Usama Bin Shafqat** [29:41]
Right

**Raiza Martin** [29:41]
... a Spotify channel called Histories of Mysteries, which is basically like he just takes, like, interesting stuff from Wikipedia and made audio overviews out of it.

### Evals

**Swyx** [29:50]
Yeah.

**Usama Bin Shafqat** [29:51]
Yeah, he became a podcaster overnight.

**Raiza Martin** [29:52]
Yeah.

**Swyx** [29:53]
Yeah.

**Raiza Martin** [29:53]
I'm, I'm here for it. I, I fully support him. I'm racking up the listens for him.

**Swyx** [29:59]
Honestly, it's useful even without the audio. You know, I, I feel like the audio does add an el- element to it, but, um, I always want, you know, paired audio and text, and it's just amazing to see what people are organically discovering.

I feel like it's because you laid the groundwork with NotebookLM, and then you, you came in and added the, the, the sort of TTS portion and made it so good, so human, which is weird. Like, it's, it's this engineering process of humans.

Oh, one thing I wanted to ask, um, do you have evals?

**Raiza Martin** [30:24]
Yeah.

**Usama Bin Shafqat** [30:24]
Yes.

**Swyx** [30:25]
What?

**Raiza Martin** [30:25]
Potatoes for chefs.

**Swyx** [30:28]
What is that? What do- what do you mean potatoes?

**Usama Bin Shafqat** [30:30]
Oh, sorry. Sorry.

**Swyx** [30:30]
Okay.

**Usama Bin Shafqat** [30:30]
We were joking.

**Swyx** [30:31]
This is in joke. Yeah, yeah.

**Usama Bin Shafqat** [30:31]
We were joking about this, like, a couple weeks ago. We were doing, like, side-by-sides, but, like, Usama sent me the file, and it was literally called Potatoes for Chefs. And I was like, "You know, my job is really serious, but like-"

**Swyx** [30:43]
That's kind of funny

**Usama Bin Shafqat** [30:43]
... you have to laugh a little bit at, like, the title of the file is, like, Potatoes for Chefs.

**Raiza Martin** [30:48]
Mm.

**Swyx** [30:48]
Is it, like, a training document for chefs?

**Usama Bin Shafqat** [30:50]
It's just a side, side-by-side-

**Raiza Martin** [30:52]
Yeah, side by sides

**Usama Bin Shafqat** [30:52]
... uh, for, like, two different kind of audio transcripts.

**Swyx** [30:55]
The question is really, like, as you iterate, the typical engineering advice is you establish some kind of tests, you have a-- or a benchmark. You're at, like, 30%, you wanna get it up to 90, right?

**Raiza Martin** [31:05]
Yeah.

**Swyx** [31:06]
What does that look like for making something sound human and interesting and voice?

**Usama Bin Shafqat** [31:11]
We have the sort of formal eval process as well, but I think, like, for this particular project, we maybe took a slightly different route to begin with. Like, there was a lot of just within the team listening sessions, a lot of, like, sort of like-

**Swyx** [31:23]
Dogfooding.

**Usama Bin Shafqat** [31:24]
Yeah, like, I think we-- the bar that we Tried to get to before even starting formal evals with raters and everything was much higher than I think other projects would. Like, 'cause that's, as you said, like, the traditional advice, right?

Like, get that ASAP. Like, what are you looking to improve on?

**Raiza Martin** [31:40]
Mm-hmm.

**Usama Bin Shafqat** [31:40]
Whatever benchmark it is. So there was a lot of just, like, critical listening, and I think a lot of making sure that those improvements actually could go into the model and, like, we're happy with that human element of it.

And then eventually we had to obviously distill those down into an eval set. But, like, still there's, like, the team is just, like, a very, very, like, avid user of the product-

**Raiza Martin** [32:02]
Yeah

**Usama Bin Shafqat** [32:02]
... at all stages.

**Raiza Martin** [32:03]
I think you just have to be really opinionated. I think that sometimes if you are, your intuition is just sharper, and you can move a lot faster on the product because it's like if you hold that bar high, right?

Like, like if you think about, like, the iterative cycle, it's like, hey, we could take, like, six months to ship this thing to get it to, like, mid where we were, or we could just, like, listen to this and be like, "Yeah, that's not it," right?

And I don't need a rater to tell me that. That's my preference, right? And collectively, like, if I have two other people listen to it, they'll probably agree. And it's just kind of this step of, like, just keep improving it to the point where you're like, "Okay, now, I think this is really impressive," and then, like, do evals.

**Usama Bin Shafqat** [32:42]
Mm-hmm.

**Raiza Martin** [32:42]
Right? And then validate that. Was the sound model done and frozen before you started doing all this, or are you also saying, "Hey, we need to improve the sound model as well"?

**Usama Bin Shafqat** [32:51]
Both.

**Raiza Martin** [32:51]
Oh, okay.

**Usama Bin Shafqat** [32:52]
Uh, yeah. We were impr- making improvements on the audio-

**Raiza Martin** [32:54]
Audio one

**Usama Bin Shafqat** [32:54]
... and, and just, like, generating those a- the, the transcript-

**Raiza Martin** [32:58]
Yeah

**Usama Bin Shafqat** [32:58]
... as well. Um, I think another weird thing here was, like, we need it to be entertaining.

**Raiza Martin** [33:02]
Yeah.

**Usama Bin Shafqat** [33:03]
And that's much harder to quantify than some of the other-

**Raiza Martin** [33:06]
Yeah

**Usama Bin Shafqat** [33:06]
... benchmarks that you can make for, like, you know, SWE-bench or, or-

**Raiza Martin** [33:09]
Mm-hmm

**Usama Bin Shafqat** [33:09]
... like, get better at this math-

**Raiza Martin** [33:11]
Yeah

**Usama Bin Shafqat** [33:11]
... prompt.

**Raiza Martin** [33:11]
Do you just have people rate one to five or, you know, or just up, thumbs up and down?

**Usama Bin Shafqat** [33:14]
For the formal rater evals, we have sort of, like, a Likert scale and, like-

**Raiza Martin** [33:18]
Yeah

**Usama Bin Shafqat** [33:18]
... a, a bunch of different dimensions there, but, um, we had to sort of break down the what makes it entertaining into, like, a bunch of different factors. But I think the team stage of that was more critical.

It was like, we need to make sure that, like, what is making it fun and engaging, like, we dial that as far as it goes. And while we're making other changes that are necessary, like obviously they shouldn't make stuff up or, you know, be-

**Raiza Martin** [33:40]
Hallucinations

**Usama Bin Shafqat** [33:40]
... insensitive, hallucinations, um-

**Raiza Martin** [33:42]
Um, safe- other safety things, right?

**Usama Bin Shafqat** [33:44]
Right. Like a bunch of-

**Raiza Martin** [33:44]
Instead of just reader safety stuff. Yeah.

**Usama Bin Shafqat** [33:46]
Yeah. Exactly. So, like, with all of that, and, like, also just, you know, just following sort of a coherent narrative and structure is really important, but, like, with all of this, we really had to make sure that that central tenet of being entertaining and engaging and something you actually want to listen to, it just doesn't go away, which takes, like, a lot of just active listening time-

**Raiza Martin** [34:04]
Yeah

**Usama Bin Shafqat** [34:04]
... 'cause you're closest to the-

**Raiza Martin** [34:06]
Yeah

**Usama Bin Shafqat** [34:06]
... the prompts, the model, and everything.

**Raiza Martin** [34:07]
I think sometimes the, the, it difficulty is because we're dealing with non-deterministic models-

**Usama Bin Shafqat** [34:12]
Yeah

**Raiza Martin** [34:13]
... sometimes you just got a bad roll of the dice.

**Usama Bin Shafqat** [34:14]
Mm-hmm.

**Raiza Martin** [34:15]
And it's always on the distribution that you could get something bad. Basically, how, how many... Do you, like, do 10 runs at a time, and then how do you get rid of the non-determinism ?

**Usama Bin Shafqat** [34:24]
Right. Yeah. That's, uh, that's the-

**Raiza Martin** [34:25]
Like bad luck.

**Usama Bin Shafqat** [34:26]
Yeah, yeah. I mean, there still will be, like, bad audio overviews. There's, like, a bunch of them that happens. Um, do you mean for, like, the rater evals?

**Raiza Martin** [34:33]
For raters, right?

**Usama Bin Shafqat** [34:34]
Yeah.

**Raiza Martin** [34:35]
Like, what if that one person just got, like, a really bad rating? You actually had a great prompt, you actually had great model, great weights, whatever, and you just, you had a bad output. Like, and that's okay, right?

I actually think, like, the way that these are constructed, if you think about, like, the different types of controls that the user has, right? Like, what can the user do today to affect it?

**Usama Bin Shafqat** [34:55]
We push a button.

**Raiza Martin** [34:55]
Just push resources. You just push a button.

**Usama Bin Shafqat** [34:56]
I, I have tried to prompt engineer by changing the title.

**Raiza Martin** [34:59]
Yeah, yeah.

**Usama Bin Shafqat** [35:00]
Mm-hmm.

**Raiza Martin** [35:00]
Changing the title.

**Usama Bin Shafqat** [35:01]
Yeah.

**Raiza Martin** [35:01]
People have found out-

**Usama Bin Shafqat** [35:02]
Title of the notebook. Yeah

**Raiza Martin** [35:03]
... the title of the notebook. People have found out you can add show notes, right?

**Usama Bin Shafqat** [35:06]
Yeah, yeah.

**Raiza Martin** [35:06]
You can get them to think like the, the show has changed-

**Usama Bin Shafqat** [35:08]
Someone changed the language of the output

**Raiza Martin** [35:08]
... sort of fundamentally. Changing the language of the output. Like, those are, are less well-tested because we focused on, like, this one aspect, so it, it did change the way that we sort of think about quality as well, right?

So it's like quality is on the dimensions of entertainment, of course, like consistency, groundedness, but in general, does it follow the structure of the Deep Dive? And I think when we talk about, like, non-determinism, it's like, well, as long as it follows, like, the structure of the Deep Dive, right, it sort of inherently meets all those other qualities.

And so it makes it a little bit easier for us to ship something with confidence to the extent that it's like I know it's gonna make a Deep Dive. It's gonna make a good Deep Dive. Whether or not the person likes it, I don't know.

But as we expand to new formats, as we open up controls, I think that's where it gets really much harder. Even with the show notes, right? Like, people don't know what they're going to get when they do that, and we see that already where it's like this is gonna be a lot harder to validate in terms of quality, where now we'll get a greater distribution.

Whereas I don't think we really got, like, varied distribution because of, like, that pre-process that Usama was talking about, and also because of the way that we'd constrained, like, what were we measuring for? Literally, just like, is it a Deep Dive?

Yes or no. And you determine what a Deep Dive is. Yeah.

**Usama Bin Shafqat** [36:20]
Everything needs a PM. Yeah, I, I, I, I have, um... This is very similar to something I've been thinking about for AI products in general. There's always, like, a chief tastemaker, and for NotebookLM, it seems like it's a combination of you and Steven.

**Raiza Martin** [36:32]
Well, okay, I, I wanna take a step back-

**Usama Bin Shafqat** [36:33]
And, and Usama. I mean, presumably for the voice stuff.

**Raiza Martin** [36:36]
Usama's, like, the, like the head chef, right, of, like, Deep Dive, I think.

**Usama Bin Shafqat** [36:40]
Potatoes.

**Raiza Martin** [36:40]
Of potatoes. And I, I say this because I think even though we are already a very opinionated team, and Steven for sure very opinionated, I think of the audio generations, like, Usama was the most-

**Usama Bin Shafqat** [36:53]
Okay

**Raiza Martin** [36:53]
... opinionated, right? And we all, we all, like, would say, like, "Hey," I remember, like, one of the first ones he sent me, I was like, "Oh, I feel like they should introduce themselves. I feel like they should say a title."

But then, like, he-- we would catch things like maybe they shouldn't say their names.

**Usama Bin Shafqat** [37:05]
Yeah, they don't say their names. That was a Steven catch.

**Raiza Martin** [37:07]
Yeah, yeah.

**Usama Bin Shafqat** [37:07]
Like, not give them names.

**Raiza Martin** [37:09]
So stuff like that is just like we all injected, like, a little bit of just like, "Hey, here's, like, my take on, like, how a podcast should be," right? And I think, like, if you're a person who, like, regularly listens to podcasts, there's probably some collective preference there that's generic enough that you can standardize into, like, the Deep Dive format.

**Usama Bin Shafqat** [37:26]
Mm-hmm.

**Raiza Martin** [37:27]
But yeah, it's the new formats where I think, like, "Oh, that's the next test."

**Usama Bin Shafqat** [37:30]
Yeah. I've tried to made a, make a clone, by the way. Of course, everyone did.

**Raiza Martin** [37:33]
Yeah.

**Usama Bin Shafqat** [37:33]
Everyone in AI was like, "Oh, you know, this is so easy. I'll just take a TTS model."

**Raiza Martin** [37:36]
Yeah.

**Usama Bin Shafqat** [37:37]
Obviously, our, our models are not as good as yours, but I tried to in-inject a consistent character backstory, like age-

**Swyx** [37:43]
... identity-

**Raiza Martin** [37:44]
Yeah

**Swyx** [37:44]
... where they went to, where they work, where they went to school, what their hobbies are. Then it just, then the models try to bring it in too much.

**Raiza Martin** [37:50]
Yeah.

**Usama Bin Shafqat** [37:50]
Yeah, yeah.

**Swyx** [37:50]
I don't know if you tried this.

**Usama Bin Shafqat** [37:51]
Yeah.

**Swyx** [37:51]
So then I'm like, "Okay, like, how do I define a personality but it doesn't keep coming up-

**Raiza Martin** [37:56]
Mm-hmm

**Swyx** [37:57]
... every single time?"

**Raiza Martin** [37:58]
Yeah, I mean, we have, like, a really, really good, like, character designer on our team.

**Swyx** [38:02]
What? Like a D&D person? Uh...

**Raiza Martin** [38:05]
Just to say, like, we, just like we had to be opinionated about the format-

**Swyx** [38:09]
Yeah

**Raiza Martin** [38:09]
... we had to be opinionated about who are those two people talking.

**Swyx** [38:12]
Okay.

**Raiza Martin** [38:12]
Right? And then to the extent that, like, you can design the format, you should be able to design the people as well.

**Swyx** [38:19]
Yeah. I would love, like, a... You know, like, when you play Baldur's Gate. Like, you s-

**Raiza Martin** [38:23]
Yeah, I love that

**Swyx** [38:23]
... you roll, you roll, like, 17 on charisma, and, like, you, you see, like, what race they are. I don't know.

**Raiza Martin** [38:27]
I recently, actually, I was just talking about character select screens.

**Swyx** [38:30]
Yeah.

**Raiza Martin** [38:30]
I was like-

**Swyx** [38:31]
People spend hours in them

**Raiza Martin** [38:32]
... I love that.

**Swyx** [38:32]
Yeah.

**Raiza Martin** [38:32]
I love that, right? And I, I was like, maybe there is something to be learned there because, like, people have fallen in love with the Deep Dive as a, as a format, as a technology, but also as just, like, those two personas.

Now, when you hear a Deep Dive-

**Swyx** [38:46]
Mm-hmm

**Raiza Martin** [38:46]
... and you've heard them, you're like, "I know those two."

**Swyx** [38:48]
Mm-hmm.

**Raiza Martin** [38:49]
Right? And people... It's so funny when I, when people are trying to find out their names. I'm like, it's a, it's a worthy task. It's a worthy goal. I know what you're doing.

**Swyx** [38:57]
Yeah.

**Raiza Martin** [38:57]
But the next step here is to sort of introduce, like, is this, like, what people want, people want to sort of edit the personas, or do they just want more of them?

### Engagement

**Swyx** [39:04]
I'm sure you're getting a lot of opinions, and they all, they all conflict with each other. Before we move on, I have to ask, because we're kind of on this topic, how do you make audio engaging? Because it's useful not just for Deep Dive, but also for us-

**Usama Bin Shafqat** [39:16]
Mm-hmm

**Swyx** [39:17]
... as podcasters.

**Raiza Martin** [39:18]
Yeah.

**Swyx** [39:18]
What is eng- what does engaging mean? Um, if you could break it down for us, that'd be great.

**Usama Bin Shafqat** [39:22]
I mean, I can try. Like I don't, don't claim to be an expert at all.

**Swyx** [39:25]
Yeah, yeah.

**Usama Bin Shafqat** [39:25]
But, uh-

**Swyx** [39:26]
So I, I'll give you some.

**Usama Bin Shafqat** [39:27]
Yeah.

**Swyx** [39:28]
Like, va- variation in tone-

**Usama Bin Shafqat** [39:29]
Right

**Swyx** [39:30]
... and speed. You know, there's this sort of writing advice where, you know, this sentence is five words, this sentence is three. That kind of advice, where you, where you vary things. You have excitement. You have laughter, all that stuff.

But I'd be curious how else you break down.

**Usama Bin Shafqat** [39:42]
So there's the basics, like obviously structure.

**Swyx** [39:44]
Structure.

**Usama Bin Shafqat** [39:44]
It can't be meandering, right? Like, you, there needs to be sort of a, a, an ultimate goal that the voices are trying to get to-

**Swyx** [39:50]
Yeah

**Usama Bin Shafqat** [39:50]
... human or artificial. I think one thing we find often is if there's just too much agreement between people, like, that's not fun to listen to, so there needs to be some sort of tension and build up. You know, withholding information, for example.

Like, as you listen to a story unfold, like, you're gonna learn more and more about it. In audio, that maybe becomes even more important because, like, you actually don't have the ability to just, like, skim to the end of something.

You're driving or something. Like, you're gonna be hooked 'cause, like, there's... And that's how, like, that's how a lot of podcasts work. Like, maybe not interviews necessarily, but a lot of true crime, a lot of-

**Raiza Martin** [40:29]
Mm-hmm

**Usama Bin Shafqat** [40:29]
... entertainment in general. There's just, like, a gradual unrolling of information, and that also, like, sort of goes back to the content transformation aspect of it. Like, maybe you are going from, let's say, the Wikipedia article of, like, one of the History of Mysteries maybe episodes.

Like, the Wikipedia article's gonna state out the information very differently. It's like, "Here's what happened," would probably be in the very first, like-

**Raiza Martin** [40:51]
Yeah

**Usama Bin Shafqat** [40:51]
... paragraph. And one approach we could have done is, like, maybe a person's just narrating that thing.

**Swyx** [40:56]
Mm-hmm.

**Usama Bin Shafqat** [40:56]
Um, and maybe that would work for, like, a certain audience. Um, or I guess that, that's how I would picture, like, a standard history lesson to unfold. But, like, because we're trying to put it in this two-person dialogue format, like, there, we, we inject, like, the fact that, you know, there's, you don't give everything at first, and then you set up, like, differing opinions of the same topic or the same...

Like, maybe you seize on a topic and go deeper into it and then try to bring yourself back out of it and go back to the, the main narrative. So that's, that's mostly from, like, the, the setting up the, the script perspective.

And then the audio, like I was saying earlier, it's trying to be as close to just human speech as possible, I think was the, what we found success with so far.

**Raiza Martin** [41:40]
Yeah. Like with, with interjections, right? Like, I think, like, when you listen to two people talk, there's a lot of like, "Yeah, yeah, right." And then there's, like, a lot of, like, that questioning.

**Swyx** [41:48]
Mm.

**Raiza Martin** [41:48]
Like, "Oh yeah? Really? What did you think?"

**Swyx** [41:50]
I noticed that. That's great.

**Raiza Martin** [41:53]
Totally.

**Swyx** [41:54]
Like, so-

**Usama Bin Shafqat** [41:55]
Exactly

**Swyx** [41:55]
... my question is, do you pull in speech experts to do this, or did you just come up with it yourselves? You can be like, okay, talk to a whole bunch of fiction writers to, to make things engaging, or comedy writers or whatever.

Stand-up comedy, right? They, they have to make audio engaging.

**Raiza Martin** [42:10]
Yeah.

**Swyx** [42:10]
Uh, but audio as well. Like, there's professional fields of studying-

**Raiza Martin** [42:13]
Yeah, yeah

**Swyx** [42:14]
... that, that, where people do this for a living, but us as AI engineers are just making this up as we go.

**Raiza Martin** [42:20]
Well, that, I mean, it's a great idea, but you definitely didn't.

**Swyx** [42:23]
Yeah.

**Raiza Martin** [42:23]
You know? And now I'm just like, "Oh."

**Swyx** [42:24]
My, my guess is, my guess is you didn't.

**Raiza Martin** [42:26]
Yeah.

**Swyx** [42:26]
There's a, a, there's a certain appeal to authority that people have. They're like, "Oh," like, "you can't do this 'cause you don't have any experience, like, making engaging audio," but that's what you literally did-

**Usama Bin Shafqat** [42:36]
Right. I mean-

**Swyx** [42:36]
... from first principles

**Usama Bin Shafqat** [42:36]
... I was, I was literally chatting with someone at Google earlier today about how some people think that, like, you need a linguistics person-

**Swyx** [42:43]
Yeah

**Usama Bin Shafqat** [42:43]
... in the room-

**Swyx** [42:44]
Yeah

**Usama Bin Shafqat** [42:44]
... for, like, making a good chatbot, but that's not actually true. 'Cause, like, this person went to school for linguistics, and according to him... He's a, he's an engineer now. According to him, like, most of his classmates were not actually good at language.

Like, they knew how to analyze language and, like, sort of the mathematical patterns and rhythms in language.

**Swyx** [43:01]
Mm-hmm.

**Usama Bin Shafqat** [43:02]
But that doesn't necessarily mean they were gonna be eloquent at, like, while speaking or writing.

**Swyx** [43:07]
Mm.

**Usama Bin Shafqat** [43:07]
So I think, yeah, a lot of, uh, we haven't invested in specialists-

**Raiza Martin** [43:11]
I know. I know

**Usama Bin Shafqat** [43:12]
... in the audio format yet, but maybe that would-

**Raiza Martin** [43:14]
I think it's, like, super interesting because I think there is, like, a very human question of, like, what makes something interesting, and there's, like, a very deep question of, like, what is it, right? Like, what, what is the quality that we are all looking for?

Is it, does somebody have to be funny? Does something have to be entertaining? Does something have to be straight to the point? And I think when you try to distill that, this is the interesting thing I think about our experiment, about this particular launch, is first, we only launched one format, and so we sort of had to squeeze everything we believed about what an interesting thing is into one package.

And as a result of it, I think we learned it's like, hey, interacting with a chatbot ... is sort of novel at first, but it's not interesting, right? It's like humans are what makes interacting with chatbots interesting. It's like, "I'm gonna try to trick it."

It's like, "That's interesting. Spell strawberry," right? This is, like, the fun that, like, people have with it, but, like, that's not the LLM being interesting.

**Usama Bin Shafqat** [44:09]
Mm.

**Raiza Martin** [44:09]
That's you just, like, kind of giving it your own flavor. But it's like what does it mean to sort of flip it on its head and say, "No, you be interesting now," right? Like, you give the chatbot the opportunity to do it, and this is not a chatbot per se, it is, like, just the audio, and it's, like, the texture I think that really brings it to life.

And it's, like, the things that we've described here, which is like, okay, now I have to, like, lead you down a path of information about, like, this commercialization deck. It's like, how do you do that? To be able to successfully do it, I do think that you need experts.

I think we'll engage with experts, like, down the road, but I think it will have to be in the context of, well, what's the next thing we're building, right? It's like, what am I trying to change here? What do I fundamentally believe needs to be improved?

And I think there's still, like, a lot more studying that we have to do in terms of, like, well, what, what are people actually using this for? And we're just in such early days. Like, it, it hasn't even been a month.

**Usama Bin Shafqat** [45:05]
Two, three weeks. Three weeks, I think.

**Raiza Martin** [45:07]
Yeah.

**Usama Bin Shafqat** [45:07]
Yeah.

**Raiza Martin** [45:07]
Yeah.

**Usama Bin Shafqat** [45:08]
Um, I think the other... w- one other element to that is the, like, the fact that you're bringing your own sources-

**Raiza Martin** [45:12]
Yeah

**Usama Bin Shafqat** [45:13]
... to it.

**Swyx** [45:13]
Mm.

**Usama Bin Shafqat** [45:13]
Like, it's your stuff. Like, you know-

**Swyx** [45:15]
Mm

**Usama Bin Shafqat** [45:15]
... this somewhat well or you care to know about this. So, like, that I think changed the equation on its head as well. It's, like, your sources and someone's telling you about it.

**Swyx** [45:25]
Yeah.

**Usama Bin Shafqat** [45:25]
So, like, you care about how that dynamic is, but you just care for it to be good enough to-

**Raiza Martin** [45:29]
Yeah

**Usama Bin Shafqat** [45:29]
... be entertaining.

**Swyx** [45:30]
Mm, yeah.

**Usama Bin Shafqat** [45:30]
Or 'cause ultimately they're talking about your mortgage deed or whatever.

**Swyx** [45:33]
Right.

**Raiza Martin** [45:34]
Yeah.

**Swyx** [45:34]
So it's interesting just from the topic itself, even taking out all the agreements and the hiding of the slow reveal of the-

**Usama Bin Shafqat** [45:41]
I mean, there's a baseline maybe

**Swyx** [45:42]
... of the topic. Yeah.

**Usama Bin Shafqat** [45:42]
Like, if it was, like, too drab. Like, if it was someone was reading it off, like, you know, that's, like, the absolute, like, worst.

**Swyx** [45:47]
Yeah.

**Usama Bin Shafqat** [45:47]
But, like, um-

**Swyx** [45:48]
Do you prompt for humor?

**Usama Bin Shafqat** [45:50]
Um-

**Swyx** [45:51]
That's a tough one, right?

**Raiza Martin** [45:52]
I think it's more of a, a generic way to bring humor out, if possible. I think humor is actually one of the hardest things.

**Swyx** [46:00]
Yeah.

**Raiza Martin** [46:00]
But I, I don't know if you saw-

**Swyx** [46:01]
That is AGI. Humor is AGI.

**Raiza Martin** [46:02]
Yeah. But did you see the chicken one?

**Swyx** [46:04]
No.

**Raiza Martin** [46:04]
Okay, if you haven't heard it-

**Swyx** [46:06]
We'll splice it in here.

**Raiza Martin** [46:07]
Okay. Yeah, yeah.

**Swyx** [46:07]
Yeah, yeah.

**Raiza Martin** [46:08]
There is a video on Threads, I think it was by Martino Wong, and, um, it's a, a PDF-

Welcome to your deep dive for today.

**Raiza Martin** [46:18]
Oh, yeah. Get ready for a fun one.

Buckle up- ... because we are diving into Chicken, Chicken, Chicken, Chicken, Chicken.

**Raiza Martin** [46:29]
You got that right.

By Doug Zunker.

**Raiza Martin** [46:31]
No.

And yes, you heard that title correctly.

**Raiza Martin** [46:34]
Titles.

Our listener today submitted this paper.

**Raiza Martin** [46:37]
Yeah, they're gonna need our help.

And I can totally see why.

**Raiza Martin** [46:39]
Absolutely.

It's dense, it's baffling.

**Raiza Martin** [46:42]
It's a lot.

And it's packed with more chicken- ... than a KFC buffet.

**Raiza Martin** [46:47]
What? That's hilarious. That's so funny. So it's, like, stuff like that that, that's, like, truly delightful, truly surprising, but it's like we didn't tell it to be funny.

**Swyx** [46:55]
Mm-hmm.

**Usama Bin Shafqat** [46:55]
Humor's contextual also, like super contextual is what we're realizing. So we're not prompting for humor, but we're prompting for maybe a lot of other things that are bringing out that humor.

**Swyx** [47:04]
I think the thing about AI-generated content, if we look at YouTube, like we do videos on YouTube, and it's like, you know, a lot of people like screaming in the thumbnails to get clicks. There's like- ... everybody... There's kind of like a meta of, like, what you need to do to get clicks.

But I think in your product, there's no actual creator on the other side investing the time, so you can actually generate a type of content that is maybe not universally appealing, you know?

**Raiza Martin** [47:29]
Yeah.

**Swyx** [47:29]
At a much-

**Usama Bin Shafqat** [47:29]
It's personal.

**Swyx** [47:30]
Yeah, exactly. I think that's the most interesting thing is, like, well, is there a way for, like... Take MrBeast.

**Usama Bin Shafqat** [47:36]
Right.

**Swyx** [47:37]
It's like MrBeast optimizes videos to reach the biggest audience and, like, the most clicks, but what if every video could be kind of, like, regenerated to be closer-

**Raiza Martin** [47:46]
Yeah. Yeah

**Swyx** [47:46]
... to your taste, you know, when you watch it?

**Raiza Martin** [47:49]
I think that's kind of the promise of AI that I, I think we are just, like, touching on, which is I think every time I've gotten information from somebody, they have delivered it to me in their preferred method, right?

Like, if somebody gives me a PDF, it's a PDF. Somebody gives me a 100-slide deck, that is the format in which I'm going to read it. But I think we are now living in the era where transformations are really possible, which is, look, like, I don't want to read your 100-slide deck, but I'll listen to a 16-minute audio overview on the drive home.

**Swyx** [48:15]
Yeah.

**Raiza Martin** [48:16]
And that, that I think is, is really novel, and that is, is paving the way in a way that, like, maybe we wanted but didn't expect, where I also think you're listening to a lot of content that normally wouldn't have had content made about it.

Like, I, I watched this TikTok where this woman uploaded her diary from 2004. For sure, right? Like, nobody was gonna make a podcast about a diary. Like, hope- hopefully not. Like, it seems kind of embarrassing.

**Swyx** [48:42]
It's kind of creepy.

**Raiza Martin** [48:42]
Yeah, it's kind of creepy. But she was, she was doing this, like, live listen of, like, oh, like, here's a podcast about my diary, and it's like, it's entertaining right now to sort of all listen to it together, but, like, the connection is personal.

It was, like, it was her interacting with, like, her information in a totally different way, and I think that's where, like... Oh, that's a super interesting space, right? Where it's like I'm creating content for myself in a way that suits the way that I wanna, I wanna consume it.

**Swyx** [49:07]
Yeah.

**Usama Bin Shafqat** [49:07]
Or people compare, like, retirement plan options.

**Raiza Martin** [49:09]
Yeah.

**Usama Bin Shafqat** [49:10]
Like, no one's gonna give you that content, like, for your personal financial situation.

**Raiza Martin** [49:15]
Yeah, yeah.

**Usama Bin Shafqat** [49:15]
And, like, even when we started out the experiment, like, a lot of the goal was to go for really obscure content and see how well we could transform that. So, like, if you look at the, the Mountain View, like, city council meeting notes, like, you're never gonna read it, but, like, if it was a three-minute summary, like, that would be interesting.

**Swyx** [49:33]
I see you have one system, one prompt that just covers everything you throw at it.

**Raiza Martin** [49:38]
Maybe.

**Usama Bin Shafqat** [49:39]
No.

**Swyx** [49:39]
I'm just, I, I am just-

**Raiza Martin** [49:40]
I'm just kidding.

**Swyx** [49:41]
It's really interesting. You know what, um, I'm trying to figure out what you nailed f- compared to others, and I think that the way that you treat your, the AI is, like, a, a little bit different than a lot of the builders I talk to.

So I don't know what is, what it is you said. I wish I had a transcript right in front of me. But it's, it's something like people treat AI as like a tool for thought, but usually it's kind of doing their bidding.

**Raiza Martin** [50:01]
Mm-hmm.

**Swyx** [50:01]
And, you know, what you're really doing is loading up these like two virtual agents. Uh, I don't... I... You've never said the word agents, I put that in your mouth, but two virtual humans or AIs, and letting them form, form their own opinion and letting them kind of just live and embody it a little bit.

**Raiza Martin** [50:17]
Mm.

**Swyx** [50:17]
Is that accurate?

**Raiza Martin** [50:18]
I think that that is as close to accurate as possible. I, I mean, in general I try to be careful about saying like, oh, you know, letting, you know-

**Swyx** [50:25]
Personification.

**Raiza Martin** [50:26]
Yeah.

**Swyx** [50:26]
Yeah.

**Raiza Martin** [50:26]
Like, these, these personas live. But I think to your earlier question of like, what makes it interesting? That's what it takes to make it interesting.

**Swyx** [50:32]
Yeah.

**Raiza Martin** [50:33]
Right? And I think to do it well is like a worthy challenge. I also think that it's interesting because they're interested, right? Like, is it interesting to compare two-

**Swyx** [50:41]
Oh, this is the Dale Carnegie thing.

**Raiza Martin** [50:42]
Yeah. Is it, is it interesting to have two retirement plans? No. But to listen to, to these two talk about it, oh my gosh, you'd think it was like the best thing ever invented, right? It's like, "Get this.

Deep dive into 401 through Chase versus..." Or, you know, whatever, right?

**Swyx** [51:01]
They do do a lot of "get this," which is funny.

**Raiza Martin** [51:03]
Yeah. I know, I know, I, I dream about it. I'm sorry.

**Swyx** [51:08]
Um, th- there's a... I- I have a, I have a few more questions on just like the engineering around this, um, and obviously some of this is just me creatively asking how, how this works. How do you make decisions between when to trust the AI overlord to decide for you?

In other words, stick it... Let's say product as it is today, it, um, you want to improve it in some way. Do you engineer it into the system, like write code to make sure it happens, or you just stick it in a prompt and hope that the LLM does it for you?

Do you know what I mean?

**Raiza Martin** [51:39]
Do you mean specifically about audio or sort of in general, as like an approach?

**Swyx** [51:41]
In general. Like w- um, designing AI products, I think this is like the one thing that people are struggling with.

**Raiza Martin** [51:47]
Yeah. Yeah.

**Swyx** [51:48]
Um, and there's, there's compound AI people, and then there's big AI people.

**Raiza Martin** [51:51]
Yeah.

**Swyx** [51:51]
So compound AI people would be like Databricks.

**Raiza Martin** [51:53]
Yeah.

**Swyx** [51:53]
Have lots of little models, chain them together-

**Raiza Martin** [51:55]
Yep, yep

**Swyx** [51:56]
... to, to make an output. It's deterministic. You control every single piece-

**Raiza Martin** [51:59]
Yeah

**Swyx** [51:59]
... and, you know, you, you produce what you produce. The open AI people, totally the opposite. Like, write one giant prompt and let the model figure it out.

**Raiza Martin** [52:06]
Yeah.

**Swyx** [52:06]
And obviously, the, the, the answer for most people is gonna be a spectrum in between those two, like big model, small model. When do you decide that?

**Raiza Martin** [52:12]
I think it depends on the task. It also depends on... Well, it depends on the task, but ultimately depends on what is your desired outcome. Like, what am I engineering for here? And I think there's like several potential outputs, and they're sort of like general categories.

Am I trying to delight somebody? Am I trying to just like meet whatever the person is trying to do? Am I trying to sort of simplify a workflow? At what layer am I implementing this? Am I trying to implement this as part of the stack to re- reduce like friction, you know, particularly for like engineers or something?

Or am I trying to engineer it so that I deliver like a super high quality thing? I think that the question of like which of those two, I think you're right, it is a spectrum, but I think fundamentally it comes down to like it's a craft.

Like it's still a craft as much as it is a science, and I think the reality is like you have to have a really strong POV about like what you want to get out of it and to be able to make that decision.

Because I think if you don't have that strong POV, like you're gonna get lost in sort of the detail of like capability, and capability is sort of the last thing that matters because it's like models will catch up, right?

Like, models will be able to do, you know, whatever in the next five years. It's gonna be insane. So I think this is like a race to like value, and it's like really having a strong opinion about like what does that look like today and how far are you gonna be able to push it?

Sorry, I think maybe that was like very like philosophical.

**Swyx** [53:31]
Philosophical.

**Raiza Martin** [53:31]
But-

**Swyx** [53:31]
Yeah, that's fine. We get, we get there.

**Usama Bin Shafqat** [53:33]
And I think that, that, that hits a lot of the points I was gonna make.

### Feature Requests

**Alessio** [53:36]
I tweeted today, or I X-posted, uh, whatever, um-

**Swyx** [53:40]
X-posted

**Alessio** [53:40]
... that we're gonna interview you and what we should ask you.

**Raiza Martin** [53:42]
Okay.

**Alessio** [53:43]
So we got-

**Swyx** [53:43]
Yes

**Alessio** [53:43]
... a list of feature requests mostly.

**Raiza Martin** [53:45]
Oh.

**Alessio** [53:46]
It's funny, nobody actually had any like specific questions about how the product was built. They just wanna know when you're releasing some feature.

**Raiza Martin** [53:52]
Okay.

**Alessio** [53:52]
So I know you cannot talk about all of these things, but I think maybe it will give people an idea of like where the product is going. So I think the most common question I think five people asked is like, "Are you gonna build an API?"

**Raiza Martin** [54:03]
Yeah.

**Alessio** [54:04]
And, you know, do you see this product as still be kind of like a full head product for like a login and do everything there, or do you want it to be a piece of infrastructure that people build on?

**Raiza Martin** [54:13]
I mean, I think why not both, right? I think we work at a place where you could have both. I think that end user products, like products that touch the hands of users, have a lot of value. For me personally, like we learn a lot about what people are trying to do and what's like actually useful and what people are ready for, and so we're gonna keep investing in that.

I think at the same time, right, like there, there are a lot of developers that are interested in using the same technology to build their own thing. We're going to look into that. How soon that's going to be ready, I can't really comment, but these are the things that like, hey, we heard it, we're trying to figure it out, and I think there's room for both.

**Swyx** [54:51]
Is there a world in which this becomes the default Gemini interface because it's technically different org?

**Raiza Martin** [54:56]
It's such a good question, and I think every, every time someone asks me it's like, "Hey, I just lead NotebookLM."

**Swyx** [55:02]
Yeah.

**Raiza Martin** [55:02]
We'll, we'll ask the, the Gemini folks what they think.

**Alessio** [55:05]
Multilingual support. I know people kind of hack this a little bit together. Any ideas for full support? But also I- I'm also interested in dialects. In Italy we have Italian obviously, but we have a lot of local dialects.

Like if you go to Rome, people don't really speak Italian, they speak a local dialect.

**Raiza Martin** [55:21]
Oh.

**Alessio** [55:21]
Do you think there's a path to which these models, especially the, the speech can learn very like niche dialects? Like how much data do you need? Can people contribute? Like, uh-

**Swyx** [55:32]
Ooh

**Alessio** [55:32]
... I- I'm curious like-

**Raiza Martin** [55:33]
Good one

**Alessio** [55:33]
... if you see this as a possibility.

**Usama Bin Shafqat** [55:35]
Yeah. Totally. So I guess high level, like we're definitely working on adding more languages. That's like top priority. We're gonna start small, but like theoretically we should be able to cover like most languages pretty soon. What I don't, don't, don't, don't want-

**Swyx** [55:47]
What a ridiculous statement by the way. That's, that's crazy.

**Usama Bin Shafqat** [55:50]
... on like the, the soon or the pretty soon part.

**Swyx** [55:53]
Yeah. No, but like, you know, a few years ago s- like a small team of like, I don't know, 10 people saying that we will support the top like 100, 200 languages is like absurd, but-

**Usama Bin Shafqat** [56:02]
Right

**Swyx** [56:02]
... you can do it.

**Raiza Martin** [56:03]
Yeah.

**Swyx** [56:03]
You can do it.

**Raiza Martin** [56:03]
And I, and I think like the speech team-

**Swyx** [56:06]
Mm-hmm

**Raiza Martin** [56:06]
... you know, w- we are a small team.

**Swyx** [56:08]
Yeah.

**Raiza Martin** [56:08]
But the speech team is another team, and the modeling team, like these folks are just like- Absolutely brilliant at what they do, and I think, like, when we've talked with them and we've said, "Hey, you know, how about more languages?

How about more voices? How about dialects," right? This is something that, like, th- they are game to do, and, like, that's, that's the roadmap for them.

**Usama Bin Shafqat** [56:26]
The speech team supports, like, a bunch of other efforts across Google. Like Gemini Live-

**Raiza Martin** [56:30]
Yeah, yeah

**Usama Bin Shafqat** [56:30]
... for example, is also the model's built by the same, like, sort of DeepMind speech team. But yeah, the, the, the thing about dialects is really interesting 'cause, like, in some of our sort of earliest testing with trying out other languages, we actually noticed that sometimes it wouldn't stick to a certain dialect.

**Swyx** [56:46]
Mm.

**Usama Bin Shafqat** [56:46]
Especially for, like, I think for French we noticed that. Like, when we presented it to, like, a native speaker it would sometimes go from, like, a Canadian person speaking French versus-

**Swyx** [56:54]
Yeah, yeah

**Usama Bin Shafqat** [56:54]
... like a, a French person-

**Swyx** [56:55]
Yeah, yeah, yeah

**Usama Bin Shafqat** [56:55]
... speaking French, or an American person speaking French, which is not what we wanted.

**Swyx** [56:58]
Right.

**Usama Bin Shafqat** [56:59]
Um, so there's a lot more sort of speech quality work that we need to do there to make sure that it works reliably in at least sort of like the, like the, the standard dialect that we want. But that does show that there's potential to sort of do the thing that you're talking about of, like, fixing a dialect that you want, maybe contribute your own voice or, like, you pick from one of the options.

There's, there's a lot more headroom there.

**Swyx** [57:20]
Yeah. Because we have movies. Like, we have old Roman movies that have, like, different-

**Raiza Martin** [57:24]
Yeah

**Swyx** [57:24]
... uh, languages, but there's not that many, you know? So I'm always like, well, I'm sure, like, the Italian is so strong in the model that, like, when you're trying to, like, pull that away from it, like you kind of need a lot, but...

**Usama Bin Shafqat** [57:36]
Right, that's, that's all sort of, like, wonderful DeepMind speech team-

**Raiza Martin** [57:39]
Yeah

**Swyx** [57:39]
Yeah, yeah, yeah

**Usama Bin Shafqat** [57:40]
... work.

**Swyx** [57:40]
Well, anyway, if you need Italian, he's got you.

**Raiza Martin** [57:42]
Yes.

**Swyx** [57:42]
Trying to be ready. He's got you.

**Usama Bin Shafqat** [57:43]
I got a mic. I got a mic, so.

**Swyx** [57:44]
Yeah, yeah. Specifically Singlish, I got you. Managing system prompts, people want a lot of that.

**Raiza Martin** [57:49]
Mm-hmm, mm.

**Swyx** [57:50]
I assume yes-ish.

**Raiza Martin** [57:51]
Definitely looking into it, uh, for just core NotebookLM.

**Swyx** [57:56]
Yeah.

**Raiza Martin** [57:56]
Like, everybody's wanted that forever, so we're working on that. I think for the audio itself, we're trying to figure out the best way to do it, so we'll launch something sooner rather than later. So we'll probably stage it, and I think, like, you know, just to be fully transparent, we'll probably launch something that's more of a fast follow than, like, a fully baked feature first.

**Swyx** [58:15]
Mm-hmm.

**Raiza Martin** [58:15]
Just because, like, I see so many people put in, like, the fake show notes. It's like, "Hey, I'll, I'll help you out. We'll just put a text box or something."

**Swyx** [58:21]
Yeah.

**Raiza Martin** [58:22]
Yeah.

**Usama Bin Shafqat** [58:22]
And I think a lot of people are like, "This is almost perfect, but, like, I just need that extra 10, 20%."

**Swyx** [58:26]
Yeah.

**Raiza Martin** [58:26]
Yeah.

**Swyx** [58:26]
I noticed that you say no a lot, I think, or you, you try to ship one thing.

**Raiza Martin** [58:31]
Yeah.

**Swyx** [58:32]
And that, that's different about you than maybe other PMs or other eng teams that try to sh- they're like, "Oh, here are all the knobs. Um, just take all my knobs," you know?

**Raiza Martin** [58:40]
Yeah, yeah.

**Swyx** [58:40]
Top P, top K, doesn't matter. I'll just put it in the docs and, like, you figure it out, right?

**Raiza Martin** [58:44]
That's right. That's right.

**Swyx** [58:45]
Uh, whereas for you it's... You, you actually just... You make one product-

**Raiza Martin** [58:49]
Yeah

**Swyx** [58:50]
... as opposed to, like, 10 you could possibly have done.

**Raiza Martin** [58:52]
Yeah.

**Swyx** [58:52]
I don't know. It's interesting.

**Raiza Martin** [58:53]
I, I think about this a lot. I think it requires a lot of discipline because I thought about the knobs. I was like, "Oh, I saw on Twitter, or, you know, on X, people want the knobs." Like, great.

Start mocking it up, making the text boxes, designing, like, the little fiddles, right? And then I looked at it and I was kind of sad. I was like, "Well," right? It's like, "Well," it's like, "This is not cool.

This is not fun. This is not magical." It is sort of exactly what you would expect knobs-

**Swyx** [59:18]
Mm-hmm

**Raiza Martin** [59:18]
... to be.

**Swyx** [59:19]
Yeah.

**Raiza Martin** [59:19]
And then, you know, it's like, well, I mean, how, how much can you, you know, design a knob?

**Swyx** [59:25]
Yeah.

**Raiza Martin** [59:25]
I thought about it, and I was like... But the thing that people really like was that there wasn't any, that they just pushed a button.

**Swyx** [59:31]
One button.

**Raiza Martin** [59:31]
And it was cool. And so I was like, how do we bring more of that, right? That still gives the user the optionality that they want. And so this is where, like, you have to have a strong POV, I think.

You have to, like, really boil down, what did I learn in, like, the months since I've launched this thing that people really want, and I can give it to them while preserving, like, that, that delightful sort of fun experience?

And I think that's actually really hard. Like, I'm, I'm not gonna come up with that by myself. Like, that's something that, like, our team thinks about every day. We all have different ideas. We're all experimenting with sort of how to get the most out of, like, the insight and also ship it quick.

So, so we'll see. We'll find out soon if people like it or not.

**Usama Bin Shafqat** [1:00:09]
I think the other interesting thing about, like, AI development now is that the knobs are not necessarily... Like, sp- going back to all the sort of, like, craft and, like, human taste and all of that that went into building it-

**Raiza Martin** [1:00:22]
Yeah

**Usama Bin Shafqat** [1:00:23]
... like, the knobs are not as easy to add as simply, like, I'm gonna add a parameter to this and it's gonna make it happen. It's like you kind of have to redo the quality process for, for everything.

**Raiza Martin** [1:00:34]
Yeah.

**Usama Bin Shafqat** [1:00:34]
So the prioritization is also different that way.

**Swyx** [1:00:36]
Mm.

**Raiza Martin** [1:00:37]
It goes back to sort of like it's a lot easier to do an eval for, like, the Deep Dive format than if, like, okay, now I'm gonna let you inject, like, these random-

**Swyx** [1:00:45]
Whatever

**Raiza Martin** [1:00:45]
... things, right?

**Swyx** [1:00:45]
Yeah.

**Raiza Martin** [1:00:45]
Okay, how am I gonna measure quality? Either I say, "Well, I don't care," because, like, you just input-

**Swyx** [1:00:50]
Yeah

**Raiza Martin** [1:00:50]
... whatever, or I say, "Actually, wait," right? Like, "I wanna help you get the best output ever. What's it going to take?"

**Usama Bin Shafqat** [1:00:56]
The knob actually needs to work reliably.

### The Future

**Raiza Martin** [1:00:58]
Yeah.

**Swyx** [1:00:59]
Yeah. A very important point. Two more things we definitely wanna talk about. I guess now people equivalate NotebookLM to, like, a podcast generator.

**Raiza Martin** [1:01:07]
Yeah.

**Swyx** [1:01:07]
But I guess, uh, you know, there's a whole product suite there.

**Raiza Martin** [1:01:10]
Yeah, yeah, yeah.

**Swyx** [1:01:10]
How should people think about that? Like, is this... A- and also, like, the future of the product as far as monetization too, you know?

**Raiza Martin** [1:01:17]
Yeah, yeah.

**Swyx** [1:01:17]
Like, is it gonna be... The voice thing gonna be a core to it? Is it just gonna be one output modality and, like, you're still looking to build, like, a broader kinda like-

**Raiza Martin** [1:01:25]
Yeah

**Swyx** [1:01:25]
... a interface with data and documents platform?

**Raiza Martin** [1:01:27]
And, I mean, that's such a, that's such a good question that I think the answer, it's I'm waiting to get more data. I think because we are still in the period where everyone's really excited about it, everyone's trying it, I think I'm getting a lot of sort of like positive feedback on the audio.

We have some early signal that says it's a really good hook, but people stay for the other features.

**Swyx** [1:01:49]
Mm.

**Raiza Martin** [1:01:49]
So that's really good too. I was making a joke yesterday. I was like, "It'd be really nice, you know, if it was just the audio, 'cause then I could just, like, simplify the train," right?

**Swyx** [1:01:58]
Mm.

**Raiza Martin** [1:01:58]
I don't have to think about all this other functionality. But I think the reality is that the framework, kind of like what we were talking about earlier that we had laid out, which is like you bring your own sources, there's something you do in the middle, and then there's an output, is a really extensible one, and it's a really interesting one, and I think, like- Particularly when we think about what a big business looks like, especially when we think about commercialization, audio is just one such modality.

But the editor itself, like the space in which you're able to do these things, is like that's the business, right? Like, maybe the audio by itself, not so much, but, like, in this big package, like, oh, I could see that.

I could see that being like a, a really big business.

**Swyx** [1:02:37]
Yep. Any thoughts on some of the alternative interact with data and documents thing like Claude Artifacts-

**Raiza Martin** [1:02:44]
Yeah

**Swyx** [1:02:44]
... like a ChatGPT Canvas.

**Raiza Martin** [1:02:46]
Yeah.

**Swyx** [1:02:46]
You know, kind of how do you see maybe where NotebookLM starts, where, like, Gemini starts, like y- you have so many amazing teams and products at Google that sometimes, like, I'm sure you have to figure that out.

**Raiza Martin** [1:02:56]
Yeah. Well, uh, I love Artifacts. I played a little bit with Canvas. I got a little dizzy using it. I was like, oh, it's like there's something... Well, you know, I- I like the idea of it fundamentally, but something about the UX was like, oh, this is, like, more disorienting than, like, Artifacts.

**Swyx** [1:03:11]
Mm.

**Raiza Martin** [1:03:11]
And I couldn't figure out what it was, and I didn't spend a lot of time thinking about it, but I love that, right? Like, the thing where you are like, "I'm working with, you know, an LLM, an agent, a chatbot, whatever, to create something new," and there's, like, the chat space, there's, like, the output space.

I love that. And the thing that I think I- I feel a little angsty about is, like- like, we've been talking about this for, like, a year, right? Like, ah, of course, like, I'm gonna say that.

**Swyx** [1:03:37]
Yeah.

**Raiza Martin** [1:03:37]
But it's like for... But, like, for a year now, I've had these, like, mocks that w- I was just like, "I wanna push the button." But we prioritized other things. We were like, "Okay, what can we, like, really win at?"

And, like, we prioritized audio, for example, instead of that. But just, like, when people were like, "Oh, what is this magic draft thing?" Oh, it's, like, 100%, right? It's, like, stuff like that that we wanna try to build into Notebook II.

And I'd made this comment on Twitter as well, where I was like, now I don't know, actually, right? I don't actually know if that is the right thing. Like, are people really getting utility out of this? I mean, from the launches, I...

it seems like people are really getting it. But I think now if we were to ship it, I have to rev on it, like, one layer more, right? I have to deliver, like, a, a differentiating value compared to, like, Artifacts or Canvas, which is, which is hard.

**Swyx** [1:04:20]
Which is because you've... you demonstrated the ability to fast follow, so you don't have to innovate every single time.

**Raiza Martin** [1:04:27]
I know. I know. I think for me it's just, like, the bar is high to ship, and when, when I say that, I think it's sort of like conceptually, like, the value that you deliver to the user. I mean, you'll, you'll see in NotebookLM there are a lot of corners that, like, that I have personally cut where it's like our UX designer is always like, "I can't believe you let, you let us ship with, like, these ugly scroll bars."

And I'm like, "No, no one noticed this, I promise." He's like, "No, every- everyone. This is a screenshot, this thing." But I, I mean, kidding aside, I- I think that's true, that it's like we do wanna be able to fast follow, but I think we wanna make sure that things also land really well, so the utility has to be there.

**Swyx** [1:05:00]
Code in, especially on our podcast, has a special place. Is code NotebookLM interesting to you? I haven't... I've never... Like, I don't see, like, a connect my GitHub to this thing, you know?

**Raiza Martin** [1:05:10]
Yeah, yeah. I- I think code, code is a big one.

**Swyx** [1:05:13]
Yeah.

**Raiza Martin** [1:05:13]
Code is a big one. I think we have been really focused, especially when we had, like, a much smaller team, we were really focused on, like, let's push, like, an end-to-end journey together. Let's prove that we can do that.

Because then once you lay the groundwork of, like, sources, do something in the chat, output, once you have that, you just scale it up from there, right? And it's like now it's just a matter of, like, scaling the inputs, scaling the outputs, scaling the capabilities of the chat.

So I think we're going to get there.

**Swyx** [1:05:38]
Mm.

**Raiza Martin** [1:05:38]
And now I also feel like I have a much better view of, like, where the investment is required, whereas previously I was like, "Hey, like, let's flesh out the story first before we put more engineers on this thing, because that's just going to slow us down."

**Usama Bin Shafqat** [1:05:50]
I mean, for what it's worth, the model still understands code.

**Raiza Martin** [1:05:52]
Yeah.

**Usama Bin Shafqat** [1:05:52]
So, like, I- I've seen at least one or two people just, like-

**Swyx** [1:05:56]
Chatting

**Usama Bin Shafqat** [1:05:56]
... download their GitHub repo, put it in there, and get, like, an audio overview of your code.

**Raiza Martin** [1:06:00]
Yeah. Yeah.

**Swyx** [1:06:01]
Oh, I've never tried that.

**Usama Bin Shafqat** [1:06:01]
This is like-

**Raiza Martin** [1:06:02]
That was crazy

**Usama Bin Shafqat** [1:06:02]
... these are all... how are all the files are connected together. 'Cause the model still understands code, like, even if we haven't, like-

**Raiza Martin** [1:06:08]
I think on, uh-

**Usama Bin Shafqat** [1:06:08]
... optimized for it

**Raiza Martin** [1:06:09]
... on sort of like the, the creepy side of things, I did watch a student, like with her permission, of course, I watched her do her homework in NotebookLM, and I didn't tell her, like, what kind of homework to bring, but she brought, like, her, um, computer science homework, and I was like, "Oh."

And she uploaded it, and she said, "Here's my homework. Read it." And it was just the instructions and, and, you know, NotebookLM was like, "Okay, I've read it." And, uh, the student was like, "Okay, here's my code so far," and she copy-pasted it from the editor, and she was like, "Check my homework."

And NotebookLM was like, "Well, number one is wrong." And I thought that was really interesting 'cause it didn't tell her what was wrong, it just said it's wrong. And she was like, "Okay, don't tell me the answer, but, like, walk me through, like, how you'd think about this."

And it was... what was interesting for me was that she didn't ask for the answer, and I asked her, I was like, "Oh, why did you do that?" And she was like, "Well, I actually wanna learn it." She was like, "'Cause I'm gonna have to take a quiz on this at some point."

And I was like, "Oh, yeah, that's a, that's a really good point." And it was interesting because, you know, NotebookLM, while the formatting wasn't perfect, like, did say, like, "Hey, have you thought about using, you know, maybe an integer instead of like this?"

And so that was, that was really interesting.

**Swyx** [1:07:17]
Mm.

**Usama Bin Shafqat** [1:07:17]
Are you adding, like, real-time chat on the output? Like, you know, there's kinda like the, uh, Deep Dive show, and then there's like the-

**Raiza Martin** [1:07:23]
Yeah

**Usama Bin Shafqat** [1:07:24]
... listeners call in-

**Raiza Martin** [1:07:25]
Yeah. Yeah

**Usama Bin Shafqat** [1:07:25]
... and say, "Hey."

**Raiza Martin** [1:07:26]
Yeah. We're actively... that's one of the things we're actively prioritizing. Actually, one of the interesting things is now we're like, why would anyone wanna do that, right? Like, what are the actual... like, kind of going back to sort of having a strong POV about the experience, it's like what is better, like, what is fundamentally better about doing that that's not just like being able to Q&A your Notebook?

How is that different from, like, a conversation? Is it just the, the fact that, like, there was a show and you wanna tweak the show? Is it because you want to participate? So I think there's a lot there that, like, we can continue to unpack, but yes, like, that's coming.

**Swyx** [1:07:58]
It's because I formed a parasocial relationship with-

**Raiza Martin** [1:08:01]
Yeah.

**Usama Bin Shafqat** [1:08:01]
Yeah

**Swyx** [1:08:02]
... your two-voice system.

**Raiza Martin** [1:08:02]
I just wanna be part of your life.

**Usama Bin Shafqat** [1:08:04]
Get this.

**Raiza Martin** [1:08:06]
Totally.

**Swyx** [1:08:07]
Uh, yeah, but it is obviously because OpenAI has just launched a real-time chat, it- it's a very-

**Raiza Martin** [1:08:12]
Yeah

**Swyx** [1:08:12]
... hot topic.

**Raiza Martin** [1:08:13]
Mm.

**Swyx** [1:08:13]
I would say one of the toughest AI engineering disciplines out there, because even their API doesn't do interruptions that well, to be honest, and, you know... Yeah, so real-time chat is, is, is tough.

**Raiza Martin** [1:08:25]
I love that thing.

**Swyx** [1:08:26]
Yeah.

**Raiza Martin** [1:08:26]
I love, I love it. Yeah.

### Outro

**Swyx** [1:08:27]
Okay. So, uh, we have, uh, a couple ways to end, either call to action or laying out one principle of AI PM-ing or engineering that you, that you really think about a lot. Is there anything that comes to mind?

**Raiza Martin** [1:08:39]
I feel like that's a test. Of course, I'm gonna say go to notebooklm-

**Swyx** [1:08:43]
Uh-huh

**Raiza Martin** [1:08:43]
... .google.com.

**Swyx** [1:08:43]
Uh-huh.

**Raiza Martin** [1:08:44]
Try it out, join the Discord, and tell us what you think.

**Swyx** [1:08:47]
Yeah. E- especially, like, when you have a technical audience, what, what do you want from a technical engineering audience?

**Raiza Martin** [1:08:52]
I mean, I think it's interesting because the technical and engineering audience typically will just say, "Hey, where's the API?"

**Swyx** [1:08:58]
Yeah.

**Raiza Martin** [1:08:59]
But, you know, and I think we addressed it, but I think what I, what I would really, uh, be interested to discover is, is this useful to you? Why is it useful? What did you do, right? Is it useful tomorrow?

How about next week? Just the most useful thing for me is if, if you do stop using it or if you do keep using it, tell me why. Because I think contextualizing it within your life, right, your background, your motivations, like it is what really helps me build really cool things.

**Swyx** [1:09:22]
And then one piece of advice for AI PMs.

**Raiza Martin** [1:09:25]
Okay. If I had to pick one, it's just always be building. Like, build things yourself. I think, like, for PMs, it's, like, such a critical skill, and just, like, take time to, like, pop your head up and see what else is new out there.

On the weekends, I try to have a lot of discipline. Like, I only use ChatGPT and, like, Claude on the weekend. I try to, like, use, like, the APIs. Occasionally I'll, I'll try to build something on, like, GCP over the weekend 'cause, like, I don't, I don't do that normally, like, at work, but it's just, like, the rigor of just trying to be, like, a builder yourself and even just, like, testing, right?

Like, you could have a, an idea of, like, how a product should work and maybe your engineers are building it, but it's like, what was your, like, proof of concept, right? Like, what gave you conviction, like that was the right thing?

**Swyx** [1:10:06]
Um, call to action?

**Usama Bin Shafqat** [1:10:07]
Yeah. I feel like consistently, like, the, the most magical moments out of, like, AI building come about for me when, like, I'm really, really, really just close to the edge of the model capability, and sometimes it's, like, farther than you think it is.

Like, I think while building this product, some of the other experiments, like, there were phases where it was, like, easy to think that you've, like, approached it, but, like, sometimes at that point what you really need is to, like, show your thing to someone and, like, they'll come up with creative ways to improve it.

Like, it's... We're all sort of, like, learning, I think. So yeah, like, I, I feel like unless you're hitting that bound of, like, this is what Gemini 1.5 can do, probably, like, the magic moment is, like, somewhere there, like, in that sort of, um, limit.

**Swyx** [1:10:49]
So push the edge of the capability.

**Usama Bin Shafqat** [1:10:51]
Yeah, totally.

**Alessio** [1:10:52]
It's funny because we had, uh, Nicolas Carlini from DeepMind on the pod and he was like, "If the model is always successful, you're probably not trying hard enough to, like, give it hard-"

**Usama Bin Shafqat** [1:11:00]
Right

**Alessio** [1:11:01]
... things. So, um, yeah, to-

**Swyx** [1:11:03]
My, my problem is, like, sometimes I'm, I'm not smart enough to challenge it.

**Alessio** [1:11:06]
Yeah, right. That's, uh, that's...

**Raiza Martin** [1:11:08]
Well, I think, I think, like, that's... I hear that a lot. Like, people are always like, "I don't know how to use it."

**Swyx** [1:11:14]
I run out. Yeah.

**Raiza Martin** [1:11:14]
Yeah, and it's, and it's hard. Like, I remember the first time I used Google Search, I was like, "What do I type?" My dad was like, "Anything." It's like, "Anything? I got nothing in my brain, Dad." Like, what do you mean?

And, and I think there's a lot of, like, for product builders, is like have a strong opinion about, like, what is the user supposed to do?

**Swyx** [1:11:30]
Yeah.

**Raiza Martin** [1:11:31]
Help them do it.

**Swyx** [1:11:32]
Principle for AI engineers or, like, just one advice that you have others?

**Usama Bin Shafqat** [1:11:37]
Um, I guess, like, in addition to pushing the bounds, and to do that, that often means, like, you're not gonna get it right in the first go. So, like, don't be afraid to just, like, batch multiple models together.

I guess that's, I'm basically describing an agent, but more thinking time equals just better results consistently, and that holds true for probably every-

**Swyx** [1:11:59]
Rebuild

**Usama Bin Shafqat** [1:11:59]
... single time that I've tried to build something.

**Swyx** [1:12:01]
Well, at some point we'll talk about the sort of longer inference paradigm. Uh, it seems like DeepMind is rumored to be coming out with something. You can't comment, of course. Yeah. But, well, thank you so much.

**Usama Bin Shafqat** [1:12:11]
Yeah.

**Swyx** [1:12:11]
You know, you, you've created, um... I, I actually said, I think you saw this, I think that NotebookLM was kind of like the ChatGPT moment-

**Raiza Martin** [1:12:18]
Oh, yeah

**Swyx** [1:12:18]
... for Google.

**Raiza Martin** [1:12:18]
That was so, that was so crazy when I saw that. I was like, "What?" Like, ChatGPT was huge for me, and I think, you know, when pe- when you said it and other people have said it, I was like, "Is it?"

**Swyx** [1:12:28]
Yeah.

**Raiza Martin** [1:12:28]
Like, that's, that's crazy.

**Swyx** [1:12:29]
Yeah.

**Raiza Martin** [1:12:29]
That's so cool.

**Swyx** [1:12:29]
People weren't, weren't, like, really cognizant of NotebookLM before and, and audio overviews, and, and NotebookLM, like, unlocked the, uh, you know, a, a use case for people in the way that I would go so far as to say cloud projects never did, and, uh, I don't know, you know...

I, I think a lot of it is composite PM-ing and engineering, uh, but also just... You know, it's, it's interesting how a lot of these projects are always, like, low-key research previews. For you, it's like you're, you're a separate org but, like, you know, you built, um, products and UI innovation on top of also imp- working with research to improve the model.

That was a success. That, that wasn't planned to be this whole big thing. You know, your TPUs were on fire, right? Like, when-

**Raiza Martin** [1:13:06]
Oh, my gosh. That was so funny. I didn't know people would, like, really catch onto the Elmo fire, but it, it was just, like, one of those things where I was like- ... you know, we had to ask for more TPUs.

**Usama Bin Shafqat** [1:13:17]
Sorry, I cut.

**Raiza Martin** [1:13:17]
Yeah. We... Many times and, you know, it was a little bit of a, of a sub tweet of like, "Hey, reminder, give us more TPUs on here."

**Swyx** [1:13:25]
Yeah. It's weird. Like, I just think, like, when people try to make big launches, then they flop, and then, like, when they're not trying and they're just, they're just trying to build a good thing, then, then they succeed.

It's, it's this fundamentally really weird magic that I, I haven't really encapsulated yet, but you've, you've done it.

**Raiza Martin** [1:13:40]
Oh, thank you.

**Swyx** [1:13:40]
So congrats.

**Raiza Martin** [1:13:41]
Thank you. And, you know, I think we'll just keep going in, like, the same way. We just keep trying, keep trying to make it better.

**Swyx** [1:13:46]
Yeah, I hope so, hope so. All right.

**Usama Bin Shafqat** [1:13:47]
Cool.

**Swyx** [1:13:48]
Thank you.

**Raiza Martin** [1:13:48]
Thank you. Thanks for having us.

**Usama Bin Shafqat** [1:13:49]
Okay, thanks.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
