# The Agent Reasoning Interface: Claude, ChatGPT Canvas, Tasks, Operator — with Karina Nguyen, OpenAI

Latent Space · 2025-02-01

<https://addtry.com/89f02fe7-9c47-40fa-b812-1c07dd6d8853>

OpenAI's Karina Nguyen, who led the creation of ChatGPT Canvas and Tasks, explains how her team trains separate models for new interaction paradigms rather than prompting the core model, enabling rapid iteration on user feedback. She details the 'behavioral design' process that shapes model personality in collaborative contexts like Canvas—deciding when to rewrite vs. edit—and how that extends to defining agents as a progression from one-off actions to fully trustworthy long-horizon delegation. Drawing on her time at Anthropic building Claude.ai from scratch and co-creating Claude 3, she contrasts the product mindsets: OpenAI takes more product risks across consumer features while Anthropic focuses on enterprise. She also shares her vision for ChatGPT evolving into a 'generative OS' where UIs adapt dynamically to user intent, and argues that human creativity—not model capability—is the current bottleneck for rethinking software interfaces.

## Questions this episode answers

### What is 'behavioral design' for AI models according to Karina Nguyen?

Karina Nguyen describes behavioral design as extending product design into model design, shaping model behavior for specific contexts like collaboration. It involves defining core values (e.g., honesty, harmlessness, helpfulness) and balancing conflicts between them. She says it is more art than science, using synthetic data to create scenarios that teach the model appropriate behavior, similar to character creation in games.

[22:09](https://addtry.com/89f02fe7-9c47-40fa-b812-1c07dd6d8853?t=1329000)

### How did the OpenAI team develop ChatGPT Canvas, and what role did model post-training play?

The Canvas team began with a prompted baseline of ChatGPT, then identified edge cases that required post-training. They retrained the entire 4o model with Canvas-specific data to handle behaviors like when to write comments, edit sections, or trigger Canvas. Deploying as a separate model allowed rapid iteration on user feedback before integrating into the core model, with the project taking about four months.

[27:37](https://addtry.com/89f02fe7-9c47-40fa-b812-1c07dd6d8853?t=1657000)

### How does Karina Nguyen envision the progression of AI agents from simple tasks to full delegation?

Nguyen sees agents progressing gradually: from one-off actions, to collaboration (like Canvas), to fully trustworthy long-horizon delegation. Trust is built through collaborative interactions, where consistency and reliability are demonstrated over time, mirroring how humans build working relationships before delegating sensitive tasks. She emphasizes that collaboration is a crucial milestone toward full delegation.

[49:09](https://addtry.com/89f02fe7-9c47-40fa-b812-1c07dd6d8853?t=2949000)

## Key moments

- **[0:00] Intro**
  - [0:27] OpenAI's Karina Nguyen leads team creating new interaction paradigms like ChatGPT Canvas and Tasks.
  - [1:34] Canvas coding improvements require RL training and new synthetic data generation methods.
  - [3:24] Karina Nguyen discovered Anthropic through Chris Olah's interpretability work.
  - [4:28] Karina Nguyen was rejected first time applying to Anthropic, then hired as first designer/frontend engineer.
- **[5:25] Anthropic Roots**
  - [5:30] Anthropic's first product was Claude in Slack, sunsetted later.
  - [7:16] Karina Nguyen built the first Claude.ai interface in two weeks after ChatGPT launched.
  - [7:43] Karina Nguyen wrote Claude.ai's first 50,000 lines of code without any reviews.
  - [8:27] Q: Why did Anthropic not launch ChatGPT first? A: Leadership lacked conviction because Claude 1.3 had hallucinations.
  - [9:11] Canvas and Tasks ideas could have existed two years ago but AI landscape was too nascent for UX focus.
  - [9:25] Karina Nguyen prototyped Claude Workspace in 2023 inspired by Harry Potter's Tom Riddle diary.
- **[11:30] Claude 3**
  - [11:34] Karina Nguyen worked on post-training for Claude 3 Haiku and wrote the model card.
  - [12:13] "Every model will have its own brain damage" — Karina Nguyen on fine-tuning variations.
  - [14:34] Anthropic first to publish GPQA benchmark numbers, highlighting high variance in evals.
  - [15:36] Model card evals never apples-to-apples; labs use different prompt settings and formatting.
  - [16:47] Stanford's Helm benchmark misprompted Claude, leading to lower scores.
  - [18:01] o1 prompting tip: Give hard constraints to help model filter candidates.
- **[21:57] Model Personality**
  - [22:01] Model behavioral design: shaping AI personality for collaboration, not just chat.
  - [24:03] Balancing honesty, helpfulness, harmlessness in model behavior is more art than science.
  - [26:58] OpenAI's Model Design team, led by Jan, shapes ChatGPT writing and Canvas collaboration.
- **[27:37] Canvas**
  - [28:03] Canvas project started on July 4th when Karina pitched to Barrett Zoph, forming new team.
  - [30:43] Canvas shipped as separate model to iterate on feedback, then integrated into core.
  - [31:56] ChatGPT writing quality improved for nonfiction emails and cover letters with targeted data.
  - [34:38] Canvas writing improvements now in core 4o, but slight differences remain between API and ChatGPT.
  - [36:30] Canvas inverts Google Docs paradigm: AI-first with document on side.
  - [37:35] Canvas to evolve into full IDE for writing and coding, with code execution.
- **[41:45] Tasks**
  - [41:54] Tasks project led by resident Vivek, developed in under 2 months, mirroring Canvas model.
  - [46:28] Nguyen predicts ChatGPT will become proactive, suggesting recurrent tasks based on user habits.
  - [49:10] Nguyen defines agents: one-off actions → collaboration → trustworthy long-horizon delegation.
  - [52:10] Computer use is core agent capability; vision model improvements make it feasible now.
- **[52:23] Computer Agents**
  - [57:05] Nguyen predicts website clicks will decline as people access internet via AI models.
  - [57:16] Swyx's benchmark for computer use: automate expense reports across multiple apps.
  - [59:29] Next-gen ChatGPT output: dynamic React apps and visualizations tailored to user.
- **[1:00:21] Two Cultures**
  - [1:00:21] OpenAI takes more product risks; Anthropic more enterprise-focused.
  - [1:03:00] OpenAI offers more creative freedom for researchers to propose and resource new directions.
- **[1:03:46] Outro**
  - [1:04:33] Nguyen urges designers to play more with models to spark creativity for new AI interfaces.

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **Karina Nguyen** (guest)

## Topics

Agents, Startups

## Mentioned

Anthropic (company), OpenAI (company), Canvas (product), Claude (product), GPT-3 (product), GPT-4 (product), Gemini (product), Google Docs (product), Operator (product), Slack (product), Tasks (product), o1 (product)

## Transcript

### Intro

**Alessio** [0:05]
Hey, everyone. Welcome to the Late in Space podcast. This is Alessio, partner and CTO at Decibel, and I'm joined by my usual co-host, Swyx.

**Swyx** [0:12]
Hey, and today we're very, very blessed to have Karina Nguyen in y- in the studio. Welcome.

**Karina Nguyen** [0:16]
Nice to meet you.

**Alessio** [0:17]
We finally made it happen.

**Karina Nguyen** [0:19]
Yes.

**Swyx** [0:19]
My-- finally made it happen. Uh, first time we tried this, you were s- working at a different company. And now, now we're here. Fortunately, you had some time, so thank you so much for joining us.

**Karina Nguyen** [0:26]
Yeah, thank you for inviting me.

**Swyx** [0:27]
Uh, Karina, you-- now y-your website says you lead a research team in OpenAI creating new interaction paradigms for reasoning interfaces and capabilities like ChatGPT Canvas and most recently ChatGPT Tasks. I don't know, is, is that what we're calling it?

Uh, streaming chain of thought for o1 models and more via novel synthetic model training. What is this research team?

**Karina Nguyen** [0:46]
Yeah, I need to, like, clarify this a little bit more.

**Swyx** [0:48]
Okay.

**Karina Nguyen** [0:49]
I think it changed a lot, like, since l- the last time we launched. So we launched Canvas, and it was, like, what? The first, like, project that I was a tech lead, basically. And then I think over time I was, like, trying to refine what my team is, and I feel like it's at an intersection of, like, human-computer interaction, defining what the next interaction paradigms might look like with some of the most recent, like, reasoning models, as well as actually trying to come up with, like, novel methods, how to improve those models for certain tasks that we want to.

So for Canvas, for example, one of the most common use cases is basically writing and coding and, like, we continually working on, like, okay, like, how do we make Canvas coding to go beyond what is possible right now?

And, like, that requires us to actually do, like, RL training and, like, coming up with, like, new methods of, like, synthetic data generation. The way I'm, like, thinking about it is that, like, my team is going from, like, very full stack from, like, training models all the way up to, like, deployment and, like, making sure that we create novel, like, product features that is coherent to what ChatGPT can become.

There are different types of, like, features like Canvas, Tasks, but all those components that go... They compose together to evolve ChatGPT into something completely new, I think, in the new year.

**Swyx** [2:09]
It's evolving. I, I liked your tweet about that. It's, like, kind of modular. You can compose it with the stocks feature, the creative writing feature.

**Karina Nguyen** [2:17]
Mm-hmm.

**Swyx** [2:17]
I forget what else. We have a list of other use cases, but we don't have to go into that yet.

**Alessio** [2:21]
Yeah. Can we maybe go back to when you first started working with LLMs? I know you had some early UX prototypes with-

**Karina Nguyen** [2:29]
Mm-hmm

**Alessio** [2:29]
... GPT-3 as well and kind of, like, maybe how that has-

**Karina Nguyen** [2:31]
Mm-hmm

**Alessio** [2:31]
... informed the way you build products.

**Karina Nguyen** [2:33]
I think my background was mostly, like, working on computer vision applications for, like, investigative journalism, uh, back when I was, like, at school at Berkeley, and I was working a lot with, like, Human Rights Center and, like, investigative journalists from r- various media, and that's how I learned more about, like, AI, like, with vision transformers.

And at that time I was working with some of the professors at Berkeley Eye Research. But-

**Swyx** [3:01]
There are some Pulitzer Prize-winning professors, right? That, that teach there.

**Karina Nguyen** [3:05]
No.

**Swyx** [3:05]
Like-

**Karina Nguyen** [3:05]
So it's mostly-

**Swyx** [3:06]
Okay

**Karina Nguyen** [3:06]
... like, was reporting for, like, teams like The New York Times, like-

**Swyx** [3:10]
Yeah

**Karina Nguyen** [3:10]
... a, the AP, Associate Press. So it was, like, all in the context of, like, Human Rights Center.

**Swyx** [3:15]
Got it.

**Karina Nguyen** [3:15]
Yeah. So that was, like, in computer vision, and then I saw Chris Olah's work around, you know, like-

**Swyx** [3:24]
Interoperability

**Karina Nguyen** [3:25]
... compatibility-

**Swyx** [3:25]
Yeah

**Karina Nguyen** [3:25]
... from Google, and that's how I found out about, like, Anthropic. And at that time I was just, like... I think it was, like, the year when, like, Ukraine's war happened, and I was, like, trying to find a full-time job, and it was kinda, like, all got distracted.

It was, like, kinda, like, spring, and I was, like, very focused on, like, figuring out, like, what to do. And then my best option at that time was just, like, continue my internship at The New York Times and convert to, like, full-time.

At The New York Times, it was just, like, working on, like, mostly, like, product engineering work around, like, R&D prototypes, kinda, like, storytelling features on the mobile experience, so, like, kinda, like, storytelling ex-experiences. And, like, at that time we were, like, thinking about, like, how do we employ, like, NLP techniques to, like, scrape some of the archives from The New York Times or something.

But then I always wanted to, like, get into, like, AI, and, like, I knew OpenAI for a while, like, since I was, like, in Berkeley. And yeah, so I kinda, like, applied to Anthropic just on the website, and I was rejected the first time.

But then at that time they were not hiring for, like, anything, like, product engineering or frontend engineering, which was something I was, like-- at that time I was, like, interested in. And then, um, there was, like, a new opening at Anthropic.

It was, like, kinda like for U-frontend engineer, and so I applied, and that's how my journey began. But, like, the earlier prototypes was mostly, like, I used, like, Clip for, like, fashion recommendation search. That was, like, one of those-

**Alessio** [4:53]
I saw that project

**Karina Nguyen** [4:54]
... successful projects, I think.

**Swyx** [4:55]
Yeah.

**Karina Nguyen** [4:55]
And I was like, uh, before even coming to Anthropic, I was, like, thinking maybe I should just, like, do my own startup. But I, I, I feel like I didn't have, like, enough confidence and conviction in myself that I could do that.

But it was, like, one of the early, like, prototypes. And I think Twitter is a good platform to, like-- for side projects-

**Swyx** [5:11]
That's fantastic

**Karina Nguyen** [5:12]
... that helped me to, like, have visibility.

**Swyx** [5:12]
Especially something visual.

**Karina Nguyen** [5:14]
Yeah.

**Swyx** [5:14]
Yeah. We'll briefly mention that the Ukrainian crisis actually hit home more for you than most people because you're from the Ukraine, and you, you moved here, like, for school, I guess.

**Karina Nguyen** [5:23]
Yeah.

**Swyx** [5:24]
Yeah, yeah. We'll come back to that if it comes up. Uh, but then you joined Anthropic, uh, not just as a frontend engineer. You were the first.

### Anthropic Roots

**Alessio** [5:30]
Is that true? It's like-

**Karina Nguyen** [5:31]
Designer.

**Swyx** [5:32]
Yeah

**Karina Nguyen** [5:33]
Yes. I think, like, I did both product design and front-end engineering together, and like at that time it was like pre-ChatGPT. It was like I think August 2022. And that was a time when Anthropic really decided to like do more product-y related things, and the vision was like, "We need to like fund research, and like building product is like the best way to like fund safety research," which I found it quite admirable.

So the really first product that Anthropic built was like Claude and Slack. And it was sunsetted not long after, but like it was like one of the first... I think I still come back to that idea of like Claude operating inside some of the organizational workplace like Slack, and it...

something magical in there. And I remember we built like ideas like summarize the thread. But you can like imagine having automated like ways of like maybe Claude should like summarize multiple channels every week custom for what you like, uh, or for what you want.

And then we built some like really cool features like tag Claude and then ask to summarize what's- what happened in a thread, suggest like new ideas. But we didn't quite double down because you could like imagine like Claude having access to like the files or like Google Drive that you can upload in Sl- uh, just connectors, like connections in the Slack.

Also, the UX was kind of constraining. At that time I was like thinking like, "Oh, we wanted to do this feature," but like Slack interface kind of like constrained us to like do that, and we didn't want to like be dependent on the platform like Slack.

And then after like ChatGPT came out, I remember the first two weeks my manager made me this challenge like can I like reproduce kinda like a similar interface in like two weeks? And one of the early mistakes being in engineering is like I said yes.

Instead I should have said like, you know, it's double, 2X the time.

**Swyx** [7:35]
Sure.

**Karina Nguyen** [7:36]
Um, and this is how like Claude.ai was kind of like born.

**Swyx** [7:40]
Oh, so you actually wrote Claude.ai as your first job?

**Karina Nguyen** [7:43]
Yeah, like I think like- ... the first like 50,000 code lines-

**Swyx** [7:48]
Yeah, yeah

**Karina Nguyen** [7:48]
... without any reviews at that time- ... because there was no one. Um, yeah, it was like a very small team. It was like six, seven team who we were called like Deployment Team. Yeah.

**Swyx** [8:00]
Oh, my- I actually interviewed for, uh, Anthropic at, around that time. I got- I was given Claude in Sheets.

**Karina Nguyen** [8:05]
Oh, cool. Yes.

**Swyx** [8:05]
And that was my other form factor. I was like, "Oh, yeah, this needs to be in a table so we can, we can just copy-paste-

**Karina Nguyen** [8:10]
Yeah

**Swyx** [8:11]
... and just span it out," uh, which is kinda cool. The other rumor that, um, we might as well just mention this, um, Raza Habib from Human Loop, uh, often says that, uh, you know, there was some- there's some version of ChatGPT in Anthropic.

**Karina Nguyen** [8:23]
Mm.

**Swyx** [8:23]
Like, you had the chat interface already. Like, you had Slack.

**Karina Nguyen** [8:26]
Mm-hmm.

**Swyx** [8:27]
Why not launch a web UI? Like, basically like how did, how did OpenAI beat Anthropic to ChatGPT, basically?

**Karina Nguyen** [8:33]
Mm. Well, at that time-

**Swyx** [8:35]
Like it seems kind of obvious to have it

**Karina Nguyen** [8:36]
... I think ChatGPT model itself came out way before than we decided to like launch Claude 2 necessarily.

**Swyx** [8:42]
Okay. Mm.

**Karina Nguyen** [8:42]
And I think like at that time, Claude 1.3 had a lot of hallucinations, actually. So I think there was like one of the concerns is like I don't think like the leadership was conven- con- had the conviction that this is the model that you need to like- you want to like deploy or something.

So there was a lot of discussions around, around that time.

**Swyx** [8:59]
Cool.

**Karina Nguyen** [8:59]
But Claude 1.3 was like I don't know if you played with that, but it's like extremely creative and was like really cool.

**Swyx** [9:08]
Nice. It's still creative.

**Alessio** [9:11]
And you had a tweet recently that you said things like Canvas and Task could have-

**Karina Nguyen** [9:14]
Mm

**Alessio** [9:14]
... happened two years ago-

**Karina Nguyen** [9:15]
Mm-hmm

**Alessio** [9:15]
... but they were not. Do you know why they were not? Was it too many researchers at the labs not focused on UX? Was it just not a priority for the labs?

**Karina Nguyen** [9:25]
Yeah. I come back to that question a lot. I guess like I was working on something similar to like Canvas-y but for Claude at that time in like 2023. It was the same similar idea of like Claude Workspace where a human and a Claude could have like a shared work, workspace, um-

**Swyx** [9:43]
Yeah, and that's Artifacts

**Karina Nguyen** [9:43]
... which would document.

**Swyx** [9:44]
Right.

**Alessio** [9:44]
No, no, no.

**Karina Nguyen** [9:45]
It kinda, I don't know if that's-

**Alessio** [9:45]
This is Claude Projects.

**Karina Nguyen** [9:47]
I don't know. I think it kinda evolved. I think like at that time I was like in product engineering team, and then I switched to like research team, and the product engineering team grew so much. They had their own ideas of like Artifacts and like Projects-

**Alessio** [10:01]
Oh

**Karina Nguyen** [10:01]
... and not necessarily... Maybe they ha- they looked at my like previous explorations, but like, you know, when I was exploring like Claude Documents or like Claude Workspace was like I don't think anybody was thinking about UX as much or like not many like researchers understood that.

And I think the inspiration actually for-- I still have like all the sketches, but the inspiration was like from the Harry Potter, like Tom Riddler-

**Swyx** [10:26]
Mm-hmm. Yeah, yeah, yeah

**Karina Nguyen** [10:27]
... uh, diary. That was inspirational like having, um, Claude writing into the document or something and communicate back.

**Swyx** [10:34]
Yeah, uh, so like in the movie you write a little bit and then it-

**Karina Nguyen** [10:37]
Yeah

**Swyx** [10:37]
... answers you.

**Karina Nguyen** [10:38]
Yeah.

**Swyx** [10:38]
Okay. Interesting.

**Karina Nguyen** [10:40]
But that was like in the-- only in the context of like writing. I think Canvas is like more-- also serves like coding, one of the most common use cases. But yeah, I think like those, those ideas could have happened like two years ago, just like maybe...

I don't think it was like a priority at that time. It was like very unclear. I think like AI landscape at that time was very nascent-

**Swyx** [11:03]
Mm

**Karina Nguyen** [11:03]
... if that makes sense. Like nobody like... Even when I would talk to like some of the designers at that time, like product designers, they were not even thinking about that at all. They did not have like AI in mind and like...

It's kinda interesting. Except for one of my designer friends.

**Swyx** [11:17]
Yeah.

**Karina Nguyen** [11:17]
His name is Jason Young. Yeah. Who was thinking about that.

**Swyx** [11:20]
And Jason now is at New Computer.

**Karina Nguyen** [11:22]
Yes.

**Swyx** [11:22]
We'll have them on at some point. I had them speak at my first summit, and we're s- you're speaking at the second one, which will be really fun.

**Karina Nguyen** [11:27]
Nice.

**Swyx** [11:28]
We'll stay on Anthropic for a bit, and then we'll move on to more recent things. I think the other big project that you were, that you were involved with was just Claude 3.

### Claude 3

**Karina Nguyen** [11:34]
Mm-hmm.

**Swyx** [11:35]
Just tell us the story. Like, what's it like to launch one of the biggest launches of the year?

**Karina Nguyen** [11:39]
Yeah. I think like I was- So Claude 3-

**Swyx** [11:44]
This is Haiku, Sonnet, Opus all at once, right?

**Karina Nguyen** [11:46]
Yes.

**Swyx** [11:46]
Yeah.

**Karina Nguyen** [11:47]
It was a Claude 3 family. I was a part of the post-training fine-tuning team. We only had, like, what? Like, 10, 12 people involved, and it was really, really fun to, like, work together as friends. So yeah, I was mostly involved in, like, Claude 3 Haiku post-training side and then evaluations, like developing new evaluations and, like, literally writing the entire, like, model card, and I had a lot of fun.

I think, like, the way you train the model is, like, very different obviously, but I think what I've learned is that, like, you will end up with, like, I don't know, like, 70 models, and every model will have its own, like, brain damage.

And, like, so many things will, like, like, kind of debug.

**Swyx** [12:28]
Like, personality-wise or, or-

**Karina Nguyen** [12:31]
Um

**Swyx** [12:31]
... or performance benchmarks?

**Karina Nguyen** [12:32]
Like, I think every model is very different.

**Swyx** [12:33]
Okay.

**Karina Nguyen** [12:33]
And I think, like, it's like one of the interesting, like, research questions is like how do you understand, like, the data interactions as you, like, train the model? It's like if you train the model on, like, contradictory data sets, how can you make sure that there won't be, like, any, like, weird, like, side effects?

And sometimes you get, like, side effects. And, like, the learning is that you have to, like, iterate very rapidly and, like, have to, like, debug and detect it and make... like, address it with, like, interventions. And actually, some of the techniques from, like, software engineering is, was very, like, useful here.

It's like how do you debug code?

**Swyx** [13:07]
Gift for data.

**Karina Nguyen** [13:09]
Yeah, exactly.

**Swyx** [13:10]
So I really empathize with this because datasets, if you put in the wrong one, you can basically kind of screw up, like, the, the-

**Karina Nguyen** [13:16]
Mm-hmm

**Swyx** [13:16]
... the past month of training. The problem with this for me is the, the existence of YOLO runs.

**Karina Nguyen** [13:21]
Mm-hmm.

**Swyx** [13:21]
I cannot square this with YOLO runs. If you're telling me, like, you're taking such care about datasets, then every day I'm gonna check in, run evals and do that stuff.

**Karina Nguyen** [13:28]
Mm-hmm.

**Swyx** [13:29]
But then we also know that YOLO runs exist.

**Karina Nguyen** [13:31]
Yes.

**Swyx** [13:32]
So how do you square that?

**Karina Nguyen** [13:33]
Well, I think it's, like, dependent on how much compute you have, right?

**Swyx** [13:37]
Yes.

**Karina Nguyen** [13:37]
So it's like, it's actually a lot of questions and, like, research around, like, how do you most effectively use the compute that you have? And maybe you can have, like, two to three runs that is only, like, YOLO runs.

**Swyx** [13:51]
Mm-hmm.

**Karina Nguyen** [13:51]
But if you don't have a luxury of that, like, you kind of need to, like, prioritize, uh, ruthlessly and, like, what are the experiments that are most important to, like, run?

**Swyx** [14:00]
Yeah.

**Karina Nguyen** [14:00]
I think this is what, like, research management is basically, is like how do you, um-

**Swyx** [14:05]
Funding efforts, yeah.

**Karina Nguyen** [14:06]
Yeah, like-

**Swyx** [14:07]
Prioritizing

**Karina Nguyen** [14:07]
... take, like, research bets and make sure that you build the conviction in those bets rapidly such that if they work out, you, like, double down on them.

**Swyx** [14:16]
Yeah. You, you almost have to, like, kind of ablate datasets too and-

**Karina Nguyen** [14:18]
Yeah

**Swyx** [14:18]
... and, like, do it on the side channel and then merge it in. Yeah, it's kind of super interesting. So tell us more, like, what's your favorite... So you, you, um... I'm... I have this in front of me, the model card.

**Karina Nguyen** [14:27]
Mm-hmm.

**Swyx** [14:27]
You say constructing this painful, this table was slightly painful.

**Karina Nguyen** [14:30]
Mm-hmm.

**Swyx** [14:30]
Just pick a benchmark and what's an interesting story behind one of them?

**Karina Nguyen** [14:34]
I would say GPGA was kind of interesting. I think it was, like, the first-- I think we were the first lab, like, Anthropic was the first lab to, like, run, to, like, publish-

**Swyx** [14:43]
Oh, 'cause it was, like, relatively new after NeuriPS?

**Karina Nguyen** [14:45]
Yeah, yeah.

**Swyx** [14:46]
Okay.

**Karina Nguyen** [14:46]
Publish GPGA, like, numbers. And I think one of the things that we've learned was that I personally learned about that, like, any evals is like, some evals are, like, very, like, high variance, and, like, GPA is, like, happened to be, like, a huge, like, high variance, like, evaluation.

So, like, one thing that we did is, like, having, like, run the average of, like-

**Swyx** [15:09]
Mm-hmm

**Karina Nguyen** [15:10]
... five and, like, s- take the average. But, like, the hardest thing about, like, the model card is, like, none of the numbers are, like, apples to apples.

**Swyx** [15:19]
Yes.

**Karina Nguyen** [15:19]
So you actually need to, like-

**Swyx** [15:20]
Will knows this.

**Karina Nguyen** [15:20]
... go back to, like, I don't know, like, GPT-4 model card-

**Swyx** [15:25]
Uh-huh

**Karina Nguyen** [15:25]
... and, like, read the appendix just to, like, make sure that, like, the settings are the same as you're running the settings too. So it's, like, never an apples to apples.

**Swyx** [15:35]
Yeah.

**Karina Nguyen** [15:36]
But it's interesting how, like, you know, when you market models as products, like, customers don't necessarily know-

**Swyx** [15:44]
Yeah

**Karina Nguyen** [15:44]
... like, um-

**Swyx** [15:44]
They're just like, "My MMLU is 99. What do you mean?"

**Karina Nguyen** [15:48]
Yeah, exactly. Um.

**Swyx** [15:49]
Why isn't there an industry standard harness, right? There's this Eleutheris thing-

**Karina Nguyen** [15:53]
Mm-hmm

**Swyx** [15:53]
... which it seems like none of the model labs use.

**Karina Nguyen** [15:55]
Mm-hmm.

**Swyx** [15:56]
And then OpenAI put out SimpleEval, and nobody uses that.

**Karina Nguyen** [15:59]
Mm-hmm.

**Swyx** [15:59]
Why isn't there just an one standard way everyone runs this? Because the, the alternative approach is you rerun your evals on their models.

**Karina Nguyen** [16:07]
Mm-hmm.

**Swyx** [16:07]
And obviously the numbers, your numbers will be lower-

**Karina Nguyen** [16:10]
Yeah

**Swyx** [16:10]
... and they'll be unhappy, so that's why you don't do that.

**Karina Nguyen** [16:13]
I think it operates on a assumption that, like, the models, the next generation of the model or the model that you produce next is gonna behave the same. So for example, like, I think the way you prompt o1 or, like, Claude 3 is gonna be very different from each other.

I feel like there is a lot of, like, prompting that you need to do to get the evals to run correctly. So sometimes the model will just, like, output, like, new lines, and the way it parsed will be, like, incorrect or something.

This has happened with, like, Stanford, I remember. Like, when, um, Stanford had this also, like they were, like, running benchmarks.

**Swyx** [16:46]
Helm?

**Karina Nguyen** [16:47]
Yeah, Helm.

**Swyx** [16:47]
Yeah.

**Karina Nguyen** [16:47]
And somehow, like, Claude was, like, always, like, not performing well, and that's because-

**Swyx** [16:52]
Oh

**Karina Nguyen** [16:52]
... like, the way they prompted it was kind of wrong. So it's, like, a lot of, like, techniques. It's just, like, very hard because, like, nobody even knows.

**Swyx** [17:01]
Has that gone away with chat models instead of, you know, just raw completion models?

**Karina Nguyen** [17:06]
Yeah. I guess, like, each eval-

**Swyx** [17:08]
Structured output as well

**Karina Nguyen** [17:09]
... also can be run in a very different way. Sometimes you can, like, the out-- ask the model to output in, like, XML tags, but some models are not really good at XML tags. So it's like-

**Swyx** [17:18]
Yeah

**Karina Nguyen** [17:18]
... do you change the formatting per model?

**Swyx** [17:21]
Mm.

**Karina Nguyen** [17:21]
Or, like, do you run the same format across all models? And then, like, the metrics themselves, right? Like, maybe, you know, accuracy is, like, one thing, but maybe you care about, like, some other metrics, like F score, like some other, like, things.

Yeah, it's, like, hard. I don't know. Yeah.

**Swyx** [17:37]
And talking about o1 prompting, we just had a o1 prompting post on the newsletter, which I think was the- Apparently it went viral within OpenAI. Yeah? I don't know. I got, I got pinged by other OpenAI people. They were like, "Is this helpful to us?"

I'm like Okay

**Karina Nguyen** [17:49]
Oh, nice.

**Swyx** [17:50]
I, I, I think it's, like, maybe one of the top three most-read posts now.

**Karina Nguyen** [17:54]
Yeah.

**Swyx** [17:54]
Um, so-

**Karina Nguyen** [17:54]
And I didn't write it.

**Swyx** [17:57]
Exactly.

**Karina Nguyen** [17:57]
Anyway, go ahead.

**Swyx** [17:58]
What are your tips on o1 versus, like-

**Karina Nguyen** [18:01]
Yeah

**Swyx** [18:01]
... Claude prompting or, like, what are things that you took away from that experience? And especially now I know that with 40 for Canvas you've done RL after on the model. So yeah, just general learning, so another think about prompting these models differently.

**Karina Nguyen** [18:13]
I actually think, like, o1 I did not even harness the magic of, like, o1 prompting, but, like, one thing that I found is that, like, if you give o1, like, hard, like, constraints of, like, what you're looking for, basically the model will be- will have a much easier time to, like, kind of like select the candidates and, uh, match, like, the candidate that is most, like, fulfill the criteria that you gave.

And I think there's a- the class of problems like this that o1 excels at. For example, if you have a question like a bio question on, like, some... Or like in chemistry, right? Like, if you have, like, very specific criteria with the protein or, like, some of the chemical bindings or something, like, then the model would be really- will be really good at, like, determining the exact candidate that will match these certain criteria.

**Swyx** [19:05]
I have often thought that we need a new IF eval for this.

**Karina Nguyen** [19:08]
Mm-hmm.

**Swyx** [19:08]
'Cause this is basically kind of instruction following, isn't it?

**Karina Nguyen** [19:10]
Yes.

**Swyx** [19:11]
But I don't think IF eval has, like, multi-step IF eval.

**Karina Nguyen** [19:14]
Yeah.

**Swyx** [19:14]
So that- that's what basically I use AI News for.

**Karina Nguyen** [19:17]
That's cool.

**Swyx** [19:17]
I have a lot of prompts and a lot of steps and a lot of criteria, and o1 just kind of checks through each kind of systematically.

**Karina Nguyen** [19:23]
That's nice.

**Swyx** [19:23]
And we don't have any evals like that.

**Karina Nguyen** [19:25]
Yeah.

**Swyx** [19:26]
Does OpenAI know how to prompt o1? I think that's kind of like the- That's the, uh... You know, Sam is always talking about incremental deployments and kind of like getting-- having people getting used to it. When you release a model, you obviously do all the safety testing, but do you feel like people internally know how to get 100% out of the model?

Or like are you also spending a lot of time learning from, like, the outside on how to better prompt o1 and, like, all these things?

**Karina Nguyen** [19:51]
Yeah. I certainly think that you learn so much from, like, external feedback too on how people use, like, o1. I think, like, a lot of people use o1 for, like, really hardcore, like, coding questions. I feel like I don't fully know how to-

**Swyx** [20:08]
Yeah, you've released the model

**Karina Nguyen** [20:09]
... best use o1 except for, like, I use o1 to just, like, do some, like, synthetic data explorations.

**Swyx** [20:15]
Mm-hmm.

**Karina Nguyen** [20:15]
But that's it.

**Swyx** [20:17]
Do people inside of OpenAI, once the model is coming out-

**Karina Nguyen** [20:21]
Mm

**Swyx** [20:21]
... do you get, like, a company-wide memo of like, "Hey, this is how you should try and prompt this," especially for people that might not be close to it during development?

**Karina Nguyen** [20:27]
Mm.

**Swyx** [20:28]
You know? Or I don't know if you can share anything, but I'm curious how internally these things kind of get shared.

**Karina Nguyen** [20:35]
I feel like I'm, like, in my own little corner in, like, research. I don't really, like- ... look at some of the Slack channels.

**Swyx** [20:40]
It's very, very big.

**Karina Nguyen** [20:43]
So I actually don't know if something like this exists. Probably. It must be exists because we need to share to, like, customers or, like, you know, like some of the guides on, like, how to use this model, so probably th- there is.

**Swyx** [20:57]
I often say this, the reason that AI engineering can exist outside of the model labs is because the model labs release models with capabilities that they don't even fully know-

**Karina Nguyen** [21:07]
Mm

**Swyx** [21:07]
... 'cause you never trained specifically for it. It's emergent.

**Karina Nguyen** [21:10]
Mm-hmm.

**Swyx** [21:10]
And you can rely on basically crowdsourcing the search of that space or the behavior space to the rest of us.

**Karina Nguyen** [21:17]
Yeah.

**Swyx** [21:18]
Yeah. So, like, you don't have to know-

**Karina Nguyen** [21:19]
Yeah

**Swyx** [21:19]
... is what I'm saying.

**Karina Nguyen** [21:20]
Yeah. I think, like, um, an interesting thing about, like, o1 is that, like, it's really... for, like, average human, sometimes I don't even know whether the model, like, produced the correct output or not. Like, it's really hard for me to, like, verify even, like, hard, like, STEM questions.

I don't know. If I'm not an expert, like, I usually don't know. So it's like the question of, like, alignment is actually more important, like, for these, like, complex reasoning models to, like, how do we help humans to, like, verify the outputs of these models is quite important.

And I feel like, yeah, like learning from external feedback is kind of cool.

**Swyx** [21:56]
For sure. Um, one last thing on Claude 3. You had a section on behavioral design.

### Model Personality

**Karina Nguyen** [22:01]
Yes.

**Swyx** [22:01]
Anthropic's very famous for the HHH goals.

**Karina Nguyen** [22:05]
Mm-hmm.

**Swyx** [22:05]
What was your insights there? Or, you know, maybe just talk a little bit about what you s- explored.

**Karina Nguyen** [22:09]
Yeah. I think, like, behavioral design is, like, a really cool... I'm glad that I, I made it, like, a section around this, and it's, like, really cool. I think, like-

**Swyx** [22:18]
Like, you weren't gonna publish one and then you insisted on it or what? Like-

**Karina Nguyen** [22:21]
Mm. I think, like- ... I just, like, put the section inside it and, like, yeah, Jared, my- like, one of my f- most favorite researchers was like, "Yeah. That's cool. Let's, let's do that, I guess." Um, yeah, like, nobody had this, like, term of, like, behavioral design necessarily for the models.

It's kind of like a new little field of, like, extending, like, product design into, like, the model design, right? Like, so how do you create a behavior for the model in certain contexts? So, uh, as for example, like, in Canvas, right, like, one of the things that we had to, like, think about is like, okay, like, now the model enters, like, more collaborative environment, more collaborative context.

So, like, what's the most appropriate behavior for the model to act like as a collaborator? Should it ask, like, more follow-up questions? Should it, like, change... Uh, what's the tone should be? Like, what is the collaborator's tone? It's different from, like, a chat, like conversationalist-

**Swyx** [23:18]
Mm

**Karina Nguyen** [23:18]
... uh, versus a collaborator. So how do you shape the persona and the personality around that? It has, like, some philosophical questions too. Like, yeah, behavioral... Like, I mean, like, I guess, like, I can talk more about, like, the methods of, like, creating the personality.

**Swyx** [23:33]
Please.

**Karina Nguyen** [23:34]
It's the same thing as like you would create like a character in a video game or something. It's kind of like-

**Swyx** [23:40]
Charisma, intelligence-

**Karina Nguyen** [23:41]
Yeah, exactly

**Swyx** [23:42]
... wisdom.

**Karina Nguyen** [23:43]
Like, what are the core principles-

**Swyx** [23:44]
Helpful, harmless, honest

**Karina Nguyen** [23:46]
... like core values?

**Swyx** [23:46]
Yeah.

**Karina Nguyen** [23:46]
And obviously for Claude this was m- is much easier than I would say, like, for ChatGPT. For Claude, it's, like, it's, like, baked in in, like, the mission, right? It's like honest, harmless-

**Swyx** [23:57]
Helpful

**Karina Nguyen** [23:58]
... helpful. But the most complicated thing about, like, the model behavior or, like, the behavioral design is that, like, sometimes two values would contradict each other. I think this happened in Claude 3. One of the main things that we were thinking about is, like, how do we balance this, like, honesty versus, like, har- harmlessness or, like, helpfulness?

It's like we don't want the model to always, like, refuse even to, like, innocuous queries like some, like, creative writing prompts. But also if you don't want the model to be-- act like a-- be harmful or something. So it's like there's always a balance between those two, and it's more like art than a science necessarily, and this is what datasets craft is, is like more of an art than a literal science.

You can definitely do, like, empirical research on this, but it's actually like-- like, this is the, the idea of, like, synthetic data. Like, if you look back to a constitution AI paper, is around, like, how do you create completions such that you would agree to certain, like, principles that you want your model to, to agree on.

So it's like, if you create the core values of the models, how do you decompose those core values into, like, specific scenarios? Or like, so how does the model need to express its honesty in a variety, uh, kind of like scenarios?

And this is where, like, generalization happens, uh, when you craft the persona of the model.

**Swyx** [25:22]
Yeah. It seems like, um, what you describe behavior modification or sh-

**Karina Nguyen** [25:27]
Mm-hmm

**Swyx** [25:27]
... shaping as a side job that was done in, done... I mean, I think Anthropic has always focused on it the first and, and the most.

**Karina Nguyen** [25:34]
Mm-hmm.

**Swyx** [25:34]
But now it's like every lab has sort of vibes officer.

**Karina Nguyen** [25:38]
Mm-hmm.

**Swyx** [25:38]
For you guys, it's Amanda.

**Karina Nguyen** [25:39]
Mm-hmm.

**Swyx** [25:40]
For OpenAI, it's Rune. Um, and then for, for Google, it's Steven Johnson and Raisa-

**Karina Nguyen** [25:45]
Mm-hmm

**Swyx** [25:46]
... who we had on the podcast. Do you think this is like a job? Like it's like a-- like every, every company needs a tastemaker?

**Karina Nguyen** [25:51]
I think the model's personality is actually the reflection of the company or the reflection of the people who create that model. So, like, for Claude's, I think Amanda was doing a lot of, like, Claude character work, and I was working with her at that time.

Um-

**Swyx** [26:05]
Yeah, but there was no team, like Claude character team or something.

**Karina Nguyen** [26:07]
Now there is a-

**Swyx** [26:08]
Yeah, right?

**Karina Nguyen** [26:08]
... little bit of a team.

**Swyx** [26:09]
Isn't that cool?

**Karina Nguyen** [26:10]
But before that, there was none.

**Swyx** [26:11]
Yeah.

**Karina Nguyen** [26:11]
I think, like, actually with Claude 3, it was like we kind of doubled down on the feedback from Claude 2. Like people... We didn't even, like, think, but, like, people said, like, Claude 2 is, like, so much better at, like, writing and, like, has certain personality even though it was, like, unintentional at all.

And we did not pay that much attention and didn't know even how to, like, productionize this property of model being h- better, like, personality until, like, with Claude 3, we kind of like had to, like, double down because we knew that we would launch, like, in chat.

We wanted to, like... Claude honesty is, like, really good for, like, enterprise customers, so we kind of wanted to, like, make sure the hallucinations went-- like factuality would, like, go up or something. We didn't have a team until or after, like, Claude 3, I guess.

**Swyx** [26:58]
Yeah. I mean, it's, it's growing now, and-

**Karina Nguyen** [26:59]
Yeah

**Swyx** [26:59]
... I think anyway, everyone's taking it seriously.

**Karina Nguyen** [27:01]
I think at OpenAI, there was a team called Model Design. It's Jan, the PM, she's leading that team, and I work very closely with those teams. Like, we were working on, like, actually writing improvements that we did with ChatGPT last year, and then I was working on, like, this collaboration, like how do we make ChatGPT act like this collaborator for, like, Canvas.

And then, yeah, we worked together with on some of the projects.

**Swyx** [27:25]
I don't think it's publicly known, his- ... his actual name other than Rune, but I-

**Karina Nguyen** [27:30]
Okay, sorry. Maybe we should like re-

**Swyx** [27:31]
He's, he's-- He's mostly, he's mostly docs.

**Karina Nguyen** [27:32]
Cut that.

**Swyx** [27:33]
We'll, we'll beep it-

**Karina Nguyen** [27:34]
The model design

**Swyx** [27:34]
... and then people can guess.

### Canvas

**Karina Nguyen** [27:37]
Yeah.

**Alessio** [27:37]
Do we wanna move on to OpenAI-

**Karina Nguyen** [27:40]
Sure

**Alessio** [27:40]
... and some of the recent work, especially you mentioned Canvas. So the, the first thing about Canvas is, like, it's not just a UX thing. You have a different model in the back end which you, um, post trained on o1 preview distilled data-

**Karina Nguyen** [27:54]
Mm

**Alessio** [27:54]
... which was pretty interesting. Can you maybe just run people through, you come up with a feature idea maybe-

**Karina Nguyen** [27:58]
Mm

**Alessio** [27:58]
... then how do you decide what goes in the model, what goes in the product-

**Karina Nguyen** [28:02]
Mm-hmm

**Alessio** [28:02]
... and just that, that process?

**Karina Nguyen** [28:03]
Yeah. I think the most unique thing about ChatGPT Canvas was that it was also the team formed out of the air. So it was like July 4th or something, uh, during the break.

**Swyx** [28:16]
Wow.

**Karina Nguyen** [28:16]
And-

**Swyx** [28:17]
Like Independence Day. They just like- Okay.

**Karina Nguyen** [28:19]
I think it was there some, like, company break or something, and I, I remember I was just, like, taking a break, and then I was, like, pitching this idea to, like, Barrett Zoph-

**Swyx** [28:29]
Barrett Zoph, yeah

**Karina Nguyen** [28:29]
... uh, who was my manager at that time. She's like, "I just wanna, like, create this, like, canvas or something." And I really didn't know how to, like, navigate OpenAI's... Uh, it was, like, my first, like, I don't know, like, first month at OpenAI, and I really didn't know how to, like, navigate, how do I get product to work with me or, like, some of the ideas, like, some of the things like this was like...

So I'm really grateful for, like, actually Barrett and Mira, who helped me to, like, start this project basically. And I think that was really cool. And it was, like, this 4th of July, and, like, Barrett was like, "Yeah."

Actually, who's, like, an engineering manager is like, "Yeah, we should, like, staff this project with, like, s- five, six engineers or something, and then Karina can be a, like, researcher on this project." And I think, like, this is how the team was formed.

This was kind of, like, out of the air. And so, like, I didn't know anyone there at that time, except for Thomas Dimson. He did, like, the first, like, initial, like, engineering prototype of the canvas, and it kind of, like, riffed off.

But I think the first... We learned a lot on the way how to work together as product and research. And I think this is one of the first projects at OpenAI where research and product worked together from the very beginning, and we just made it, like, a successful project in my opinion, is because, like, designers, engineers, PM, and research team were all together, and we would, like, push back on each other, like, if, like, it doesn't make sense to do it on a model side.

W- like, we had to, like, collaborate with, like, applied engineers to, like, make sure this is being handled on the applied side. But the idea is you can go that far with, like, prompted baseline. Prompted ChatGPT was kind of, like, the first thing that we tried, was like a canvas as a tool or something.

So how do we define the behavior of the canvas? But then, like, we, we've, we've found a bunch of, like, different, like, edge cases that we wanted to, like, fix, and the only way to, like, fix some of these edge cases is actually through post-training.

So we actually-- What we did was actually retrain the entire 40 plus our Canvas stuff. And this is like-- There are, like, two reasons why we did this. It's because, like, the first one is that we wanted to ship this as a better model in the dropdown menu.

We could, like, rapidly iterate on users' feedback as we ship it and not going through the entire, like, integration process into, like, this, like, new one model or something, which took some time, right? So, like, from beta to, like, GA it took, I think, three months.

So we kind of wanted to, like, ship our own model with that feature to, like, learn from the user feedback very quickly. So that was, like, one of the decisions that we made. And then with Canvas itself, we just, like, had a lot of, like, different, like, behavioral...

It's, again, like, it's behavioral engineering. It's kind of like various behavioral craft around, like, when does Canvas needs to write comment? When does it need to, like, update or, like, edit the document? When does it need to edit the entire-- like rewrite the entire document versus, like, edit very specific section that the user asks?

And when does it need to, like, trigger the canvas itself? Was one of those, those, like, behavioral engineering questions that we had. At that time, I was also working on, like, writing quality, so that was, like, the perfect way for us to, like, literally both teach the model how to use Canvas, but also, like, improve writing quality if writing was, like, one of the main use cases for ChatGPT.

So I think that was, like, the reasoning around that.

**Swyx** [31:56]
There's so many questions. Oh my God.

**Karina Nguyen** [31:57]
Okay, yeah.

**Swyx** [31:58]
Quick one. What does improve writing quality mean? Uh, what are the evals?

**Karina Nguyen** [32:02]
What are the evals?

**Swyx** [32:03]
How do you improve it?

**Karina Nguyen** [32:03]
Yeah. So the way I'm thinking about it is, like, have two various directions. The first direction is, like, how do you improve the quality of the writing of the current use cases of ChatGPT? And those-- most of the use cases are mostly, like, nonfiction writings.

It's like email writing or, like, some of the maybe blog posts. Cover letters is, like, one of the main use cases. But then the second one is, like, how do we teach the model to literally think more creatively or, like, write in a more creative manner such that it will, like, just create novel forms writing?

And I think the second one is, like, much of a longer-term, like, research question, while the first one is more like, okay, we just need to improve data quality for the writing use cases that between the models are.

It is more straightforward question, but the way we evaluated the writing quality... So actually, I worked with Jan's team on the model design, so they had a team of, like, model writers, and we would work together, and it's just like a human eval.

It's like internal human eval where we would just like-

**Swyx** [33:11]
Always like that.

**Karina Nguyen** [33:12]
Yeah, on the prompt distribution that we cared about. Like, we want to make sure that the models that we, like, used, that we trained were always, like, better or something. Um-

**Swyx** [33:21]
Yeah. So, like, some test set of, like, 100 prompts that-

**Karina Nguyen** [33:24]
Yes

**Swyx** [33:24]
... you want to make sure you're good on. I don't know how big the prompt distribution needs to be, 'cause it-- you, you are literally catering to everyone.

**Karina Nguyen** [33:32]
Right. Yeah. I think it was much more opinionated way of, like, improving writing quality because we worked together with, like, model designers to, like, come up with, like, core principles of what makes this particular writing good. Like, what does make email writing good?

And we had to, like, craft, like, some of the literal, like, rubric on, like, what makes this good and then, uh, make sure during the eval we check the marks on this, like, rubric.

**Swyx** [33:58]
Yeah. That, that's what I do.

**Karina Nguyen** [34:00]
Yeah.

**Swyx** [34:00]
That's what school-

**Karina Nguyen** [34:00]
It's like most common practice

**Swyx** [34:01]
... that's what school teachers do.

**Karina Nguyen** [34:02]
Yeah. Yeah.

**Swyx** [34:03]
It's really funny. Like, yeah, that's exactly how we grade essays.

**Karina Nguyen** [34:06]
Yes.

**Alessio** [34:07]
Yeah, I guess my question is: When do you work the improvements back in the model? So the Canvas model is better at writing. Why not just make the core model better, too? So for example, I built this small podcaster thing-

**Karina Nguyen** [34:19]
Mm-hmm

**Alessio** [34:19]
... for a podcast and I have the 4o API and I asked it to write-

**Karina Nguyen** [34:23]
Mm-hmm

**Alessio** [34:23]
... a write-up about the episode based on the transcript, and then I done the same in Canvas.

**Karina Nguyen** [34:26]
Mm-hmm.

**Alessio** [34:27]
The Canvas one is a lot better.

**Karina Nguyen** [34:28]
Mm-hmm.

**Alessio** [34:28]
Like, the one from the raw 4o, it starts, "The podcast delves," and I was like, "No, not delve. Then the, their word." Why not put them back in 4o core or is there just like-

**Karina Nguyen** [34:38]
I think we put it back in the core now.

**Alessio** [34:41]
Yeah? So, like-

**Karina Nguyen** [34:41]
In the Canvas

**Alessio** [34:41]
... so the 4o Canvas now is the same-

**Karina Nguyen** [34:43]
Yes

**Alessio** [34:44]
... as 4o?

**Swyx** [34:45]
Yeah. You, you must have missed that update.

**Alessio** [34:46]
Yeah. What's the, what's the-

**Karina Nguyen** [34:48]
But I think the models-

**Alessio** [34:48]
What's the process to-

**Karina Nguyen** [34:49]
... are still a little bit different.

**Swyx** [34:50]
I think it's just like an A/B test almost, right?

**Alessio** [34:52]
To me it feels-- I mean, I've only tried it, like, three times.

**Karina Nguyen** [34:55]
Mm-hmm.

**Alessio** [34:56]
But it feels-

**Karina Nguyen** [34:57]
Different

**Alessio** [34:57]
... the Canvas, the Canvas output feels very different than the API output.

**Karina Nguyen** [35:02]
Yeah. Yeah. I think, like, there's always, like, a difference in the model quality. I would say, like, the original better model that we released with Canvas was actually much more creative than even right now when I use, like, 4o with Canvas.

I think it's just, like, the complexity of, like, the data and the complexity of the... It's kind of like versioning issues right here. It's like, okay, like, your version 11 will be very different from, like, version 8, right?

It's like, ugh, even though, like, the stuff that you put in is, like, the same or something. Um-

**Swyx** [35:32]
It's a good time to, to say that I have used it a lot more than three times. I'm a huge-

**Karina Nguyen** [35:37]
Mm-hmm

**Swyx** [35:37]
... fan of Canvas. I think it is, um... yeah, ev- like, it's weird when I talk to my other friends. They, they don't really get it yet-

**Karina Nguyen** [35:44]
Mm-hmm

**Swyx** [35:44]
... or they don't really use it yet, I think because it's maybe sold as, like, sort of writing help.

**Karina Nguyen** [35:48]
Mm-hmm.

**Swyx** [35:48]
When really, like, it's kind- it's the scratch pad.

**Karina Nguyen** [35:51]
Mm-hmm.

**Alessio** [35:51]
Yeah. What are the core use cases or like, yeah.

**Karina Nguyen** [35:53]
Oh, yeah. I'm curious to learn.

**Swyx** [35:54]
Literally drafting anything. Like, I wanna draft, like, copy for my conference that I'm running. Like, I'll put it there first and then I, I, like... It'll just have the canvas up, and I'll just say what I don't like about it and it changes.

I will maybe edit stuff here and paste in... So, so for example, like, I wanted to draft a brainstorm list of reasons, of signs that you may be an NPC.

**Karina Nguyen** [36:15]
Mm-hmm.

**Swyx** [36:15]
Just for fun. Just like a blog post for fun.

**Karina Nguyen** [36:17]
Nice.

**Swyx** [36:17]
And I was like, "Okay, I'll do 10 of these, and then I want you to generate the next 10." So I wrote 10. I pasted it into, to ChatGPT, and it generated the next 10, and they all sucked.

All horrible. But it also spun up the canvas with, with the blog post. And I was like, "Okay-" Self-critique why your output sucks-

**Karina Nguyen** [36:35]
Yeah

**Swyx** [36:35]
... and then try again. And it, and it just kind of just iterates-

**Karina Nguyen** [36:38]
Mm-hmm

**Swyx** [36:38]
... on the blog post with me as a writing partner, and it is so much better than, I don't know, s- like intermediate steps. It's like that would be my primary use case. Like literally-

**Karina Nguyen** [36:48]
Mm

**Swyx** [36:48]
... drafting anything. I think the other way that I, I'll put it, I'm not putting words in your mouth, like this is how I view what Canvas is and why-

**Karina Nguyen** [36:55]
Mm-hmm

**Swyx** [36:55]
... it's so important, it's basically an inversion of what Google Docs is-

**Karina Nguyen** [36:59]
Mm-hmm

**Swyx** [36:59]
... wants to do with Gemini. It's like Google Docs on the main screen and then Gemini on the side.

**Karina Nguyen** [37:03]
Mm-hmm.

**Swyx** [37:04]
And write... What- not what ChatGPT has done is do the chat thing first and then the docs on the side.

**Karina Nguyen** [37:08]
Mm-hmm.

**Swyx** [37:08]
But it's kind of like a reversal of, of what is the-

**Karina Nguyen** [37:10]
Mm

**Swyx** [37:10]
... main thing. Like Google Docs starts with the canvas first that you can edit and whatever, and then you maybe sometimes you call in the AI assistance.

**Karina Nguyen** [37:16]
Mm-hmm.

**Swyx** [37:16]
But ChatGPT, what you are now is you're kind of AI first-

**Karina Nguyen** [37:20]
Mm-hmm

**Swyx** [37:20]
... with the s- the side output being Google Docs.

**Karina Nguyen** [37:23]
I think we definitely want to improve like writing use case in terms of like how do we make it easier for people to format or like do some of the editing. I think there is still like a lot of r- room for improvement, to be honest.

I think another thing is like coding, right? I feel like one of the-

**Swyx** [37:38]
Yes

**Karina Nguyen** [37:38]
... uh, things that'd be like doubling down is actually like executing code inside the canvas, since there is a lot of questions like how do we evolve this? It's kind of like IDE for both. And I feel like this is where-

**Swyx** [37:50]
Yeah

**Karina Nguyen** [37:50]
... I'm, I'm coming, coming from is like the ChatGPT evolves into this blank interface which can morph itself in whatever you trying... Like the model should try to like derive your true intent and then modify the interface based on your intent, and then if you like writing, it should become like the most powerful like writing ID- IDE possible.

If it's like coding, it should become like a coding IDE or something.

**Swyx** [38:15]
I think it's a little bit of a odd decision for me to call those two things the same product name-

**Karina Nguyen** [38:19]
Mm

**Swyx** [38:19]
... because they're basically two different UIs.

**Karina Nguyen** [38:22]
Mm-hmm.

**Swyx** [38:22]
Like one-

**Karina Nguyen** [38:23]
Yes

**Swyx** [38:23]
... one is code interpreter plus plus-

**Karina Nguyen** [38:24]
Yeah, yeah

**Swyx** [38:24]
... and the other one's Canvas.

**Karina Nguyen** [38:26]
Yes.

**Swyx** [38:26]
I don't know if you have other thoughts on Canvas.

**Alessio** [38:28]
No, I'm just curious, maybe some of the harder things. So when I was reading, for example, forcing the model to do targeted edits-

**Karina Nguyen** [38:34]
Mm-hmm

**Alessio** [38:34]
... versus like full rewrite, sounds like it was like really hard. In the AI engineer mind, maybe sometimes it's like just pass one sentence in the prompt, it's just gonna rewrite that sentence, right? But obviously it's harder than that.

What are maybe some of the like hard things that people don't understand from the outside in building products like this?

**Karina Nguyen** [38:51]
I think it's always hard with any new like product feature, like Canvas or Tasks or like any other new features, like you don't know how people would use this feature, and so how do you even like build evaluations that would simulate how people would use this feature?

And it's always like really hard for us, therefore like we, we try to like lean on to like iterative deployment this in order to like learn from user feedback as much as possible. Again, it's like we didn't r- know that like code diffs was very difficult for a model, for example.

Again, it's like do we go back to like fundamentally improve like code diffs as a model capability, or do you like do a workaround where the model will just like rewrite the entire document, which is yield to like higher accuracy?

And so those are like some of the decisions that we had to like make as, yeah, how do you like improve the bar to the product quality but also make sure the model quality is also a part of it, and like what kind of like trade-offs you're okay to do?

Again, I think, I think this is like new way of product development. It's more like product research. Model training and like product development goes like together hand in hand. This is like one of the hardest things. Like defining the entire like model behaviors.

Uh, I think just like there's so many edge cases that might happen, especially when you like do Canvas with like other tools, right? Like Canvas plus DALL-E, Canvas plus search. If you like select certain section and then like ask for search, like how do you build such evals?

Like what kind of like features or like behaviors that you care the most about, and this is how you build evals. Um-

**Swyx** [40:36]
You tested against every feature of ChatGPT?

**Karina Nguyen** [40:38]
Uh, no.

**Swyx** [40:39]
Oh, okay.

**Karina Nguyen** [40:40]
Yes.

**Swyx** [40:40]
I, I, I mean, I don't think there's that many that you can-

**Karina Nguyen** [40:43]
Right

**Swyx** [40:43]
... that would take forever, but-

**Karina Nguyen** [40:45]
But it's the same as any decision boundary between like Python ADA advanced data analysis-

**Swyx** [40:52]
Mm-hmm

**Karina Nguyen** [40:52]
... versus Canvas is one of the most trickiest like decision boundary behaviors that we had to like figure out. Like how do you derive the intent from the human user query?

**Swyx** [41:03]
Yeah.

**Karina Nguyen** [41:03]
And how do I say this? Deriving the intent, meaning does the user expect Canvas or some other tool, and then like make sure that it's like maximally like the intent was, is like actually still one of the hardest problems.

**Swyx** [41:21]
Yeah.

**Karina Nguyen** [41:21]
Especially with like agents, right? Like you don't want like agents to go for like five minutes and do something on the background and then come back with like some mid answer that you could have gotten from like a normal model- ...

or like the answers that you didn't even want because it didn't have enough context, so it didn't like follow up correctly or...

**Swyx** [41:41]
You said the magic word. We have to take a shot every time you say it. You said agents.

**Karina Nguyen** [41:44]
Agents, yeah.

### Tasks

**Swyx** [41:46]
So let's move to Tasks. You just launched Tasks.

**Karina Nguyen** [41:49]
Mm-hmm.

**Swyx** [41:49]
What was that like? What was the story? I mean, it's, it's your, it's your baby, so...

**Karina Nguyen** [41:54]
Now that I have a team, I actually like T- Tasks was purely like my residence project. I was-

**Swyx** [42:01]
Oh

**Karina Nguyen** [42:01]
... mostly a supervisor, so I kind of like delegated a lot of things to my resident. His name is like Vivek. And I think this is like one of the projects where I learned management I would say.

**Swyx** [42:15]
Yeah.

**Karina Nguyen** [42:16]
But it was really cool. I think it's very similar model. I'm, I'm trying to replicate Canvas operational model, how do we operate with product people or like product applied orgs with research. And the same happened, I was trying to replicate like the methods and replicate the operational process with Tasks.

And actually Tasks was developed less than like two months.

**Swyx** [42:38]
Mm.

**Karina Nguyen** [42:39]
So if Canvas took like, I don't know, four months, then Tasks took like two months. And I think again, like it's kind of very similar process of like how do we build evals? You know, some people like ask- For like reminders in actual ChatGPT, but then like obviously-

**Swyx** [42:56]
Even though they know it doesn't work

**Karina Nguyen** [42:57]
... yeah, it doesn't work.

**Swyx** [42:59]
Yeah.

**Karina Nguyen** [42:59]
So like there is some like demand or like desire from users to like do this. And actually I feel like Task is like simple feature in my opinion. It's-

**Swyx** [43:06]
Chrome job

**Karina Nguyen** [43:06]
... something that you would want from any model, right? But then the magic is like when... A- actually because the model is so general, it knows how to use search or like Canvas, or like create sci-fi stories, and create Python puzzles.

When coupled with Task, this actually becomes like really, really powerful. It was like the same ideas of like, how do we shape the behavior of the model? Again, we shipped it as like as a beta model in the model dropdown, and then we are working towards like making that feature integrated in like the core model.

So I feel like the, the principle is that like everything should be like in one model, but because of some of the operational difficulties, it's, it's much easier to like deploy as a separate model first to like learn from the user feedback, and then iterate very quickly, and then improve into the core model, basically.

Again, this is a project was also like together at the beginning, from the very beginning, designers, engineers, researchers were working all together. And together with model designers we were like trying to like come up with like evals, evaluations, and like testing and like bug bashing, and it's like a lot of cool like synergy.

**Swyx** [44:13]
Evals, bug bashing. I'm trying to distill-

**Karina Nguyen** [44:16]
Okay

**Swyx** [44:16]
... I, I would love a canvas for this, for distill what the ideal product management or research management process is, right?

**Karina Nguyen** [44:22]
Mm-hmm.

**Swyx** [44:23]
Start from like, do you have a PRD? Do you have a doc that like lists these things?

**Karina Nguyen** [44:27]
Yes.

**Swyx** [44:27]
And then from PRD, you get funding maybe-

**Karina Nguyen** [44:30]
Um-

**Swyx** [44:31]
... or like f- you know, staffing resources, whatever.

**Karina Nguyen** [44:33]
Yes.

**Swyx** [44:34]
And then prototype maybe. Uh-

**Karina Nguyen** [44:37]
Yeah, prototype. I would say like prototype with prompted baseline. It's all, all-

**Swyx** [44:42]
Always prompted baseline

**Karina Nguyen** [44:42]
... everything starts with like prompted baseline.

**Swyx** [44:44]
Yeah.

**Karina Nguyen** [44:44]
And then like we craft like certain like evaluations that we want to like capture.

**Swyx** [44:48]
Okay.

**Karina Nguyen** [44:48]
That we want to like measure progress at least-

**Swyx** [44:50]
Yeah

**Karina Nguyen** [44:50]
... with the model.

**Swyx** [44:50]
Yeah.

**Karina Nguyen** [44:51]
And then make sure the evals are good, and make sure that the prompted baseline actually fails on those like evals, because then you have like AVL to like hill climb on. And then once you start iterating on the model training, it actually very iterative.

So like every time you train the model or you like look at the bench- or like look at your evals and that goes up, it's like good, but then also you don't want to like... You wanna make sure it, it's not like super over-fitting.

Like that's where you run on other evals, right, like intelligence evals or something and then like-

**Swyx** [45:20]
You, you don't want regressions on the other stuff.

**Karina Nguyen** [45:22]
Right. Yes.

**Swyx** [45:22]
Okay. Is that your job or is that like the rest of the company's job to do?

**Karina Nguyen** [45:26]
Um, I think it's mainly my like-

**Swyx** [45:29]
Really?

**Karina Nguyen** [45:29]
... the job of the people who like-

**Swyx** [45:31]
Because regressions are going to happen and you don't-

**Karina Nguyen** [45:33]
Yes

**Swyx** [45:33]
... necessarily own the data for the other stuff.

**Karina Nguyen** [45:35]
What's happening right now is that like you basically you only like upload your, your data sets, right? So it's like you compare on the baseline, you compare like the regressions on the baseline model.

**Swyx** [45:48]
Model training and then bug bash, and that's, that's about it, and then ship.

**Karina Nguyen** [45:51]
Actually, I did the course with Andrew Ng, who-

**Swyx** [45:53]
Yes

**Karina Nguyen** [45:54]
... um, there's like one little lesson around this.

**Swyx** [45:57]
Okay. I haven't seen-

**Karina Nguyen** [45:59]
Um, like product research lifecycle.

**Swyx** [46:00]
You tweeted a picture with him- ... and it wasn't clear if you were working on a course. I mean, it looked like the standard course picture with Andrew Ng.

**Karina Nguyen** [46:06]
Yes.

**Swyx** [46:06]
Uh, okay. It was a course with him. What was that like working with him?

**Karina Nguyen** [46:09]
No, I'm not working with him. Like I just like did the course with him.

**Swyx** [46:11]
Yeah, yeah.

**Karina Nguyen** [46:12]
Yeah.

**Alessio** [46:12]
How do you think about the tasks? So I started creating a bunch of them. Like do you see this as being, going back to like the composability, like composable together later?

**Karina Nguyen** [46:22]
Mm-hmm.

**Alessio** [46:22]
Like, uh, you're gonna be scheduled one task that does-

**Karina Nguyen** [46:25]
Mm-hmm

**Alessio** [46:25]
... multiple tasks chained together. What's the vision?

**Karina Nguyen** [46:28]
I would say Task is like a foundational module. Obviously to generalize to all sorts of like behaviors that you want. Like sometimes like I see like people have like three tasks in one query, and right now I don't think like the model handles this very well.

I think that ideally we learn from like the user behavior, and ideally the model will just be more proactive in suggesting of like, "Oh, I can either do this for you every day because I've observed that you do that every day," or something.

So it's like more becomes like a proactive behavior. I think right now you have to be more explicit like, "Oh yeah, like every day, like remind me this." But I think like the, the ideally the model will always think about you on the background and like kinda suggest, "Okay, like I noticed you've been reading some, uh, s- this particular like Hacker News articles.

Maybe I can try to suggest you like every day or something." So it like, it's just like much more like of a natural like friend, I think.

**Swyx** [47:36]
Well, there is an actual startup called Friend that is trying to do that .

**Karina Nguyen** [47:39]
That's cool. Yes.

**Swyx** [47:40]
We'll have-- we'll interview Avi at some point.

**Karina Nguyen** [47:42]
Yeah.

**Swyx** [47:42]
But like the, it sounds like the guiding principle is just what is useful to you. It's a little bit B2C, um-

**Karina Nguyen** [47:47]
Mm-hmm

**Swyx** [47:47]
... you know. Is there any B2B push at all or you don't think about that?

**Karina Nguyen** [47:51]
I personally don't think about that as much, but I definitely feel like B2B is cool. Again, I, I come back to like Claude and Slack as like one of the f- like the first like interfaces where like the model was operating inside your organization, right?

It would be very cool for the model to like handle, to like become like a productive member of your organization.

**Swyx** [48:14]
Mm-hmm.

**Karina Nguyen** [48:14]
And then either like even like even process... Like, uh, right now, like I'm thinking like processing like user feedback. I think it'd be very cool if the model would just like start doing this for us and like we don't have to hire a new person on this- ...

just for this or something.

**Swyx** [48:31]
Yeah.

**Karina Nguyen** [48:31]
And like we have like- Very simple, like data analysis, so like data analytics or like how this feature is like-

**Swyx** [48:37]
Do you do this analysis yourself, or do you have a data science team that tells you insights?

**Karina Nguyen** [48:41]
I think there are some data scientists.

**Swyx** [48:42]
Okay.

**Karina Nguyen** [48:43]
But-

**Swyx** [48:43]
'Cause I, I've often wondered, I think there should be some startup or something that does automated data insights.

**Karina Nguyen** [48:48]
Yeah.

**Swyx** [48:48]
Like, I just throw you my data, you t- you tell me-

**Karina Nguyen** [48:50]
Yeah

**Swyx** [48:51]
... like what to

**Karina Nguyen** [48:51]
Yeah, exactly. Yeah.

**Swyx** [48:52]
'Cause that's what a data team at any company does.

**Karina Nguyen** [48:54]
Right.

**Swyx** [48:55]
Which is just give us your data, we'll, we'll like m- make PowerPoints.

**Karina Nguyen** [48:58]
Yeah.

**Swyx** [48:59]
And

**Karina Nguyen** [48:59]
Yeah, that's, that'd be very cool.

**Swyx** [49:01]
That's, I think that's a, that's a really good vision. You had thoughts on agents in general. There's some more proactive stuff. Uh, you actually had tweeted a definition which is-

**Karina Nguyen** [49:09]
Oh

**Swyx** [49:09]
... kind of interesting.

**Karina Nguyen** [49:10]
I did?

**Swyx** [49:10]
Or... Well, I'll, I'll read it out to you. You tell me-

**Karina Nguyen** [49:13]
Okay

**Swyx** [49:13]
... uh, you can double-click on it

**Alessio** [49:14]
If you still agree with yourself.

**Swyx** [49:16]
Yeah, yeah. Um, this is five days ago. Agents are gradual progressional tasks-

**Karina Nguyen** [49:19]
Oh, yeah

**Swyx** [49:19]
... starting off with one-off actions-

**Karina Nguyen** [49:21]
Okay

**Swyx** [49:21]
... moving to collaboration. Ultimately, fully trustworthy long horizon. I know it's, I know it's uncomfortable to have your tweets read to you. I, I have had this done to me. Ultimately, fully trustworthy long-horizon delegation in complex environments like multiplayer, multi-agents, tasks, and Canvas fall within the first two.

**Karina Nguyen** [49:34]
Mm-hmm.

**Swyx** [49:34]
What is the third form factor?

**Karina Nguyen** [49:34]
One of my weaknesses is like I, I like writing long sentences. I feel like I need to like-

**Swyx** [49:38]
No, no

**Karina Nguyen** [49:39]
... go back to

**Swyx** [49:39]
That's fine. That's fine. Is that your definition of agents? Like what are you looking for?

**Karina Nguyen** [49:44]
Um, I'm not sure if this is my definition of agents, but I feel like it's more like how I think it makes sense, right? Like I feel like for me to like trust an agent with my passwords or my credit card, I actually need to build trust with that agent that it will handle my tasks correctly and reliably.

And the way I would go about this is how I would naturally like collaborate with other people is that like we first... Even with any project, right? Like we first came-- when we first come, like we don't even know each other.

Like we don't know how each other's like working style. Like what I prefer, what do they prefer, how do they prefer to communicate, et cetera, et cetera. So like you spend like the first like, I don't know, like two weeks to just like learn their style of working and then like over time you adapt to their working style, and then this is how you create the collaboration.

And then like at the beginning you don't have much trust, so like how do you build more trust? Especially like... It's the same thing as like with a manager, right? Like it's like how do you build trust with your manager?

What does he need to know about you? What do you need to know about them? Over time, as you build trust and trust builds either through collaboration, which is why I feel like building Canvas was kind of like the first steps towards like more collaborative agents.

I think with humans, like you can- you should- need to show a consistent effort to each other, like consistent effort that you care about each other, you, that you like work together very well or something. So consistency and like collaboration is like what creates trust.

And then I will naturally will try to delegate tasks to a model because I know the model will not fail me or something. So it's kinda like building out like the intuition for the form factor of like new agents.

Because sometimes I feel like a lot of researchers or like people in AI community are like so into like, "Yeah, agents delegate everything," like blah, blah. But like on the way towards that, I think like collaboration is actually one of the main roadblocks or like milestones to get over, because then you will learn some of the implicit preferences that would help you, that would help towards like this full delegation model.

**Swyx** [51:55]
Yeah. Trust, very important. I have an AGI working for me- ... and I, I, we're, we're still working out the trust issues.

**Karina Nguyen** [52:00]
Okay.

**Swyx** [52:01]
Um, we are recording this just before the launch of Operator. The other side of agents that is very topical recently is computer use. Anthropic launch-

**Karina Nguyen** [52:11]
Mm-hmm

**Swyx** [52:12]
... uh, computer use re- recently. Um, you know, you're not saying this, but OpenAI is rumored to be working on things and like there's, a lot of labs are like exploring this-

**Karina Nguyen** [52:19]
Mm-hmm

**Swyx** [52:19]
... like sort of drive a computer generally.

**Karina Nguyen** [52:22]
Mm-hmm.

**Swyx** [52:22]
Um, how important is that for agents?

### Computer Agents

**Karina Nguyen** [52:24]
I think it would be one of the core capabilities of agents. Yeah, computer using or agents using desktop or like your computer is like the delegation part, like when you might wanna like delegate an agent to like order a book for me or like order a flight or like search for a flight and then order things for me.

And I feel like this idea was flying around like for a long time, since at least like 2022 or something. And-

**Swyx** [52:53]
Yeah

**Karina Nguyen** [52:53]
... like finally we are here. And just like there's a lot of like lag between idea and like full execution in the order of like two to three years.

**Swyx** [53:02]
Yeah.

**Karina Nguyen** [53:02]
So it's kind of cool to see-

**Swyx** [53:02]
The vision, the vision models had to get better-

**Karina Nguyen** [53:04]
Yeah

**Swyx** [53:04]
... a lot better.

**Karina Nguyen** [53:05]
Like perception and something. But I think like it, it's really cool. I feel like it's, it has like implications for like consumers definitely, like delegation. But I guess again, like I think like latency is like one of the most important factors here is like you kinda wanna make sure that the model correctly understands what you want, and then if it doesn't understand or if it doesn't know like full context, it should like ask for a follow-up question and then like use that to perform the task.

Like the agent should know if it has enough information to complete the task at the maximal, if it's the maximal success or not. And I think this is like still an open kind of like research question, I feel like.

Yeah, and the second idea is that like I think it also enables new class of like research questions of like computer use agents. Like can we use it in RL? Right? Like this is kind of like very cool, like nascent area of like research.

**Swyx** [54:00]
What's one thing that you think by the end of this year people will be using computer use agents a lot for?

**Karina Nguyen** [54:06]
I don't know. It's really hard to predict. Um-

**Swyx** [54:09]
I'm trying to look for-

**Karina Nguyen** [54:09]
Maybe for coding. I don't know. Like-

**Swyx** [54:12]
For coding?

**Karina Nguyen** [54:13]
I think like right now, like with Canvas, we are thinking about like this paradigm of like real-time collaboration to like asynchronous collaboration. So it's like it would be cool if I can just delegate to a model like, "Okay, can you figure out like how to do this feature or something?"

And then the model can just like test out that feature in its own like virtual environment or something. I don't know, like maybe this is a weird idea. Obviously, there will be a lot of use cases around like consumers, consumer use cases like, "Hey, like shop for me," or something.

**Swyx** [54:43]
I was gonna say, everyone goes to booking-

**Karina Nguyen** [54:45]
Like stay home

**Swyx** [54:45]
... booking plane tickets. That's like the worst example because you only book the plane tickets, what, two, three times a year, you know? Like-

**Karina Nguyen** [54:50]
Or concert tickets.

**Swyx** [54:51]
Yeah, yeah.

**Karina Nguyen** [54:51]
I don't know. Yeah.

**Swyx** [54:51]
Concert tickets, yeah.

**Karina Nguyen** [54:52]
Like Taylor Swift or-

**Swyx** [54:53]
I want a Facebook Marketplace bot that just-

**Karina Nguyen** [54:55]
Yeah

**Swyx** [54:55]
... scrolls Facebook Marketplace for free stuff.

**Karina Nguyen** [54:57]
Yeah.

**Swyx** [54:58]
And then just go and get it.

**Karina Nguyen** [54:59]
Yeah. I don't know. What do you think?

**Swyx** [55:02]
I have been very bearish in computer use-

**Karina Nguyen** [55:05]
Oh, cool

**Swyx** [55:05]
... because they're, they're slow, they're expensive-

**Karina Nguyen** [55:07]
Mm-hmm

**Swyx** [55:07]
... they're imprecise. Like, the-

**Karina Nguyen** [55:08]
Yeah

**Swyx** [55:08]
... the accuracy is horrible.

**Karina Nguyen** [55:10]
Mm-hmm.

**Swyx** [55:10]
Uh, still, even with Anth- Anthropic's new stuff, I'm really waiting to see what OpenAI might do to change my opinions. And really what I'm trying to do is, like, Jan last year versus December last year, I changed a lot of opinions.

What am I wrong about today? And computer use is probably one of them, where I'm like, I don't think... I don't know if by end of the year we'll still be using them. Will my ChatGPT have-- Like, every GPT instance, will they, will they have a virtual computer maybe?

I don't know.

**Karina Nguyen** [55:35]
Mm-hmm.

**Swyx** [55:35]
Coding, yes, 'cause he, he invested in a company-

**Karina Nguyen** [55:38]
Mm. Mm

**Swyx** [55:38]
... that, that does, does that.

**Karina Nguyen** [55:39]
That's cool.

**Swyx** [55:39]
The, the code sandboxes. There, there are a bunch of code sandbox companies. E2B is the name. But then, like, in browsers, yes.

**Karina Nguyen** [55:44]
Mm-hmm.

**Swyx** [55:45]
Computer use is, like, coding plus browsers plus everything else.

**Karina Nguyen** [55:47]
Mm-hmm. Mm-hmm.

**Swyx** [55:48]
There's a whole operating system, and it's very... Like, you have to be pixel precise. You have to OCR. Well, I think OCR is basically solved.

**Karina Nguyen** [55:54]
Mm-hmm.

**Swyx** [55:54]
But, like, pixel precise and, like, understand the UI of what you're operating.

**Karina Nguyen** [55:57]
Mm-hmm.

**Swyx** [55:58]
And, like, I don't know if the models are there yet. I don't know.

**Karina Nguyen** [56:02]
Yeah. Yeah. Two questions. Like, do you think the progress of, like, mini models like o3-mini or like o1-mini... I guess like it's came back to, like, the Claude, Claude 3 Haiku, Claude 1.2 Instant, like, this, like, gradual progression of, like, small models becoming really powerful, which are very also, like, fast.

Like, I'm sure, like, the computer use agents, like, would be able to, like, couple with, like, those, like, small models. That will solve some of the agency issues in my opinion. I think in terms of, like, other operating system, I think a lot about it.

I-- These days, it's just like we're entering this, like, task-oriented, like, operating system or something where also generative OS. Like, in my opinion, like, people in, like, few years will click on, like, websites way less. I wanna see the plot of, like, website clicks over time.

But then my prediction is, like, it will go down and, like, people's access to the internet will be through the model's lens.

**Swyx** [57:05]
Mm-hmm.

**Karina Nguyen** [57:05]
Either you see what the model's doing or you don't see what the model's doing-

**Swyx** [57:10]
Mm-hmm

**Karina Nguyen** [57:10]
... on the internet.

**Swyx** [57:11]
Yeah. I think my personal benchmark for computer use this year is expense reports.

**Karina Nguyen** [57:16]
Mm-hmm.

**Swyx** [57:16]
So I have to do my expense report every month.

**Karina Nguyen** [57:18]
Oh, yeah.

**Swyx** [57:19]
But what you need to do... So for example, I expense a lunch.

**Karina Nguyen** [57:23]
Mm-hmm.

**Swyx** [57:23]
I have to go back on the calendar and see who I was having lunch with.

**Karina Nguyen** [57:25]
Mm-hmm.

**Swyx** [57:26]
Then I need to upload the receipt-

**Karina Nguyen** [57:28]
Yeah

**Swyx** [57:28]
... of the lunch, and I need to tag the person, the expense report, blah, blah, blah.

**Karina Nguyen** [57:31]
Yeah.

**Swyx** [57:32]
It's very simple on a task-by-task basis.

**Karina Nguyen** [57:34]
Yeah.

**Swyx** [57:34]
But, like, you have to go to every app-

**Karina Nguyen** [57:36]
Right

**Swyx** [57:37]
... that I use. You have to go to, like, the, you know, Uber app. You have to go to the-

**Karina Nguyen** [57:40]
Yeah

**Swyx** [57:40]
... camera roll to get-

**Karina Nguyen** [57:41]
Yeah

**Swyx** [57:41]
... the photo of the receipt and all these things. It's not-- You cannot actually do it today, but it feels like a tractable problem-

**Karina Nguyen** [57:47]
Mm-hmm

**Swyx** [57:47]
... you know, that probably by the end of the year we should-

**Karina Nguyen** [57:49]
Right

**Swyx** [57:49]
... be able to do it.

**Karina Nguyen** [57:50]
Yeah. This reminds me of, like, the idea of you kind of want to show to computer use agents how you would want, how you want or how you like booking your flights. It's kind of like a few shot-

**Swyx** [58:04]
Yeah, demonstration and learning

**Karina Nguyen** [58:05]
... uh, example demonstrations of, like-

**Swyx** [58:06]
Yeah

**Karina Nguyen** [58:06]
... maybe there is more efficient way that you do things that the model should learn to do it in, in that way. And so it's kinda like... Again, c-comes back to, like, personalized tasks too, is like right now a task is just, like, very, like, rudimentary.

But in the future, tasks should become, like, much more personalized for your preferences.

**Swyx** [58:28]
Okay. Well, we mentioned that... Oh, I'll, I'll also say that I think one takeaway I got from your con- this conversation is that ChatGPT will have to integrate a lot w-more with my life.

**Karina Nguyen** [58:37]
Mm-hmm.

**Swyx** [58:37]
Um, like, you, you, you will need my calendar. You will need my email.

**Karina Nguyen** [58:40]
Yes.

**Swyx** [58:40]
Like, for sure, and maybe use MCP. I don't know. Have, have you looked at MCP?

**Karina Nguyen** [58:44]
No, I haven't.

**Swyx** [58:45]
Uh, it's good.

**Karina Nguyen** [58:46]
Cool.

**Swyx** [58:46]
It's, it's got a lot of adoption.

**Karina Nguyen** [58:47]
Okay.

**Swyx** [58:47]
Anything else that we're forgetting about or, like, maybe something that people should use more, yeah, I don't know, before we wrap on, like, the OpenAI side of things?

**Karina Nguyen** [58:56]
I think, like, search product is kinda cool, like ChatGPT Search. I think this idea of like, you know, like, right now I'm thinking a lot of was like, you know, the magic of ChatGPT when it first came out was like, you know, you ask something, any, like, instruction, and then, like, it would, like, follow the instruction that you gave to a model.

Like, like, "Write a poem," and it would give you a poem. But I think, like, the magic of the next generation of ChatGPT is, like, actually... And we are, like, we are marching towards that, is like when you ask a question, it's not just gonna be in the text output.

The ideal output might be, like, in some form of, like, a React app on the fly or something. So, like, this is happening with, like, search, right? Like, give me, like, Apple stock, and then it gives you the chart, and it gives you, like, this, like, generative UI.

And I feel like this is what I mean by, like, the evolution of ChatGPT becomes, like, more of a generative OS with a task orientation or something. So it's like... And then UI will adapt to what you like.

So, like, if you really like 3D visualizations, I think the model should give you as much visualization as possible. Like, you know, if you really like certain way of, like, the UIs, like, maybe you like round corners or I don't know, it's just, like, some color schemes that you like.

It's just, like, the UI becomes, like, more dynamic and, like, becomes like a custom, custom model, like personal model, right? Like, from personal computer to, like, a personal model, I think.

**Swyx** [1:00:21]
Yeah. Takes overall, uh, you are one of the very few people, actually maybe not that rare, uh-

### Two Cultures

**Karina Nguyen** [1:00:27]
Not anymore

**Swyx** [1:00:27]
... to, uh, to work at both OpenAI and Anthropic.

**Karina Nguyen** [1:00:29]
Not anymore, yeah.

**Swyx** [1:00:31]
Um, uh, what's the cultural difference? What are general takes that people, like, only like you see?

**Karina Nguyen** [1:00:36]
I love both places. I think I've learned so much at Anthropic, and I'm really, really grateful to the people, and I'm still, like, friends with a lot of people there. And I was really sad when John left OpenAI, uh, because I, I came to OpenAI because I wanted to work with him- ...

the most or something. Um-

**Swyx** [1:00:52]
What's he, what's he doing? What's he doing now?

**Karina Nguyen** [1:00:54]
But I think it, it changed a lot. So I think, like, when I first joined Anthropic there were, like, I don't know, 60, 70 people. When I left, there were, like, 700, like, people, so it's, like, a massive, like, growth.

OpenAI and Anthropic is different in terms of, like, more, like, maybe, like, product mindsets. Maybe OpenAI is much more willing to take some of the product risks and explore different bets. And I think Anthropic is much more focused, and they have...

I think it's, it's fine. Like, they have to, like, prioritize, but they're definitely double downing on, like, enterprise maybe m-more than, like, consumers or something. I don't know. It's just, like, some of the product mindsets might be different.

I would say, like, research, I've enjoyed, like, both, like, research cultures, both at Anthropic and, like, OpenAI. And I feel like they are more... On the daily basis, I feel like it's more f-similar than different.

**Swyx** [1:01:51]
I mean, no surprise-

**Karina Nguyen** [1:01:52]
Like, how you run experiments is kind of, like, very similar. I don't know.

**Swyx** [1:01:55]
I'm sure the Anthropic-- I mean, you know, Dario used to be VP Research, right? So-

**Karina Nguyen** [1:01:58]
Mm.

**Swyx** [1:01:59]
He set the culture at Open- OpenAI.

**Karina Nguyen** [1:02:01]
Yeah.

**Swyx** [1:02:01]
So yeah, it makes sense. Maybe quick takes on, on people that you, you mentioned, uh, Barrett, you mentioned Mira.

**Karina Nguyen** [1:02:06]
Mm-hmm.

**Swyx** [1:02:06]
Like, what's one thing you learned from Barrett, Mira, Sam, maybe?

**Karina Nguyen** [1:02:09]
Mm.

**Swyx** [1:02:09]
Something like that. Like, one lesson that you would share to others.

**Karina Nguyen** [1:02:14]
I wish I, like, worked with them way longer. I think what I've learned from Mira is actually her, like, interdisciplinary mindset. She's just really good at, like, connecting dots between, like, product and, like, kind of balancing, like, product research and, like, create this, like, comprehensive, like, coherent story because sometimes, like, there are, like, researchers who, like, really hate doing product, and there are researchers who really love doing product, and it's, like, kind of dichotomy between two.

And also, like, safety is, like, a part of this process. So kind of you, you c- you kind of want to, like, create this coherent... Like, think from, like, systems perspective or, like, think about, like, bigger picture, and I think I learned a lot from her on that.

I definitely feel like I have much more creative freedom at OpenAI, and that's because the environment that the leaders set, like, enables me to do that. So it's like, if I have an idea, if I want-

**Swyx** [1:03:10]
Just propose it, yeah.

**Karina Nguyen** [1:03:11]
Yeah, exactly.

**Swyx** [1:03:11]
On your first month.

**Karina Nguyen** [1:03:12]
So it's like there's, like, more, like, creative freedom and, like, resource reallocation, especially when researchers, like, being adaptable to, like, new technologies and, like, change your views based on, like, empirical results or kind of, like, change research directions.

I've seen a lot of, like... Sometimes I've seen researchers who would just, like, s- get stuck on the same directions for, like, two to three years, and they-- it would never, like, work out or something, but they would still be, like, stubborn.

So it's, like, adaptability to, like, new directions and, like, new paradigms is kind of, like, one of those things that I learned.

**Swyx** [1:03:43]
Th-this is a Barrett thing, or is it a general culture thing?

**Karina Nguyen** [1:03:45]
The general kind of culture-

**Swyx** [1:03:46]
Okay

### Outro

**Karina Nguyen** [1:03:46]
... I think.

**Swyx** [1:03:47]
Cool. Yeah, and just to wrap up, we just usually have a call to action. Founders usually want people to work at their companies. Do you want people to give you feedback? Do you want people to join your team?

**Karina Nguyen** [1:03:57]
Oh, yeah, of course. I'm definitely hiring for, like, research engineers who are, like, more pl- uh, product-minded people. So it's, like, people who know how to train the models but also, like, interested in, like, deploying into, like, the products and developing, like, new product features.

I'm definitely looking for tha- those archetypes of, like, research engineers or, like, research scientists. So yeah, if you're, like, looking for a job , if you're, like, interested in joining my team, I'm, like, really happy to. Just reach out, I guess.

**Swyx** [1:04:25]
And then just, like, generally, what do you want people to do more of in the world, whether or not-

**Karina Nguyen** [1:04:29]
Oh

**Swyx** [1:04:29]
... they, they work with you? Like, you know, call to action as in, like, "Everyone should be doing this."

**Karina Nguyen** [1:04:33]
I think this is something that I tell to a lot of, like, designers is that, like, I think people should, like, spend more time just, like, play around with the models. And the more you play with the model, the more creative ideas you'll get around, like, what kind of, like, new potential features of the products or, like, new kind of interaction paradigms that you might wanna create with those models.

I feel like we are bottlenecked by, like, human creativity on, like, completely changing the way we think about the internet or, like, some of the, the way we think about software. Like, AI right now pushes us to, like, rethink everything that we've done before, in my view, and I feel like not enough people either double down on, like, those ideas or I'm just, like, not seeing a lot of, like, human creativity in this, like, interface design or, like, product design mindsets.

So I feel like it'd be really great for people to just, like, do that, and especially right now with, like, research. Some research becomes, like, much more product oriented, so it's like you actually can train the models for the things that you want to do in the product or something.

**Swyx** [1:05:41]
Yeah. And you define the process now. Now this is my go-to for-

**Karina Nguyen** [1:05:45]
Okay.

**Swyx** [1:05:45]
... how to-

**Karina Nguyen** [1:05:46]
Yeah

**Swyx** [1:05:46]
... manage a process. I think it's pretty common sense, but it's nice to hear-

**Karina Nguyen** [1:05:49]
Mm-hmm

**Swyx** [1:05:49]
... from you that-- 'cause you actually did it. That, that's nice. Thank you for driving innovation des- uh, interface design in, in the new models at OpenAI and Anthropic, and, uh, we're looking forward to what you're gonna talk about in, in New York.

**Karina Nguyen** [1:06:01]
Yeah. Thank you so much for inviting me here.

**Swyx** [1:06:04]
Yeah.

**Karina Nguyen** [1:06:04]
I hope my job will not be automated by the time I come to New York.

**Swyx** [1:06:08]
Well, I hope you automate your, yourself, and-

**Karina Nguyen** [1:06:10]
Yeah, I hope so too. Yeah

**Swyx** [1:06:10]
... you can go do whatever else you wanna do. That's it. Thank you.

**Karina Nguyen** [1:06:13]
Awesome. Thanks.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
