# Cooking with OpenAI’s Research Chief: AGI, o1, Evals, and Scaling Laws — Mark Chen

Latent Space · 2026-06-25

<https://addtry.com/e9cba669-0846-49ee-a486-90cb75e2bc65>

Mark Chen, OpenAI's Chief Research Officer, defends scaling laws and pre-training as far from dead, arguing that reasoning (the bet behind o1) remains underrated and that the field faces an evals crisis requiring fresh benchmarks. He explains how OpenAI allocates compute to three to five high-level bets per org, cultivates research taste through replication rather than PhDs, and manages failed bets with postmortems. Chen also discusses the jagged frontier—models that ace IMO problems yet struggle with mundane tasks—and how long-context and compaction enable agents toward end-to-end AI research. Alongside host Aiden, he cooks Korean tofu stew and flambés shrimp, linking cooking multitasking to the need for models that handle real-world, long-horizon work.

## Questions this episode answers

### Why does Mark Chen think pre-training is not dead?

Mark Chen, OpenAI’s Chief Research Officer, firmly disagrees with the “pre-training is dead” narrative, citing scaling laws that have held for nearly ten orders of magnitude. He argues that claims of insurmountable bottlenecks are recurring, but better engineering and research insights always break through. With more careful data engineering and scaling, he believes pre-training will continue unlocking new capabilities, just as it did when reasoning models like o1 overcame initial skepticism.

[9:13](https://addtry.com/e9cba669-0846-49ee-a486-90cb75e2bc65?t=553000)

### How does Mark Chen suggest developing research taste without a formal machine learning background?

Mark Chen advises replicating admired papers like ResNet or PixelCNN, matching training curves exactly to learn unspoken techniques. He credits watching AlphaGo beat Lee Sedol as his turning point, inspiring his first big project of getting a DQN working. He says OpenAI values creative problem-solving and the ability to think outside the box over formal PhD training, and has a strong track record of training people up internally.

[3:33](https://addtry.com/e9cba669-0846-49ee-a486-90cb75e2bc65?t=213000)

### How does OpenAI allocate compute to research projects?

Mark Chen explains that OpenAI uses directive compute allocation, assigning large compute blocks to top-priority bets while giving research leads flexible compute pools they can freely allocate to promising ideas. The high-level roadmap stays stable, but implementation details and resourcing are regularly reassessed during compute allocation cycles. This balances top-down steering with bottom-up creativity, focusing on a few bets per org while allowing researchers to pursue high-conviction projects.

[13:33](https://addtry.com/e9cba669-0846-49ee-a486-90cb75e2bc65?t=813000)

## Key moments

- **[0:00] Soup Story**
  - [0:28] Mark Chen's soup story: He brought soup to OpenAI researchers after Meta's poaching attempts with soup.
- **[1:52] Trading to AI**
  - [1:52] Q: How does trading prepare someone for AI research? Mark Chen: Trading teaches attention to detail and optimization, transferable skills.
  - [3:21] Mark Chen advises replicating papers like ResNet and PixelCNN to develop research taste.
  - [4:26] Mark Chen says AlphaGo's Move 37 against Lee Sedol was a turning point for many, leading him to first implement DQN.
- **[5:23] RL & Evals**
  - [5:23] Mark Chen: RL struggles in subjective domains like creative writing where grading is hard, but excels in math and coding.
  - [6:59] Mark Chen: To evaluate superhuman models, we move to real-world research; models now discover novel theorems and connect fields.
- **[8:54] Scaling & Reasoning**
  - [8:54] Mark Chen: 'Pre-training is not dead' — scaling laws have held for ten orders of magnitude and will continue.
  - [10:26] Mark Chen recounts how the o1 reasoning bet was hard to launch internally despite pre-training's success.
  - [11:12] Mark Chen: OpenAI research managers steer with earned respect; bottom-up ideas backed by evidence also shape the roadmap.
- **[12:33] Research Roadmap**
  - [12:33] Mark Chen: OpenAI's high-level research roadmap stays stable, but implementation details and resource allocation adapt.
  - [13:33] Mark Chen describes OpenAI's compute allocation: directive bets get large chunks, but managers have flexible pools.
- **[15:48] Great Researchers**
  - [15:48] Mark Chen: It's hard to spot great researchers at interview; after 6-12 months, intuition emerges, and not all impact is the same.
  - [17:44] Mark Chen: Top researchers differ from engineers by needing research taste and the ability to convince others.
- **[19:33] Evals Crisis**
  - [19:33] Mark Chen: AI is in an 'evals crisis'; canonical benchmarks are saturated, and overfitting on distributions is easy.
  - [21:32] Mark Chen: OpenAI separates evals team from model team to avoid co-incentives; partners with external orgs for gold-standard evals.
  - [23:40] Mark Chen quotes Jakub Pachocki: 'It feels like I have an army of really dumb IOI gold medalists' as researchers.
  - [24:34] Mark Chen on jagged intelligence: Models lack human context, making them miss easy tasks despite excelling at hard ones.
- **[25:56] Long Context**
  - [25:56] Mark Chen: Naively increasing context windows isn't enough; compaction techniques shortcut long-horizon learning challenges.
  - [28:36] Mark Chen predicts few big bets remain before AGI; models will soon do self-sustained research.
- **[28:37] Future Research**
  - [31:32] Mark Chen: OpenAI prefers one architecture for all modalities to reduce infrastructure cost and enable core research carryover.
  - [32:36] Mark Chen: The future of research is 'vibe researching' — models execute ideas, but researchers still provide taste.
- **[34:36] Failed Bets**
  - [34:36] Mark Chen: OpenAI's alpha is high-risk bets; failed bets yield valuable write-ups that prevent others from repeating mistakes.
- **[37:53] Overrated/Underrated**
  - [37:53] Mark Chen calls pre-training 'underrated' and ties between research primitives and real-world products also underrated.
- **[38:42] Taste Test**

## Speakers

- **Aiden** (host)
- **Sean** (host)
- **Mark Chen** (guest)

## Topics

Evals, Reasoning, Agent Infrastructure

## Mentioned

Meta (company), OpenAI (company), AlphaGo (product), ChatGPT (product), Codex (product), o1 (product)

## Transcript

### Soup Story

**Aiden** [0:01]
Woo. Yeah, I was like-

**Mark Chen** [0:04]
Thank you. That, I know.

**Aiden** [0:06]
That feels like the current situation I'm in. Like... Cheers.

**Mark Chen** [0:11]
Cheers.

**Aiden** [0:13]
Hey, guys. Welcome to the Latent Space cooking series, where we invite founders and researchers and just let them cook. Today, we have a very special guest, the chief research officer of OpenAI, Mark Chen. Welcome, Mark.

**Mark Chen** [0:25]
Thanks for inviting me, Aiden.

**Aiden** [0:26]
Yeah, thank you for coming. I mean, to begin-

**Mark Chen** [0:28]
Mm-hmm

**Aiden** [0:28]
... this all started from the inspiration after hearing a story that Mark Zuckerberg would make soup to try to poach researchers, and in response, you brought soup to researchers. Is this true? Did this happen? Did it work?

**Mark Chen** [0:39]
Oh, you know-

**Aiden** [0:40]
Can you elaborate?

**Mark Chen** [0:40]
... it's absolutely a true story. Uh-

**Aiden** [0:42]
Okay

**Mark Chen** [0:42]
... and I have brought soup to our own researchers. Um, I think that Meta's calmed down a little bit.

**Aiden** [0:46]
Yeah.

**Mark Chen** [0:46]
I think we came out on top, but, um, yeah, still a very funny story, and the craziness of how AI has evolved. Mm-hmm.

**Aiden** [0:52]
How often do you cook? Is this something you're familiar with?

**Mark Chen** [0:54]
Well, you know, I, I do enjoy cooking.

**Aiden** [0:56]
Mm-hmm.

**Mark Chen** [0:56]
But I don't have the luxury of doing that so often, so I, uh, usually have a work dinner every night of the week, and, you know, maybe post-AGI this is gonna be my hobby. I've always joked- ... you know, I'm gonna start a noodle stand once, uh, once it's, it's all over.

**Aiden** [1:11]
Yeah, yeah. You know, post-AGI, hopefully that'll, you know, still be there.

**Mark Chen** [1:14]
Mm-hmm.

**Aiden** [1:14]
But great. And I guess looking at what we have in front of us, do you have an idea generally of what we'll probably making?

**Mark Chen** [1:20]
Uh, Korean tofu soup maybe.

**Aiden** [1:22]
Yeah, yeah. That's, that's generally what it is.

**Mark Chen** [1:24]
Great. Okay.

**Aiden** [1:24]
So we inspired off of the story-

**Mark Chen** [1:26]
Okay

**Aiden** [1:26]
... of you, you know, bringing soup to researchers.

**Mark Chen** [1:28]
Mm-hmm, mm-hmm.

**Aiden** [1:28]
So we're making a tofu Korean stew.

**Mark Chen** [1:30]
Okay.

**Aiden** [1:30]
And then we have problems that we'll be cooking. Are you ready to go?

**Mark Chen** [1:32]
Yeah, let's do it. Let's do it.

**Aiden** [1:33]
Great. Okay.

**Mark Chen** [1:34]
Mm-hmm.

**Aiden** [1:34]
So the first thing we should probably do is we'll separate the veggies, and then we can cut them.

**Mark Chen** [1:38]
Okay.

**Aiden** [1:38]
And basically what we wanna do is just, um, cut the dirty part off with the dirt and then, yeah, separate that across-

**Mark Chen** [1:45]
That one, that I know.

**Aiden** [1:46]
Yeah. Okay. And then just have that there, so-

**Mark Chen** [1:49]
Okay

**Aiden** [1:49]
... you could do that. And while that's going, I guess I could ask more about your background.

### Trading to AI

**Mark Chen** [1:52]
Mm-hmm.

**Aiden** [1:53]
So in a previous life, you were once a trader, and even Sam, I think, last year in April, also tweeted about-

**Mark Chen** [1:59]
Mm-hmm

**Aiden** [1:59]
... how if you're a high-frequency trader, you should consider joining OpenAI-

**Mark Chen** [2:03]
Mm-hmm

**Aiden** [2:03]
... because, you know, build AGI.

**Mark Chen** [2:05]
Mm-hmm.

**Aiden** [2:05]
So do you think there's a relation between, you know, being a trader and being a researcher, or do you think it's just, like, a very technical and competitive area where a lot of great employees can come from? Can you expand on that background?

**Mark Chen** [2:15]
I think really the most important thing is, um-

**Aiden** [2:17]
Hmm

**Mark Chen** [2:17]
... there are a lot of researchers who just, uh, started out without a formal training in machine learning or AI research.

**Aiden** [2:23]
Gotcha.

**Mark Chen** [2:23]
Um, we've very much believed in training people up to do this. I, I think the real hard thing is the ability to creatively solve problems and think outside of the box.

**Aiden** [2:33]
Yeah.

**Mark Chen** [2:33]
And it's not so much, you know, you have to do a PhD, even though that does bring a valuable skill set.

**Aiden** [2:38]
Mm-hmm.

**Mark Chen** [2:38]
Um, with trading in particular, I mean, I, I don't know that it's that special of a profession. Like, um-

**Aiden** [2:44]
Yeah

**Mark Chen** [2:44]
... I kinda think of it as, uh, you know, we've had great mathematicians join, we have had great, you know, physicists join, but trading is something where, you know, it's like, it's very unhackable.

**Aiden** [2:54]
Mm-hmm.

**Mark Chen** [2:54]
You know, you're, um... How... You, you can't kind of cheat the real world, right?

**Aiden** [2:58]
Yeah.

**Mark Chen** [2:59]
Like, uh, you know, it's a, it, it's a hard metric to optimize.

**Aiden** [3:01]
Mm-hmm.

**Mark Chen** [3:01]
Um, and there's also a lot of characteristics, like, uh, it's, it's a field where attention to detail really matters.

**Aiden** [3:07]
Yeah.

**Mark Chen** [3:07]
And, um, you know, it, it's kind of the brutal hard optimization and squeezing out the juice of a system.

**Aiden** [3:14]
Hmm.

**Mark Chen** [3:14]
Um, and some of those, uh, those skills transfer over.

**Aiden** [3:18]
Gotcha. Yeah.

**Mark Chen** [3:18]
Mm-hmm.

**Aiden** [3:18]
And I guess for people who want to get into research-

**Mark Chen** [3:21]
Mm-hmm

**Aiden** [3:21]
... who, let's say, don't have a PhD, what do you think are the main attributes or things that they can learn to develop research taste? Because-

**Mark Chen** [3:28]
Mm-hmm

**Aiden** [3:28]
... I guess that's the main part of, um, getting into this field that may be very foreign to them.

**Mark Chen** [3:33]
Yeah. I mean, I think it's a little bit overrated.

**Aiden** [3:36]
Hmm.

**Mark Chen** [3:36]
Um, it is something you have to develop, but, um, the best mechanism I've found for developing that is really just, um, replication. So I think you should take papers that you really look up to-

**Aiden** [3:48]
Yeah

**Mark Chen** [3:48]
... and just try to fully replicate it.

**Aiden** [3:50]
Gotcha.

**Mark Chen** [3:50]
Um, I, like, I, I think a lot of replications stood out in mind, my mind.

**Aiden** [3:54]
Hmm.

**Mark Chen** [3:54]
Um, you know, back in 2018, there was, um, uh, you know, like ResNet, there were PixelCNNs, and-

**Aiden** [4:02]
Mm-hmm

**Mark Chen** [4:02]
... I think I learned so much just, um, trying to replicate the training curves exactly, get to the exact amount of, you know, like, training loss or perplexity that the paper is, um, hinted towards.

**Aiden** [4:11]
Yeah.

**Mark Chen** [4:11]
It teaches you a lot of techniques, right-

**Aiden** [4:13]
Mm-hmm

**Mark Chen** [4:13]
... that, um, people don't really kinda talk about. But, you know, once, once you dive in a couple layers deeper, um, you learn those techniques. And, um, yeah, I think really the first thing, too, that got me into the field was-

**Aiden** [4:26]
Yeah

**Mark Chen** [4:26]
... when AlphaGo played Lee Sedol.

**Aiden** [4:28]
Hmm.

**Mark Chen** [4:29]
And, you know, I think that was a turning point for so many people.

**Aiden** [4:31]
Yeah.

**Mark Chen** [4:32]
And, yeah, I, I mean, it was, it was, it was inspirational, and it, it... The, the first big project that I, I really went after was, um, can I get a DQN working?

**Aiden** [4:40]
Hmm.

**Mark Chen** [4:40]
Yeah.

**Aiden** [4:41]
Yeah. Yeah, that's true.

**Mark Chen** [4:42]
Mm-hmm.

**Aiden** [4:42]
I think it was Move 37 or-

**Mark Chen** [4:43]
Yeah. Yeah

**Aiden** [4:44]
... from, like, when the games was pretty, uh, insane watching it happen, and-

**Mark Chen** [4:47]
Yeah. Yeah

**Aiden** [4:47]
... seeing all that develop, and see also where we have gotten to today, especially with research.

**Mark Chen** [4:51]
I mean, isn't it crazy that you're seeing Move 37s in almost every field now?

**Aiden** [4:56]
Yeah.

**Mark Chen** [4:56]
It's like there's Move 37s in, in math. There's, uh, in computer science and coding. Um-

**Aiden** [5:02]
Mm-hmm

**Mark Chen** [5:02]
... I think even... Yeah, just it feels like a lot of people woke up at the start of this year and were like, "Man, agents are working in my profession."

**Aiden** [5:10]
Yeah.

**Mark Chen** [5:10]
And, um, you know, they're, they're essentially realizing that these models can just do long-horizon meaningful work for them.

**Aiden** [5:17]
Yeah. No, that's true.

**Mark Chen** [5:18]
Mm-hmm.

**Aiden** [5:18]
It is, it is very impressive to see-

**Mark Chen** [5:20]
Yeah

**Aiden** [5:20]
... like, even just using it in my own work. But okay, the next thing we could do is just as simple-

### RL & Evals

**Mark Chen** [5:23]
Yeah

**Aiden** [5:23]
... just cutting the onion.

**Mark Chen** [5:24]
Okay. Great.

**Aiden** [5:24]
So what we have to do here is just, like, dicing it.

**Mark Chen** [5:26]
Yeah.

**Aiden** [5:26]
Do you think, um, there's jobs that RL maybe will have, like, a much harder time to kind of break into? So for example-

**Mark Chen** [5:35]
Mm-hmm

**Aiden** [5:35]
... coding may be easier since a lot of the context is accessible, whether it be the code bases or even the work you're trying to do. But-

**Mark Chen** [5:41]
Mm-hmm

**Aiden** [5:41]
... let's say if you're trying to do the job that a junior consultant may do-

**Mark Chen** [5:44]
Mm-hmm

**Aiden** [5:44]
... where all the context is a little scattered-

**Mark Chen** [5:46]
Mm-hmm

**Aiden** [5:46]
... maybe a little more difficult, how do you view through, like, those different scenarios? Is there a way that you kind of assess what can be the right approach?

**Mark Chen** [5:54]
Yeah, I mean, I, I think it's... RL's traditionally had, um, headwinds when it's come to fields that, you know, it's more, um, kind of- ... subjective than objective.

**Aiden** [6:08]
Mm.

**Mark Chen** [6:08]
So if you kind of think of like, you know, one, one kind of, you know, uh, example of this is creative writing.

**Aiden** [6:15]
Yeah.

**Mark Chen** [6:15]
Where, you know, you can take two pieces of creative writing, and two experts can have wildly different opinions.

**Aiden** [6:21]
Yeah.

**Mark Chen** [6:21]
So it's, it's these fields where things are hard to grade.

**Aiden** [6:24]
Mm-hmm.

**Mark Chen** [6:24]
Um, where, you know, RL has the least amount of ability to kind of go and, um, and directly apply there.

**Aiden** [6:31]
Yeah.

**Mark Chen** [6:31]
I know a lot of people are developing techniques to apply RL in these, um, these settings, but, um, for now, it's just where there's cold, hard truth-

**Aiden** [6:39]
Mm-hmm

**Mark Chen** [6:39]
... things like math and computer science, where-

**Aiden** [6:41]
Yeah

**Mark Chen** [6:41]
... you implement it correctly or wrong. Um-

**Aiden** [6:43]
Mm

**Mark Chen** [6:43]
... that's where you kind of see it really taking off.

**Aiden** [6:46]
Yeah. No, that actually-

**Mark Chen** [6:48]
Mm

**Aiden** [6:48]
... brings up a thought on in terms of evaluating-

**Mark Chen** [6:51]
Mm-hmm

**Aiden** [6:52]
... those fields, so-

**Mark Chen** [6:53]
Yep

**Aiden** [6:53]
... um, you know, as models get much pow- much more powerful-

**Mark Chen** [6:56]
Mm-hmm

**Aiden** [6:56]
... and even saturate, for example, solving-

**Mark Chen** [6:58]
Mm

**Aiden** [6:58]
... like the IMO questions.

**Mark Chen** [6:59]
Yeah, yeah.

**Aiden** [6:59]
Um, how do you view evaluating like superhuman intelligence, like get to a point where it's so good at things that even the top, what, .01% of humans can do?

**Mark Chen** [7:08]
Yeah.

**Aiden** [7:08]
But like, you know, how can we push past that frontier of intelligence?

**Mark Chen** [7:12]
No, it's kind of, it's kind of crazy, and I feel like, um, a lot of it centers in on in, on kind of interfacing with the real world. And-

**Aiden** [7:19]
Gotcha

**Mark Chen** [7:19]
... um, when, when we've thought about how to evolve past things like programming contests in the past-

**Aiden** [7:24]
Mm-hmm

**Mark Chen** [7:24]
... um, I think a lot of the initial direction we took was you should move it to real-world research, right? And-

**Aiden** [7:31]
Hmm

**Mark Chen** [7:31]
... we've seen that the models, uh, they've gotten a lot better at, uh, just kind of discovering novel theorems and, uh, pushing the frontiers of, of hard sciences.

**Aiden** [7:39]
Yeah.

**Mark Chen** [7:39]
But even today, right, that's no longer a surprise. I think like you- we, we almost take it for granted now that, um, these, these models can solve very, very difficult problems. They can make contributions and even kind of draw relationships between, um, fields that, you know, um, that are, that are novel and insightful.

**Aiden** [7:57]
Yeah.

**Mark Chen** [7:57]
So I think, um, you know, we, we think of coding, co-working, um-

**Aiden** [8:02]
Mm-hmm

**Mark Chen** [8:03]
... as really a, a domain for, that, that tests if our models can learn in high-context settings-

**Aiden** [8:09]
Mm

**Mark Chen** [8:09]
... and in real-world long horizon settings.

**Aiden** [8:11]
Gotcha.

**Mark Chen** [8:12]
Mm-hmm.

**Aiden** [8:12]
Okay. Yeah, that makes sense. And since you're done with all the vegetables-

**Mark Chen** [8:15]
Mm-hmm

**Aiden** [8:15]
... we can now do the next step, which is-

**Mark Chen** [8:16]
Okay, great

**Aiden** [8:17]
... sautéing it, so.

**Mark Chen** [8:18]
Cool.

**Aiden** [8:18]
Yeah, we can use the Impulse Stove, which-

**Mark Chen** [8:20]
Yep

**Aiden** [8:20]
... we've seen before, and it's very powerful. Let me-

**Mark Chen** [8:22]
Mm-hmm

**Aiden** [8:23]
... turn it on. Um, and yeah, so we'll just sauté it-

**Mark Chen** [8:26]
Great

**Aiden** [8:26]
... with some oil. So yeah, put the pan in the front burner and then yeah, all-

**Mark Chen** [8:30]
Super cool stoves.

**Aiden** [8:31]
Yeah. You could use oil to pour some in.

**Mark Chen** [8:33]
Okay, great. Mm-hmm.

**Aiden** [8:33]
And then we can also... Yeah, just a good dollup, perfect.

**Mark Chen** [8:38]
Nice.

**Aiden** [8:38]
And yeah, and then we could turn on the stove. So just press it, and then-

**Mark Chen** [8:41]
Great

**Aiden** [8:41]
... um, yeah, spin the knob.

Great, perfect. And then while that heats up-

**Mark Chen** [8:49]
Mm-hmm

**Aiden** [8:49]
... we can just wait and then add the vegetables. But yeah, I guess more so on views for research.

**Mark Chen** [8:54]
Mm-hmm.

**Aiden** [8:54]
Are there, I guess, you know, commonly accepted ideas that are, you know, you disagree with, whether it be like pre-training is dead or language models will never get us to AGI? I think there's a lot of takes out there that-

### Scaling & Reasoning

**Mark Chen** [9:06]
Mm-hmm

**Aiden** [9:06]
... are very ambiguous-

**Mark Chen** [9:07]
Yeah

**Aiden** [9:07]
... and obviously haven't been proven out yet. And I guess from your perspective as, like, the research, like leading things in OpenAI, like think through those.

**Mark Chen** [9:13]
Yeah. I mean, I, uh, I firmly believe in exponent- uh, being on the exponential and in scaling laws.

**Aiden** [9:19]
Yeah.

**Mark Chen** [9:19]
So I think any of these bear takes, um, I fairly strongly disagree with.

**Aiden** [9:24]
Mm.

**Mark Chen** [9:25]
Um, you know, when it comes to pre-training is dead, I... I mean, I think the, the funny thing is this narrative only started spreading more widely after, let's say, um, the last one or two years or so.

**Aiden** [9:37]
Yeah.

**Mark Chen** [9:37]
But in many times, uh, i- in the history of, uh, developing LLMs, people have been saying this, right?

**Aiden** [9:42]
Mm.

**Mark Chen** [9:43]
And, you know, um, there, there have always been some, some bottlenecks that people, "Well, you can't scale past this because of this bottleneck." Um, and we've always found some kind of technique, whether it be better engineering or some new research insight that helps-

**Aiden** [9:56]
Yeah

**Mark Chen** [9:56]
... you break past the boundary. And so I think it's just more and more of the same, right? Like more careful research engineering, more careful data engineering-

**Aiden** [10:03]
Mm-hmm

**Mark Chen** [10:03]
... more careful scaling, and it always unlocks that next ability to scale further.

**Aiden** [10:08]
Yeah.

**Mark Chen** [10:08]
So I, I mean, it's held for, you know, almost ten orders of magnitude, but there's no reason it should- ... not keep, keep holding.

**Aiden** [10:16]
Yeah, that's a very fair point.

**Mark Chen** [10:17]
Yeah. Yeah.

**Aiden** [10:17]
And I guess on research bets that have-

**Mark Chen** [10:18]
Mm-hmm

**Aiden** [10:18]
... helped you scale beyond-

**Mark Chen** [10:20]
Mm-hmm

**Aiden** [10:20]
... were there specific ideas that you can even remember in the early days that everyone was, was saying that is not gonna work? Or-

**Mark Chen** [10:26]
Well, yeah, I mean, I think of reasoning as one of the biggest-

**Aiden** [10:29]
Okay, reasoning, yeah

**Mark Chen** [10:29]
... examples of this. Yeah, and, um, you know, the, the first breakthrough that we launched to the world here was o1.

**Aiden** [10:34]
Mm-hmm.

**Mark Chen** [10:34]
But it wasn't easy to get that off the ground because, one, the world we were back, living in back then-

**Aiden** [10:40]
Yeah

**Mark Chen** [10:40]
... it was one where pre-training plus post-training, right?

**Aiden** [10:43]
Mm-hmm.

**Mark Chen** [10:43]
That felt like such a promising paradigm.

**Aiden** [10:45]
Yeah.

**Mark Chen** [10:46]
Um, and so even at a company like OpenAI-

**Aiden** [10:49]
Mm-hmm

**Mark Chen** [10:49]
... you would have people ask naturally, "Why do something when you have a machine that works?"

**Aiden** [10:55]
Mm.

**Mark Chen** [10:56]
And fundamentally, you know, it's to the credit of, you know, Jakub, Ilya-

**Aiden** [11:00]
Yeah

**Mark Chen** [11:00]
... many of the people who really had conviction and vision in this space-

**Aiden** [11:04]
Mm-hmm

**Mark Chen** [11:04]
... um, that we started pushing on this in earnest. And even then, it took a lot of steering to get-

**Aiden** [11:08]
Mm

**Mark Chen** [11:09]
... the whole company behind this as a, as a fundamental bet.

**Aiden** [11:12]
Gotcha.

**Mark Chen** [11:12]
So, yeah.

**Aiden** [11:12]
And how do you kind of develop that ability to motivate researchers? 'Cause I assume that's a big part of, you know, taking a lot of bets, and some-

**Mark Chen** [11:19]
Mm-hmm

**Aiden** [11:19]
... will pan out, but still building the trust in the team to know that eventually some of these will actually have, you know, power law effects.

**Mark Chen** [11:25]
You know, what's, what's really cool about OpenAI is, um, research, it feels like a meritocracy.

**Aiden** [11:29]
Mm.

**Mark Chen** [11:29]
So, um, oftentimes, the research managers are the people who, um-

**Aiden** [11:35]
Do the actual-

**Mark Chen** [11:36]
... who have done the best research-

**Aiden** [11:37]
Yeah

**Mark Chen** [11:37]
... in the past. And so-

**Aiden** [11:38]
Okay

**Mark Chen** [11:38]
... I think a lot of steering can come top-down, right? Like, if your manager says, "Hey, you know, I'm like really convinced this is the path forward"-

**Aiden** [11:46]
Yeah

**Mark Chen** [11:46]
... um, generally, people will take that into heavy consideration, right?

**Aiden** [11:51]
Yeah.

**Mark Chen** [11:51]
Um, it's like, you know, this person who you've respected for their research taste and execution for so long-

**Aiden** [11:56]
Yeah

**Mark Chen** [11:56]
... is, like, now very excited by this idea.

**Aiden** [11:58]
Mm-hmm.

**Mark Chen** [11:58]
Um, it's, it's definitely something that, uh, that, um, yeah, you, you... people take into account. So-

**Aiden** [12:04]
Mm

**Mark Chen** [12:04]
... I think there, there's good top-down steering. At the same time, you know, I think one really cool thing about OpenAI is, um, there are bottom-up elements.

**Aiden** [12:11]
Mm.

**Mark Chen** [12:11]
Like, we like to be convinced that-

**Aiden** [12:12]
Yeah

**Mark Chen** [12:13]
... um, uh- You know, that we're wrong, right?

**Aiden** [12:16]
Mm-hmm.

**Mark Chen** [12:17]
And, and someone can just come with cold, hard evidence.

**Aiden** [12:19]
Mm-hmm.

**Mark Chen** [12:19]
And many things like that have turned into core parts of our research roadmap. Just things that no one was really kind of trying to steer, but some researcher on the ground had a heavy conviction in.

**Aiden** [12:29]
Yeah.

**Mark Chen** [12:30]
Um, and, and that's also a really d- big delight to see.

**Aiden** [12:32]
Yeah, no.

**Mark Chen** [12:33]
Yeah.

### Research Roadmap

**Aiden** [12:33]
Absolutely. I heard in a recent interview-

**Mark Chen** [12:36]
Mm-hmm

**Aiden** [12:36]
... that you gave that your internal research roadmap hasn't really changed.

**Mark Chen** [12:39]
Mm-hmm.

**Aiden** [12:40]
Um, even through all that we've seen-

**Mark Chen** [12:42]
Mm-hmm

**Aiden** [12:42]
... you know, with model development and even other companies.

**Mark Chen** [12:45]
Yep.

**Aiden** [12:45]
I guess, how often do you guys assess that, reassess that, even, like, act proactively? I assume it's not a lot of reactive-

**Mark Chen** [12:52]
Mm-hmm

**Aiden** [12:52]
... you know, decision making-

**Mark Chen** [12:53]
Yep

**Aiden** [12:53]
... as, like, other models come out.

**Mark Chen** [12:54]
Yep, yep.

**Aiden** [12:55]
But how do you, like, think through that process, especially as everything around you just continues to get better?

**Mark Chen** [12:59]
Yeah. So I think the thing is, um, the high level researcher map should be staying, right?

**Aiden** [13:03]
Okay.

**Mark Chen** [13:03]
I think people need something to ground in. People need to see a path to what we're building. Um, and I've been very happy that we've stayed the course for, for a while.

**Aiden** [13:12]
Yeah.

**Mark Chen** [13:12]
But the implementation details can change over time, right?

**Aiden** [13:15]
Yeah.

**Mark Chen** [13:15]
And I think, um, it's important to kind of... Like, the, the sequencing will matter, right?

**Aiden** [13:19]
Yeah.

**Mark Chen** [13:19]
The relative resourcing will matter, and it, the, the kind of exact vets on the ground will matter.

**Aiden** [13:24]
Yeah.

**Mark Chen** [13:24]
So what we do is, um, I think we have kind of points in time that force us to reconsider these things. So, uh, one example is when we do compute.

**Aiden** [13:33]
Yeah.

**Mark Chen** [13:33]
Um, one of the parts of the job is just figuring out how to allocate compute to projects.

**Aiden** [13:37]
Okay.

**Mark Chen** [13:38]
And, um, it's, it's a time to kind of question, like, are we really putting compute to use and people to use at the highest priority bets?

**Aiden** [13:46]
Yeah. And I guess could you clarify more what you mean by-

**Mark Chen** [13:50]
Need some oil.

**Aiden** [13:50]
Oh, yeah. Yeah.

**Mark Chen** [13:51]
Thank you.

**Aiden** [13:51]
I was like, having coffee, adding oil. But clarify what you mean by the higher level versus, like, the more implementation details. Like-

**Mark Chen** [13:58]
Yeah, yeah

**Aiden** [13:58]
... as high level, as general as, like, AGI.

**Mark Chen** [14:00]
Mm-hmm.

**Aiden** [14:00]
That's like our North Star, or is it more, like, granular than that?

**Mark Chen** [14:03]
Um, well, yeah, I mean, at the very highest level, right?

**Aiden** [14:05]
Yeah.

**Mark Chen** [14:05]
We have an org that focuses on pre-training, right?

**Aiden** [14:08]
Mm-hmm.

**Mark Chen** [14:08]
Which is, you know, giving models a lot of world knowledge.

**Aiden** [14:11]
Yeah.

**Mark Chen** [14:11]
We focus on RL, like, teaching the models how to reason with that knowledge, how to chain the little insights together.

**Aiden** [14:16]
Yeah.

**Mark Chen** [14:16]
And then finally, um, alignment and post-training, right?

**Aiden** [14:19]
Mm.

**Mark Chen** [14:19]
And, um, we're always looking at both, like, how to scale the mainline in each of these domains and also new bets that fundamentally unlock either, like, different scaling properties or more aggressive scaling properties.

**Aiden** [14:32]
Gotcha.

**Mark Chen** [14:33]
Mm-hmm.

**Aiden** [14:33]
And so even in that, I heard that every one to two months you go through what, like, 300 projects, like, research projects that could be, um, you know, followed through on. Is there a way that you kind of hone that decision-making, I assume, as you, like, decide what to actually double down on and what not to, since I assume there's a lot of talented researchers who provide possible ideas-

**Mark Chen** [14:51]
Yeah

**Aiden** [14:52]
... to pursue?

**Mark Chen** [14:52]
Yeah, so I think really in the spirit of, um, of focus, so one-

**Aiden** [14:56]
Yeah

**Mark Chen** [14:56]
... one narrative you might have heard is, you know, we're, we're really focusing our bets at OpenAI, and-

**Aiden** [15:00]
Yeah. I've heard

**Mark Chen** [15:01]
... um, we're also trying to do a little bit more, uh, directive compute allocation as well. So-

**Aiden** [15:06]
Okay

**Mark Chen** [15:06]
... um, I don't like micromanaging my managers. I think one important thing is to empower them.

**Aiden** [15:12]
Yeah.

**Mark Chen** [15:12]
But to just kind of give compute, big swaths of compute to the big bets you wanna make. And then, um-

**Aiden** [15:17]
That's what you mean by directive, like-

**Mark Chen** [15:20]
Yeah, yeah, yeah

**Aiden** [15:20]
... how to-

**Mark Chen** [15:21]
But, but then also give them kind of flexible pools of compute-

**Aiden** [15:23]
Mm-hmm

**Mark Chen** [15:23]
... which they can, you know, freely allocate to things that, that they believe in or-

**Aiden** [15:27]
Yeah

**Mark Chen** [15:27]
... just kind of, uh, fudge with the, the allocations that-

**Aiden** [15:30]
Mm-hmm

**Mark Chen** [15:31]
... um, that we prescribe. So I think, um, yeah, I, I think it's just tying, let's say, like, a small number of bets, say, three to five bets from each org-

**Aiden** [15:40]
Mm-hmm

**Mark Chen** [15:41]
... um, into the main research roadmap, and then really letting the, the managers, um, and org leads take things from there.

**Aiden** [15:47]
Gotcha. Okay.

**Mark Chen** [15:48]
Mm-hmm.

**Aiden** [15:48]
That makes sense. And I guess for rising researchers-

### Great Researchers

**Mark Chen** [15:50]
Mm-hmm

**Aiden** [15:50]
... so, um, let's say in an interview setting-

**Mark Chen** [15:53]
Yeah

**Aiden** [15:53]
... are there specific tells or ways that you can identify, okay, this person has some, you know, potential of becoming a researcher to impact an org in a specific way? Or is it, like, just looking at the previous research that they've done, and then that is what heavily dictates whether they can actually continue on?

**Mark Chen** [16:10]
Um, it's a hard problem before someone comes to OpenAI.

**Aiden** [16:13]
Yeah.

**Mark Chen** [16:13]
Um, I think that's, that's genuinely true. Um, I, I think, um, for a lot of the best research managers-

**Aiden** [16:20]
Yeah

**Mark Chen** [16:20]
... you know, they work with so many researchers over time-

**Aiden** [16:23]
Mm-hmm

**Mark Chen** [16:23]
... um, where you kind of develop an intuition. Like-

**Aiden** [16:26]
Yeah

**Mark Chen** [16:26]
... the things that they say, the ideas that they bring up.

**Aiden** [16:28]
Yeah.

**Mark Chen** [16:29]
Um-

**Aiden** [16:29]
That's important as well

**Mark Chen** [16:30]
... yep. Are, are those kind of... Do they hit, hit the same mark? Or, like, w- are they the things that you would be thinking about personally too?

**Aiden** [16:37]
Yeah.

**Mark Chen** [16:37]
And so there's this gut check of, like, you know, does their intuition match, um, the same intuition that you have?

**Aiden** [16:43]
Yeah.

**Mark Chen** [16:43]
Um, but it is really hard to tell out, out of the gate. Usually in, you know, let's say six months to a year, it's-

**Aiden** [16:51]
Yeah

**Mark Chen** [16:52]
... pretty clear who, who's, you know, ha- has the strongest trajectory and who's gonna make a lot of impact. Um, so yeah, honestly, um, I, I think it's a hard problem, but just having seen a lot of people, you know, go through research development, um, at, at OpenAI, you develop an intuition for, um, you know, who's more pee-pee in different areas.

**Aiden** [17:14]
Yeah.

**Mark Chen** [17:15]
And one thing to kind of, like, just mention there is not every researcher is the same.

**Aiden** [17:18]
Mm-hmm.

**Mark Chen** [17:19]
I think there's a lot of different types of impact.

**Aiden** [17:20]
Yeah.

**Mark Chen** [17:21]
There are the people who just take an idea, it's very clear, and they'll just implement it before anyone else.

**Aiden** [17:26]
Mm.

**Mark Chen** [17:26]
There are also the people who just come up with the kind of like crazy, almost too crazy, but-

**Aiden** [17:32]
Moonshot type of-

**Mark Chen** [17:33]
Yeah, but somehow not that crazy and-

**Aiden** [17:35]
Yeah

**Mark Chen** [17:35]
... and, and they, they really convince you in a, in a different way of seeing the world or, or, or, or another completely different type of project. So there's a lot of ways to make impact.

**Aiden** [17:43]
Yeah. No, that's helpful.

**Mark Chen** [17:44]
Mm-hmm.

**Aiden** [17:44]
And so I guess elaborating on that-

**Mark Chen** [17:47]
Yeah

**Aiden** [17:47]
... would you say there are similarities that you would see between, let's say, like top engineers-

**Mark Chen** [17:51]
Yep

**Aiden** [17:52]
... um, and top researchers? Like I often hear top engineers-

**Mark Chen** [17:55]
Mm-hmm

**Aiden** [17:55]
... even at like small companies and startups are ones-

**Mark Chen** [17:57]
Mm-hmm

**Aiden** [17:57]
... who can take an iota of an idea-

**Mark Chen** [17:59]
Mm-hmm

**Aiden** [17:59]
... all the way to production.

**Mark Chen** [18:00]
Mm-hmm.

**Aiden** [18:00]
I assume in research it's like coming up with the idea all the way to how it's delivered to the end user, um-

**Mark Chen** [18:06]
Mm-hmm

**Aiden** [18:06]
... through like the product. Or do you think it's more so they're focusing solely on the research and not considering like the end design, how it's used by the customer?

**Mark Chen** [18:14]
Well, I, yeah, I mean, I guess the thing about research is many times the path forward is unclear, and so what-

**Aiden** [18:20]
Gotcha

**Mark Chen** [18:20]
... differentiates researchers is how- Often they're pointed in the right direction.

**Aiden** [18:26]
Mm-hmm.

**Mark Chen** [18:26]
How, like... Like you say, taste, right?

**Aiden** [18:28]
Yeah.

**Mark Chen** [18:28]
I think in engineering there are certain patterns that work. Like, you know, if you want to build a product that looks this way-

**Aiden** [18:33]
Mm-hmm

**Mark Chen** [18:34]
...um, the engineering principles can be pretty similar.

**Aiden** [18:37]
Yeah.

**Mark Chen** [18:37]
Um, but for research, I think the thing that's slightly different is just this ability to, you know, have good research tastes to convince other people-

**Aiden** [18:46]
Yeah

**Mark Chen** [18:46]
...that, um, what you're doing is promising.

**Aiden** [18:48]
Mm-hmm.

**Mark Chen** [18:49]
Um, and then, yeah, um, again, to just kind of integrate it into the core research roadmap.

**Aiden** [18:54]
Gotcha.

**Mark Chen** [18:54]
Yeah.

**Aiden** [18:55]
Great. Okay, it seems like we're done with the vegetables.

**Mark Chen** [18:58]
Awesome. Mm-hmm.

**Aiden** [18:58]
So now we have to multitask.

**Mark Chen** [19:00]
Okay.

**Aiden** [19:00]
So we're gonna pour some water into our pots to get the base of the soup going. So in the top right. Yeah, here. And then just twisting this off.

**Mark Chen** [19:08]
Cool.

**Aiden** [19:08]
So pour some here. Um, you can use some as well. And while we have this simmer and we'll add the veg, we'll cook our prawns here.

**Mark Chen** [19:16]
Okay.

**Aiden** [19:16]
Um, yeah, so let me clean this up real quick.

Looks looking great so far. I feel like saute got some color on the, um-

**Mark Chen** [19:26]
Yeah. Yeah, yeah

**Aiden** [19:26]
...onions and, and mushrooms.

**Mark Chen** [19:28]
Yeah.

**Aiden** [19:28]
So let's turn this on. I guess one aspect or area that seems very interesting are evals.

### Evals Crisis

**Mark Chen** [19:33]
Mm-hmm.

**Aiden** [19:34]
Um, and more specifically, have there been instances where you've seen, like through just vibe checks that it's, it was really good, but on the actual benchmarks, like performs very poorly? Or do you think it's like heavily correlated that, you know, if your Suitebench Pro is, you know, a high number, then it's like your vibe check on it doing coding tasks is also really, really high.

**Mark Chen** [19:53]
No, no. I mean, I think there is this phenomenon, um-

**Aiden** [19:56]
Mm

**Mark Chen** [19:56]
...you know, I, I think internally, I, I'm not sure if this is a externally used word, but yeah, just like benchmarking-

**Aiden** [20:02]
Yeah

**Mark Chen** [20:02]
...you know? Um-

**Aiden** [20:03]
Good high school

**Mark Chen** [20:03]
...I, yeah, I think, I think you can kind of overfit onto certain distributions.

**Aiden** [20:07]
Mm-hmm.

**Mark Chen** [20:08]
Um, and it, it won't be reflective how you, how well you generalize, right? Because-

**Aiden** [20:13]
Gotcha

**Mark Chen** [20:13]
...um, I mean, easy ways to do this are, you know, you take a benchmark and-

**Aiden** [20:16]
Mm-hmm

**Mark Chen** [20:16]
...you just find like very, very, very similar types of instances to the benchmark, and you overtrain on those instances.

**Aiden** [20:22]
Yeah.

**Mark Chen** [20:22]
Um, so I think, um, beyond that, the, the other scary thing in the field is the, the number of canonical gold standard benchmarks is low.

**Aiden** [20:31]
Yeah.

**Mark Chen** [20:32]
And we really are kind of in an evals crisis, right?

**Aiden** [20:36]
Mm-hmm.

**Mark Chen** [20:37]
Where all the really great, uh, evals that we all know, like growing up, like taking the SAT or, um-

**Aiden** [20:43]
Yeah

**Mark Chen** [20:43]
...those, those are all fully saturated.

**Aiden** [20:45]
Yeah.

**Mark Chen** [20:45]
And, um, we really need to find good new ways to benchmark the models. I think one great thing about tools like Codex is they've really enabled the fast iteration of, of evals. Like we're able to just kind of have one person just very quickly put together a very high quality eval.

**Aiden** [21:02]
Mm-hmm.

**Mark Chen** [21:03]
Um, another kind of interesting thing of just being able to deploy your models is you can just see them eval as people are doing things with them. Right?

**Aiden** [21:10]
Yeah.

**Mark Chen** [21:10]
Um, one of the great things is, you know, in math, and coding, and software, like you get a sense for like where, where they fall over, what the task horizon they can do from-

**Aiden** [21:19]
Yeah

**Mark Chen** [21:19]
...from this like general, very broad-based deployment.

**Aiden** [21:21]
Mm-hmm.

**Mark Chen** [21:22]
So.

**Aiden** [21:22]
Yeah. No, that's-

**Mark Chen** [21:23]
Mm-hmm

**Aiden** [21:23]
...that's helpful. Now we'll just add the prawns to-

**Mark Chen** [21:25]
Great

**Aiden** [21:25]
...the oil and get some color on it.

**Mark Chen** [21:27]
Mm-hmm. Cool.

**Aiden** [21:28]
Yeah. And so I guess double clicking onto that-

**Mark Chen** [21:32]
Mm

**Aiden** [21:33]
...um, how do you balance both doing well on these benchmarks-

**Mark Chen** [21:36]
Mm-hmm

**Aiden** [21:37]
...but also not, you know, like benchmark maxing, as you said? 'Cause I assume you want to be most honest and like-

**Mark Chen** [21:43]
Mm-hmm. Mm-hmm

**Aiden** [21:43]
...not kind of cheat the system.

**Mark Chen** [21:45]
Yep.

**Aiden** [21:45]
But if you have like lower scores, let's say than a competitor or from other models-

**Mark Chen** [21:49]
Mm-hmm

**Aiden** [21:49]
...to the consumer, maybe like, wait, your scores aren't that great, so the model just probably is not good. Like, how do you balance both of those dichotomies?

**Mark Chen** [21:56]
Yeah, I mean, I think the thing is you just really have to operate over representative mixtures of evals and-

**Aiden** [22:03]
Yeah

**Mark Chen** [22:03]
...um, always invest in creating new evals.

**Aiden** [22:06]
Mm-hmm.

**Mark Chen** [22:06]
Um, and yeah, really just like there's this philosophy of once an eval's out in the world, um, then it's, it's just already not a good eval.

**Aiden** [22:15]
Yeah. Right. Yeah.

**Mark Chen** [22:15]
Um, and I think one, one thing is, um, also just kind of partnering with external organizations to create evals.

**Aiden** [22:23]
Mm-hmm.

**Mark Chen** [22:23]
So, you know, in, in many of the kind of hard math and science evals, um, we've partnered with external organizations and, um, they've been able to kind of craft gold standards there for us.

**Aiden** [22:34]
Gotcha.

**Mark Chen** [22:34]
So yeah, I think, um, there's a kind of interesting philosophy of separate the teams that are creating the evals from the teams that are optimizing-

**Aiden** [22:41]
Building. Okay

**Mark Chen** [22:42]
...the, the models themselves.

**Aiden** [22:43]
Yeah. The models.

**Mark Chen** [22:43]
Because that way you don't like co-incentivize them, right?

**Aiden** [22:46]
Yeah.

**Mark Chen** [22:46]
Like the, the way the evals team can work is they're trying to build evals that are hard for the model. So there's this inherently adversarial process-

**Aiden** [22:53]
Yeah

**Mark Chen** [22:53]
...where, um, you're, you're not kind of cheating yourself, right?

**Aiden** [22:57]
Yeah.

**Mark Chen** [22:57]
The, the incentives are somewhat, um, aligned in the right way-

**Aiden** [23:01]
Mm-hmm

**Mark Chen** [23:01]
...uh, between the two teams.

**Aiden** [23:02]
Yeah.

**Mark Chen** [23:02]
Yeah.

**Aiden** [23:02]
And do you kind of also contribute and help in the ideation process or even deciding-

**Mark Chen** [23:07]
Mm-hmm

**Aiden** [23:07]
...what evals, you know, you should work with a third party on to develop?

**Mark Chen** [23:11]
Yeah. Yeah. So I mean, I think a, a lot of the work that Jakub and I do also involves just kind of us steering the direction the evals go.

**Aiden** [23:17]
Yeah.

**Mark Chen** [23:17]
I think we'll notice certain gaps, right? Or certain kind of capabilities, um, that we want, and every capability on the flip side is an eval, right?

**Aiden** [23:24]
Mm-hmm.

**Mark Chen** [23:24]
You need some kind of eval that measures if you've elicited that capability well. So yeah, I think, uh, yeah, it's, um, it's takes a lot of steering and just to get everyone on the same page with evals is also a lot of prep work.

**Aiden** [23:38]
Yeah.

**Mark Chen** [23:38]
Yeah.

**Aiden** [23:38]
No, that's, that's fair.

**Mark Chen** [23:39]
Mm-hmm.

**Aiden** [23:39]
I guess on Jakub-

**Mark Chen** [23:41]
Mm-hmm

**Aiden** [23:41]
...you said in a previous interview that he's a very funny guy.

**Mark Chen** [23:44]
Yeah.

**Aiden** [23:44]
Do you have any fun stories that maybe you haven't shared about working with him? 'Cause you also stated that you guys align very well, so your discussions even on research, um, are very efficient and help a lot when like driving towards the frontier.

I guess like on the opposite side of being very funny, are there things that you're, you can share?

**Mark Chen** [24:00]
Oh, you, you asked about like a funny story. Um-

**Aiden** [24:02]
Yeah

**Mark Chen** [24:02]
...well, he told me this joke yesterday, which I thought was very funny. Um, I mean, in many ways we kind of, uh, jointly manage the, the research efforts and-

**Aiden** [24:10]
Yeah

**Mark Chen** [24:10]
...um, you know, apparently some researcher came up to him and was like, you know, um, "It feels like I now just have an army of, you know, really dumb IOI like-" "...gold medalists."

**Aiden** [24:23]
Yeah.

**Mark Chen** [24:23]
And Jakub was like That feels like already the situation I'm in - ... in real life. So, uh, yeah, no, he- he's just, like, brutally sarcastic and funny.

**Aiden** [24:33]
Yeah.

**Mark Chen** [24:34]
Yeah.

**Aiden** [24:34]
No, that's great.

**Mark Chen** [24:34]
Yeah.

**Aiden** [24:34]
It's good to have humor in the workplace-

**Mark Chen** [24:36]
Yeah

**Aiden** [24:36]
... you know, to balance out, especially as you're pushing the frontier on very important work. Um, but that also brings to mind one kind of weird scenario of how models can perform very well on, let's say, the IMO or even-

**Mark Chen** [24:49]
Mm-hmm

**Aiden** [24:49]
... on the IOI, but may struggle with some more mundane tasks that a human can easily do.

**Mark Chen** [24:54]
Mm-hmm.

**Aiden** [24:54]
So I guess, how do you deal with that?

**Mark Chen** [24:55]
Yeah, yeah. I mean, u- ultimately, I think what's intuitive for the models is often not, um, that intuitive for the humans. Like, uh, there's, there's a lot made of this jagged frontier-

**Aiden** [25:05]
Mm-hmm

**Mark Chen** [25:05]
... analogy where, um, there's some things that the models just inherently, you know, based on maybe the data it sees or, um, kind of the, the things that we, we can teach it more easily-

**Aiden** [25:16]
Yeah

**Mark Chen** [25:17]
... it's just good at. Um, I actually think, you know, a lot of, lot of it boils down to also just context, right? Um-

**Aiden** [25:23]
Okay

**Mark Chen** [25:23]
... the models don't have a lot of context that a human has.

**Aiden** [25:25]
Mm-hmm.

**Mark Chen** [25:26]
Um, vision, of course, is something that's more naturally biologically wired for humans. Um, and so, yeah, I think there, there's just certain kind of jagged capabilities that models are better at than humans and vice versa. Um, but I also think, you know, context just, um, being able to take a single task, learn lessons from it, and apply them to future tasks-

**Aiden** [25:47]
Mm-hmm

**Mark Chen** [25:47]
... um, that capability is something that, you know, a lot of people in AI are working towards right now.

**Aiden** [25:52]
Yeah.

**Mark Chen** [25:53]
Um, but it's, yeah, very natural for humans.

**Aiden** [25:55]
Yeah.

**Mark Chen** [25:55]
Mm-hmm.

**Aiden** [25:56]
And on the context point-

### Long Context

**Mark Chen** [25:58]
Mm-hmm

**Aiden** [25:58]
... um, a very low-hanging fruit example that many people say-

**Mark Chen** [26:01]
Yep

**Aiden** [26:01]
... is just to increase the context window to provide more, um, examples-

**Mark Chen** [26:04]
Yeah, yeah, yeah

**Aiden** [26:04]
... so the model can perform.

**Mark Chen** [26:05]
Yeah.

**Aiden** [26:05]
But do you think... I assume there's more complexity on how to actually enable even with a large context window and a lot of context. There could be bloat or even just-

**Mark Chen** [26:13]
Yeah

**Aiden** [26:13]
... a lot of, like, context rot, as people have said.

**Mark Chen** [26:16]
Yeah, yeah.

**Aiden** [26:16]
So how do you go through that process of navigating that?

**Mark Chen** [26:19]
Yes. I think, I think there's kind of the canonical way you would solve for very long horizon learning-

**Aiden** [26:25]
Yeah

**Mark Chen** [26:25]
... which is, you know, you just naively increase your context window, right?

**Aiden** [26:29]
Mm-hmm.

**Mark Chen** [26:29]
And, um, I mean, that makes sense. I think there's a difference between implementing long context and implementing long context well, like you said.

**Aiden** [26:37]
Yeah.

**Mark Chen** [26:37]
Um, and, you know, there's a lot of kind of like needle in the haystack style evals so they can measure that. Um, but I do think beyond that, um, there are also a lot of, in some sense, like engineering and research shortcuts that you could take.

**Aiden** [26:51]
Mm-hmm.

**Mark Chen** [26:51]
Um, so, you know, like, uh, many, many coding products today have features like compaction, right?

**Aiden** [26:57]
Yeah.

**Mark Chen** [26:57]
Where, um, you can compress kind of, uh, either insights or, um, working state.

**Aiden** [27:03]
Mm-hmm, mm-hmm.

**Mark Chen** [27:03]
And, um, stuff like that, you know, it just shortcuts a lot of the, the very brutally difficult and expensive, um, primitives that you have to build with just native long context.

**Aiden** [27:12]
Gotcha.

**Mark Chen** [27:12]
Right.

**Aiden** [27:12]
Great. Okay, now we're gonna do the fun part.

**Mark Chen** [27:15]
Mm-hmm.

**Aiden** [27:15]
So, um, let's lower the heat a little bit.

**Mark Chen** [27:17]
Okay, yeah.

**Aiden** [27:17]
And then add a little bit more oil to the pan.

**Mark Chen** [27:19]
Okay.

**Aiden** [27:20]
Um, and then we'll torch the, the shrimp-

**Mark Chen** [27:22]
Amazing

**Aiden** [27:23]
... pecans.

**Mark Chen** [27:23]
Okay.

**Aiden** [27:24]
To get a little more flavor in there.

**Mark Chen** [27:25]
Yep.

**Aiden** [27:25]
So I'll first do it on, on my pan to show you-

**Mark Chen** [27:27]
Yeah, okay

**Aiden** [27:28]
... generally what it looks like. But-

**Mark Chen** [27:29]
One shot learning.

**Aiden** [27:30]
Yeah.

**Mark Chen** [27:30]
Yeah.

**Aiden** [27:30]
Indeed. Oh, wait. I, oh, I didn't pour any bourbon. Okay, wait. So let's pour, like, a fourth. Okay.

And then pour it in. Heat is off.

**Mark Chen** [27:46]
Ooh.

**Aiden** [27:47]
And then torch it.

**Mark Chen** [27:49]
Awesome.

**Aiden** [27:50]
Great.

**Mark Chen** [27:51]
All right. Okay.

**Aiden** [27:51]
And then once that's out-

**Mark Chen** [27:52]
I, I think I got this.

**Aiden** [27:52]
Okay, yeah. So do, do you wanna do it with your... Yeah.

**Mark Chen** [27:54]
There you go.

**Aiden** [27:55]
So pour it to, like, half of the fourth cup.

**Mark Chen** [27:58]
Okay.

**Aiden** [27:58]
And then once you have that, I can give this to you.

**Mark Chen** [28:01]
Perfect.

**Aiden** [28:01]
Great. So it's off. For the one sec. Oh, we can turn it back on. Perfect. Okay. And then now do you wanna hold on this and just press this button-

**Mark Chen** [28:10]
Yep

**Aiden** [28:10]
... to, uh, fire it up. Yeah.

**Mark Chen** [28:13]
Cool.

**Aiden** [28:14]
Yeah. Yeah, perfect. Great. Okay. Flambéing it. It's a little light, but yeah. Great, and then we could turn on the heat again.

**Mark Chen** [28:22]
Okay.

**Aiden** [28:23]
And then we'll just cook off the alcohol.

**Mark Chen** [28:24]
Okay, great.

**Aiden** [28:25]
But-

**Mark Chen** [28:25]
Great, great

**Aiden** [28:26]
... great. How, how you feeling? You know, we're-

**Mark Chen** [28:28]
Great, great

**Aiden** [28:28]
... basically there, cooking everything.

**Mark Chen** [28:30]
Yeah.

**Aiden** [28:31]
So... Okay. Awesome. Yeah, and I guess in terms of research ideas and-

### Future Research

**Mark Chen** [28:37]
Mm-hmm

**Aiden** [28:37]
... what to work towards, do you think there's still a lot of low-hanging fruit or ideas that can still be improved a lot through just optimizing small parts of already implemented work? Or do you think right now there has to be a lot of research that are completely new bets that people take?

**Mark Chen** [28:51]
Um, yeah, that's a really great question. I feel like there are new bets, but probably not that many.

**Aiden** [28:58]
Yeah.

**Mark Chen** [28:58]
Um, in some sense, like, uh, hopefully you feel like, you know, AGI is coming soon, right?

**Aiden** [29:03]
Yeah.

**Mark Chen** [29:03]
And, um, I think everyone sees that these models are getting really capable. And-

**Aiden** [29:07]
Yeah

**Mark Chen** [29:08]
... I think if you really imagine the implications of that, we're getting closer and closer to a world where the models can come up with more of the innovations on their own.

**Aiden** [29:17]
Yeah.

**Mark Chen** [29:17]
They can kind of do self-sustained research. This is one of the big org goals that we've set for, for our research org.

**Aiden** [29:23]
Mm-hmm.

**Mark Chen** [29:24]
And, and so I think, like, you know, what really matters is are there big bets before that point in time?

**Aiden** [29:29]
Gotcha.

**Mark Chen** [29:30]
And, um, I, I think the window is small, but there are still, like, some fairly significant ideas we're trying out.

**Aiden** [29:35]
Yeah. I mean, there have been some researchers who have stated that-

**Mark Chen** [29:40]
Mm-hmm

**Aiden** [29:40]
... to get to AGI, we still need, let's say, like two or three more breakthroughs. It'd be like continual learning or some other ideas.

**Mark Chen** [29:45]
Mm-hmm.

**Aiden** [29:46]
Um, do you follow that same view perspective, or do you think it's kind of more so, like, not as drastic as coming up with, like, three completely different paradigms for research?

**Mark Chen** [29:55]
Um, yeah, I mean, I, I don't know. I, I don't know if I have that same framing. Like-

**Aiden** [30:00]
Yeah

**Mark Chen** [30:00]
... continual learning is a, a basic primitive that you have to unlock.

**Aiden** [30:03]
Yeah.

**Mark Chen** [30:04]
Um, uh, there's so many different techniques. I, I don't know, um, that... Yeah, I think, you know, we're, we're trying a lot of, you know, permutations of it. I don't know what would consider as a breakthrough versus not-

**Aiden** [30:14]
Yeah

**Mark Chen** [30:14]
... but I think there are clearly many shots on goal, and I-

**Aiden** [30:16]
Yeah

**Mark Chen** [30:17]
... I am pretty sure they'll work.

**Aiden** [30:19]
Great.

**Mark Chen** [30:19]
Yeah.

**Aiden** [30:19]
Okay.

**Mark Chen** [30:19]
Mm-hmm.

**Aiden** [30:20]
Great. Okay, so the shrimp is basically done.

**Mark Chen** [30:22]
Awesome.

**Aiden** [30:22]
Okay, so do you wanna do the flambé thing again to get more color?

**Mark Chen** [30:26]
Let's do it.

**Aiden** [30:27]
'Cause I feel like...

**Mark Chen** [30:27]
Yeah. I'll, I'll-

**Aiden** [30:28]
Yeah, we'll turn on the heat a little bit, and then let me get some more oil.

**Mark Chen** [30:31]
Great.

**Aiden** [30:31]
Yeah, 'cause we wanna get, like, some dark color like Big Lion has-

**Mark Chen** [30:34]
Mm-hmm

**Aiden** [30:35]
... in here. Great.

**Mark Chen** [30:37]
Okay.

**Aiden** [30:37]
And then we can hopefully get another shot.

**Mark Chen** [30:40]
Perfect.

**Aiden** [30:41]
Yes.

**Mark Chen** [30:42]
Do you want, do you want to go first?

**Aiden** [30:43]
Um-

**Mark Chen** [30:44]
You can-

**Aiden** [30:44]
Same, same amount as you.

**Mark Chen** [30:45]
Yeah, same amount. We could hopefully get... 'Cause I think the heat wasn't as... Yeah.

**Aiden** [30:50]
Okay.

**Mark Chen** [30:51]
So you can put it in. Yeah, and then there you go.

Just press a button. You can...

**Aiden** [30:59]
Great. Wow.

**Mark Chen** [30:59]
There we go.

**Aiden** [31:00]
There it is.

**Mark Chen** [31:01]
It's good stress relief.

**Aiden** [31:02]
Yeah. Indeed. And it's like adds some good flavor here. Let me try it over here.

**Mark Chen** [31:07]
Yep.

**Aiden** [31:08]
Yeah, let's see. Great. All right, let's see. Let's go for... Whoo. There it is. All right. But we're in the final stretch.

**Mark Chen** [31:21]
Okay.

**Aiden** [31:21]
We have our shrimp-

**Mark Chen** [31:23]
Mm-hmm

**Aiden** [31:23]
... all cooked.

**Mark Chen** [31:24]
Yep.

**Aiden** [31:24]
With some fire.

**Mark Chen** [31:25]
Mm-hmm.

**Aiden** [31:25]
So now we can kind of, yeah, cook it off a little bit.

**Mark Chen** [31:29]
Yeah.

**Aiden** [31:29]
And then, um, we should add our veg to the water.

**Mark Chen** [31:32]
I'm impressed by your multitasking abilities. You know, I think that's actually one thing, um, we need our models to get better at. Like, it should just be able to do some thread like this and also just have a conversation with you in the background.

**Aiden** [31:43]
Yeah, yeah.

**Mark Chen** [31:44]
Yeah.

**Aiden** [31:44]
No, I, uh, that also reminds me-

**Mark Chen** [31:46]
Mm-hmm

**Aiden** [31:46]
... do you think images and audio and video and even text-

**Mark Chen** [31:49]
Mm-hmm

**Aiden** [31:49]
... like, that should all be one under one model or do you think it'll, like, break through, like, specific specialized, like, audio model or...?

**Mark Chen** [31:55]
Well, yeah, I mean, I think for, for a research lab-

**Aiden** [32:00]
Yeah

**Mark Chen** [32:00]
... I think there are a lot of advantages for it to being under one.

**Aiden** [32:02]
Under one, yeah.

**Mark Chen** [32:03]
So, um, you just have to maintain one infrastructure stack, for instance. Um, I think the cost of, like, maintaining and scaling many infrastructure stacks at once-

**Aiden** [32:11]
Yeah

**Mark Chen** [32:11]
... um, I think that's something that you shouldn't underestimate.

**Aiden** [32:15]
Mm-hmm.

**Mark Chen** [32:16]
So I think there are a lot of benefits to just, like, you know, you do some core research in, in your, like, in your fundamental stack-

**Aiden** [32:23]
Yeah

**Mark Chen** [32:23]
... and that just carries over to whatever modality or whatever thing that you want.

**Aiden** [32:26]
Mm-hmm.

**Mark Chen** [32:26]
So, um, I, I think there's a strong bias for us to keep it in as, as few different arch- uh, architectures as possible.

**Aiden** [32:34]
Gotcha.

**Mark Chen** [32:35]
Yep.

**Aiden** [32:35]
Great. No, that, that makes a lot of sense.

**Mark Chen** [32:37]
Yeah.

**Aiden** [32:37]
I think, like, the architecture as well is something that isn't often considered-

**Mark Chen** [32:40]
Yeah

**Aiden** [32:40]
... but it's very important. But one also term that I've been seeing a lot that you've also kind of mentioned is a vibe researcher. You know, we have vibe coders obviously.

**Mark Chen** [32:49]
Yeah, yeah, yeah.

**Aiden** [32:49]
But I guess on vibe researching, what do you think is, like, the end state? Do you think the main value add of a vibe researcher is just the research taste of coming up with the right idea, or do you think it's more so the execution of going through and following through on the actual research?

**Mark Chen** [33:04]
Um, yeah, so I, I, I think we're actually moving towards this world very quickly, right?

**Aiden** [33:08]
Mm.

**Mark Chen** [33:08]
Um, I think both at OpenAI and at other labs-

**Aiden** [33:11]
Yeah

**Mark Chen** [33:12]
... you're starting to see a lot of the work become mostly orchestration focused, right? Like, um-

**Aiden** [33:17]
Gotcha

**Mark Chen** [33:17]
... the, the researcher's coming up with the ideas, um, and the model's great enough to do the implementation execution by itself.

**Aiden** [33:25]
Mm.

**Mark Chen** [33:25]
Um, so I think, you know, when you, when it comes down to, like, uh, you know, is, is the value of coming up with ideas versus execution, um, yeah, both are still important-

**Aiden** [33:35]
Yeah

**Mark Chen** [33:35]
... but it does feel like there's, there's this market shift-

**Aiden** [33:38]
Mm-hmm

**Mark Chen** [33:38]
... um, towards just kind of being able to come up with a lot of ideas, and then, um, the model can actually do the, the execution-

**Aiden** [33:44]
Mm-hmm

**Mark Chen** [33:45]
... and orchestration for you. So, um, I think it's very much going to be the future of doing research.

**Aiden** [33:50]
Yeah.

**Mark Chen** [33:51]
Um, we also said earlier, you know, like, the models don't quite have the taste yet.

**Aiden** [33:56]
Yeah.

**Mark Chen** [33:56]
And, um, that's why you still need the researchers coming up with the ideas.

**Aiden** [34:01]
Yeah.

**Mark Chen** [34:01]
Um, it, it's gonna be hard to teach the models good taste.

**Aiden** [34:04]
Yeah.

**Mark Chen** [34:04]
Uh, we noticed that. But in terms of actually accelerating the research, there's clear tangible benefits already.

**Aiden** [34:10]
Yeah. Do you think there'll ever be parity in terms of research taste with models at some point?

**Mark Chen** [34:15]
I think so. I mean, when we look at our kind of three-year roadmap, right?

**Aiden** [34:18]
Yeah.

**Mark Chen** [34:18]
Um, the end goal that we want to reach is one where, you know, the, the models are just doing end-to-end research.

**Aiden** [34:24]
Mm-hmm.

**Mark Chen** [34:24]
And, um, I think a part of that problem is just being able to have the model come up with good taste. You point at some, you know, just generic benchmark or something, and it finds the right solutions.

**Aiden** [34:34]
Yeah, yeah. No, that's helpful.

**Mark Chen** [34:36]
Mm-hmm.

### Failed Bets

**Aiden** [34:36]
And in terms of research done by humans-

**Mark Chen** [34:39]
Mm-hmm

**Aiden** [34:39]
... um, at OpenAI, how do you guys go about, I guess, the postmortem process of, let's say, a research bet that didn't turn out well?

**Mark Chen** [34:45]
Mm-hmm, mm-hmm.

**Aiden** [34:46]
'Cause I assume that a lot of it is taking these, you know, vast bets and-

**Mark Chen** [34:49]
Yeah

**Aiden** [34:50]
... some don't turn out well.

**Mark Chen** [34:50]
Well, I would say that is a big part of OpenAI's alpha-

**Aiden** [34:53]
Okay

**Mark Chen** [34:53]
... because I think one thing that differentiates us from other labs-

**Aiden** [34:59]
Yeah

**Mark Chen** [34:59]
... is we take a lot of high-risk bets.

**Aiden** [35:01]
Mm.

**Mark Chen** [35:01]
And I think it's what's allowed us to stay at the frontier-

**Aiden** [35:04]
Gotcha

**Mark Chen** [35:04]
... uh, so consistently over time. Um, but it also means that some of the bets are not gonna pan out.

**Aiden** [35:08]
Mm-hmm.

**Mark Chen** [35:09]
And, um, a hard corollary of that is when a, when a bet doesn't pan out-

**Aiden** [35:15]
Yeah

**Mark Chen** [35:15]
... um, you have to, you know, not delude yourself into thinking that, you know, this is something that will work and, and kind of, uh, disconnect from it.

**Aiden** [35:24]
Yeah.

**Mark Chen** [35:24]
So I think there are certain calls that you have to make, right? Like, uh, kind of look back and be like, "Well, um, this was a promising idea at the time, but actually it's less important than we thought.

You know, it's, uh... there's some other approach that works better."

**Aiden** [35:35]
Yeah.

**Mark Chen** [35:35]
Or, you know, um, there's some other kind of, uh, some- something that we discovered.

**Aiden** [35:40]
Yeah.

**Mark Chen** [35:40]
But I, I think many of that, much of that work is also very fruitful.

**Aiden** [35:44]
Mm-hmm.

**Mark Chen** [35:44]
So what we realized is, like, uh, even sometimes when people fail-

**Aiden** [35:48]
Mm-hmm

**Mark Chen** [35:48]
... at, um, proving out a technique, uh, their write-ups are very import- important.

**Aiden** [35:52]
Mm-hmm.

**Mark Chen** [35:52]
Because, um, they'll kind of a lot... It, it'll often be a natural idea.

**Aiden** [35:55]
Yeah.

**Mark Chen** [35:56]
And you can kind of save a lot of people from going through the same pain.

**Aiden** [35:59]
Yeah. No, that's helpful.

**Mark Chen** [36:00]
Yeah.

**Aiden** [36:01]
So I guess when it comes to this positive view on failure, how do you balance that with, you know, a researcher, let's say-

**Mark Chen** [36:07]
Mm-hmm

**Aiden** [36:07]
... who takes a lot of bets, consecutive bets-

**Mark Chen** [36:09]
Mm-hmm

**Aiden** [36:09]
... and none of them pan out?

**Mark Chen** [36:10]
Mm-hmm.

**Aiden** [36:10]
'Cause I assume at a certain point you'd want a researcher to eventually have contributions that are actually beneficial-

**Mark Chen** [36:16]
Yeah, yeah

**Aiden** [36:16]
... compared to only taking bets that maybe pan out to being not a good taste.

**Mark Chen** [36:21]
You know, just through experience, I-

**Aiden** [36:22]
Yeah

**Mark Chen** [36:22]
... I've definitely seen some people fall into this. Um, but I've also had several cases where, you know, they... it's just, like, bet after bet, it doesn't pan out, and just when you're, like, at the brink of frustration- ...

you have something that's, like, a mega hit.

**Aiden** [36:34]
Mm.

**Mark Chen** [36:35]
And, um, this happened enough, so it, it really just depends on kind of are, are the ideas themselves sound? Uh-

**Aiden** [36:43]
Gotcha

**Mark Chen** [36:43]
... they can be ambitious- But they still have to be sound. And, um, there's a certain kind of person who will, will just take a lot of those ideas, and it's okay-

**Aiden** [36:53]
Mm-hmm

**Mark Chen** [36:53]
... because they're somewhat on the riskier frontier.

**Aiden** [36:55]
Mm-hmm.

**Mark Chen** [36:56]
But they only have to justify it once in a while for it to make sense. Like, maybe like a very trading-

**Aiden** [37:01]
Yeah.

**Mark Chen** [37:01]
... like, kind of lens on the world, but-

**Aiden** [37:03]
That's fair

**Mark Chen** [37:03]
... yeah, it's just on, on expectation, right? Like they-

**Aiden** [37:05]
Right

**Mark Chen** [37:05]
... may need to add value.

**Aiden** [37:06]
Yeah.

**Mark Chen** [37:06]
So.

**Aiden** [37:06]
No, that's, that's great.

**Mark Chen** [37:07]
Yeah.

**Aiden** [37:08]
Okay. So we're basically assembled. Now it's the finishing touches.

**Mark Chen** [37:10]
Mm-hmm.

**Aiden** [37:10]
So you can taste your soup-

**Mark Chen** [37:12]
Oh, great. Okay

**Aiden** [37:12]
... and then we just add soy sauce if it's not salty enough.

**Mark Chen** [37:15]
Great. Great. Great. Okay.

**Aiden** [37:15]
And if it's too salty, we can add some water.

**Mark Chen** [37:17]
Mm-hmm.

**Aiden** [37:18]
Need to lower this down, but let's see our final creation.

How is it?

**Mark Chen** [37:25]
It's pretty good.

**Aiden** [37:26]
Good?

**Mark Chen** [37:26]
Yeah.

**Aiden** [37:26]
I'll tell you. Okay.

**Mark Chen** [37:27]
Mine's really good. Yeah.

**Aiden** [37:28]
Mine needs a little bit of water. Could you pass the water?

**Mark Chen** [37:30]
Yeah, yeah, absolutely.

**Aiden** [37:31]
Great. So how was that? How did that feel?

**Mark Chen** [37:34]
Um-

**Aiden** [37:35]
Great naturalist?

**Mark Chen** [37:35]
... I think this is a student distillation. Um, you, you're clearly better than I am at this point.

**Aiden** [37:39]
No, no, no. I feel like you did a very great job, especially with, even with the shrimp and flambéing. Yeah. Smells great.

Wow, okay. That sounds good. I guess just more generally-

### Overrated/Underrated

**Mark Chen** [37:53]
Yeah

**Aiden** [37:53]
... I'm kind of curious.

**Mark Chen** [37:54]
Mm-hmm.

**Aiden** [37:54]
Are there areas in research or topics that you think are right now overrated and underrated? Like, what would you categorize under?

**Mark Chen** [38:01]
Hmm. Um, well, I think if you still have a pre-training is dead view of the world-

**Aiden** [38:05]
Yeah

**Mark Chen** [38:05]
... um, I think, I think, uh, pre-training is definitely, uh-

**Aiden** [38:10]
Not...

**Mark Chen** [38:11]
Yeah, yeah, yeah.

**Aiden** [38:12]
Not a lot more.

**Mark Chen** [38:12]
Not, not dead. It's, it's, um, it's underrated.

**Aiden** [38:15]
Yeah.

**Mark Chen** [38:15]
Um, yeah. And honestly, um, I think products and kind of thinking about end uses-

**Aiden** [38:25]
Mm-hmm

**Mark Chen** [38:25]
... and, you know, how you tie all the primitives you build in research to, you know, real agentic use cases in the world, that's also underrated.

**Aiden** [38:32]
Mm-hmm.

**Mark Chen** [38:32]
I think, um, you really can't just kind of build everything in a vacuum and not connect things to utility.

**Aiden** [38:38]
Yeah.

**Mark Chen** [38:38]
Mm-hmm.

**Aiden** [38:38]
No, that's great. Great.

**Mark Chen** [38:39]
Yeah.

**Aiden** [38:39]
Awesome. I think we are ready to taste, so do we want to give it a go?

### Taste Test

**Mark Chen** [38:43]
Let's do it.

**Aiden** [38:43]
We can move this.

**Mark Chen** [38:44]
Yeah.

**Aiden** [38:44]
And I think there should be some plating. Yeah. Do you want to take those plates over here?

**Mark Chen** [38:49]
Okay.

**Aiden** [38:50]
Okay. Shrimp looks great, but everything else over here. Great, and then we could just use these pots.

So we have our shrimp.

**Mark Chen** [39:00]
Yep.

**Aiden** [39:01]
Do you want to try it? Cheers.

**Mark Chen** [39:02]
Sure. Cheers.

**Aiden** [39:02]
Okay, let's see. It may be a little sweet.

**Mark Chen** [39:04]
Yeah.

**Aiden** [39:07]
That's a little too sweet 'cause we flambéed twice, but-

**Mark Chen** [39:09]
It's good for me. I have a sweet tooth.

**Aiden** [39:11]
Yeah. It's good. Okay, great. Um, Sean, do you wanna come and taste our tofu soup?

**Sean** [39:16]
Sure.

**Mark Chen** [39:16]
Yeah.

**Sean** [39:16]
That was so good, by the way. You guys can't smell it, but, um, it's-

**Aiden** [39:19]
We'll pretend you're a researcher I've tried to approach. And I'm Zuck, and he's, uh, you know, trying to get you. So do you wanna grab this spoon over here?

**Sean** [39:26]
The soup, soup is really gonna sway the decision here. The quality of soup.

**Aiden** [39:31]
Yeah. Great.

**Sean** [39:32]
All right.

**Mark Chen** [39:33]
Try both.

**Sean** [39:35]
Right. What was the artistic, uh, direction here?

**Mark Chen** [39:38]
Um, artistic direction. Well, it was mimicry. I think it was great art in mimicry.

**Aiden** [39:44]
Yeah. Just letting them cook.

**Mark Chen** [39:46]
Mm.

**Aiden** [39:46]
Good?

**Sean** [39:47]
Wow. Yeah.

**Aiden** [39:48]
Strong?

**Sean** [39:52]
I mean, I think, um, o- one thing I, I... Definitely, it's, like, savory and spice that goes together.

**Mark Chen** [39:57]
Mm-hmm. Yeah.

**Sean** [39:57]
But then also, like, the sort of sea- seafood-ness-

**Mark Chen** [39:59]
Mm-hmm. Yeah

**Sean** [40:00]
... um, kind of really goes into it.

**Aiden** [40:02]
Great.

**Mark Chen** [40:02]
Mm-hmm.

**Aiden** [40:02]
Okay. Mine, mine's very balanced.

**Sean** [40:05]
Am I supposed to pick a winner or what is, what's the, what's the deal?

**Aiden** [40:05]
No, no, no. You're just supposed to try it and then see if it-

**Mark Chen** [40:07]
No, you do pick a winner. Yes, you do.

**Aiden** [40:10]
Okay.

**Sean** [40:10]
This is the evals?

**Mark Chen** [40:11]
Yes.

**Aiden** [40:12]
Yeah. Swift bench.

**Mark Chen** [40:13]
External evals.

**Aiden** [40:15]
Swift bench.

**Sean** [40:15]
Okay. I gotta say, I feel like there's too much water in this. I think, like, the-

**Mark Chen** [40:19]
Oh

**Sean** [40:19]
... the, the- ... the concentration-

**Aiden** [40:20]
I, I think it's also a pot. 'Cause this is a big pot.

**Sean** [40:22]
Yeah. Wait, okay.

**Aiden** [40:24]
I feel like.

**Sean** [40:24]
Um, I, I would s- I have to go for this. Um-

**Mark Chen** [40:26]
Okay. Okay.

**Sean** [40:26]
Just, like, you're our respected guest, but I wanna be objective.

**Mark Chen** [40:30]
Of course. Of course. Yes, yes, yes.

**Sean** [40:31]
For the, for the food.

**Mark Chen** [40:31]
Mm-hmm.

**Sean** [40:31]
Um, and, like, yeah, the density, I think really, uh, uh, flavor really, um, swings it.

**Mark Chen** [40:36]
I- I do ha- half the water. Half the water.

**Sean** [40:39]
Probably.

**Mark Chen** [40:39]
Okay. Make solid sense. Yeah.

**Sean** [40:40]
I mean, I think it's very personal, right? Like-

**Mark Chen** [40:42]
Mm-hmm.

**Aiden** [40:42]
Yeah, I think it's also very personal. Taste, you know, you mentioned-

**Sean** [40:44]
You do a lot of cooking

**Aiden** [40:45]
... research taste and-

**Mark Chen** [40:46]
No. No, no, no. Okay. I know a couple recipes.

**Sean** [40:49]
Yeah.

**Mark Chen** [40:49]
Um, I follow them to the T. I can't... Like, if you tell me, "Oh, cook something slightly different," I have no... I'm completely lost.

**Sean** [40:56]
Right, right, right.

**Mark Chen** [40:56]
Yeah.

**Sean** [40:57]
The, the, well-

**Aiden** [40:57]
Busy cooking and doing research

**Sean** [40:58]
... ChatGPT can tell you.

**Mark Chen** [40:59]
Mm-hmm. Yeah, yeah. Oh, I, I'm not gonna lie, I kind of looked up in ChatGPT a couple of things beforehand, like, um- ... just as prep, but-

**Aiden** [41:07]
No worries. But yeah.

**Mark Chen** [41:08]
Yeah.

**Aiden** [41:08]
It was, it was great having you. I feel like-

**Mark Chen** [41:10]
Mm-hmm

**Aiden** [41:10]
... you're always leading the field with a lot of research taste as well, and it's-

**Mark Chen** [41:13]
Yeah, appreciate you

**Aiden** [41:13]
... great seeing-

**Mark Chen** [41:14]
Yeah

**Aiden** [41:14]
... the work. So hopefully-

**Mark Chen** [41:15]
Absolutely

**Aiden** [41:15]
... this was fun.

**Mark Chen** [41:16]
Yeah, a lot of fun.

**Aiden** [41:16]
Thanks for coming on.

**Mark Chen** [41:16]
Yeah. Thanks so much.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
