LALatent SpaceJun 19, 2025· 1:17:47

Scaling Test Time Compute to Multi-Agent Civilizations — Noam Brown, OpenAI

Noam Brown, OpenAI researcher behind Cicero and reasoning models, joins hosts Alessio and Puix to discuss how test-time compute scaling drives multi-agent AI and the limits of the system 1/2 analogy. He explains that his Diplomacy bot Cicero reached top 10% of humans in 2022, and later he won the 2025 world championship. Brown argues that reasoning models like o3 succeed in unverifiable domains (e.g., Deep Research) and that harnesses and routers will be washed away by scale. He reveals OpenAI's multi-agent team is pursuing a principled approach, not heuristic, and envisions AI civilizations collaborating. He also shares his coding stack (Codex, Windsurf) and predicts test-time compute will hit cost and wall-clock bottlenecks.

  1. 0:00Diplomacy
  2. 5:02Reasoning
  3. 20:37RFT
  4. 29:33Data Efficiency
  5. 33:21Coding
  6. 41:35Multi-Agent
  7. 57:30Test-Time Limits
  8. 1:04:33Rapid Fire

Powered by PodHood

Transcript

Diplomacy0:00

Alessio0:05

Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, partner and CTO of Decibel, and I'm joined by my co-host, Puix, founder of Small AI.

Puix0:13

Hello, hello. And we're here recording on a holiday Monday With Noam Brown from OpenAI. Welcome.

Noam Brown0:19

Thank you.

Puix0:19

So glad to have you finally join us.

Noam Brown0:21

Yeah.

Puix0:21

Uh, a lot, a lot of people have heard you. You've been rather generous of your time on, on the podcast, um, Lex Fridman, and you've done a, you've done a TED Talk recently just talking about the thinking paradigm.

But I think maybe pers- perhaps your most interesting recent achievement is winning the World Diplomacy Championship.

Noam Brown0:39

Yeah.

Puix0:40

In 2022, you, you built, like, sort of Cicero, which was, uh, top 10% of human players. I guess my opening question is: how has your Diplomacy playing changed since working on Cicero and then now personally playing it?

Noam Brown0:53

When you work on these games, you kinda have to understand the game well enough to, like, be able to debug your bot. Because if the bot does something that's, like, really radical and, like, that ty- humans typically wouldn't do, you're not sure if that's, like, a mistake or if that's just-

Puix1:05

Mm

Noam Brown1:05

... uh, like, if it's a bug in the system or it's actually just, like, the bot being brilliant. When we were working on Diplomacy, I kinda, like, did this deep dive, like, trying to understand the game better.

Puix1:13

Yep.

Noam Brown1:14

I played in, in tournaments, I, like, watched a lot of, like, tutorial videos and commentary videos on games, and over that process, I got better. And then also seeing the bot, like, the way it would behave in these games.

Like, sometimes it would do things that humans typically wouldn't do, and that taught me about the game as well. When we released Cicero, we announced it in, like, late 2022. I still found the game, like, really fascinating, and so I, like, kept up with it, I, like, continued to play, and that led to me winning the championship in the world championship in 2025, so just a couple months ago.

Puix1:47

There's always a question of, like, centaur systems where humans and machines work together. Like, was there an equivalent of what happened in Go where you updated your play style? Because-

Noam Brown1:56

Yeah, if you're asking if I used Cicero when I played in the tournament, the answer is the answer is no. But seeing the way the bot played and, like, taking inspiration from that, I think did help me in, in the tournament, yeah.

Puix2:06

Yeah. Uh, do people now ask Turing questions every single time when they're playing Diplomacy?

Noam Brown2:12

Ask- To try to, to try to tell, like-

Puix2:14

Because it's public now

Noam Brown2:15

... uh, if the person they're playing with is a bot or a human?

Puix2:17

Yeah, like, that's the one thing you were worried about when you started.

Noam Brown2:20

It was really interesting when we, when we were working on Cicero because, like, you know, we didn't have the best language models. We were really bottlenecked on the quality of the language models, and sometimes the bot would do...

would say, like, bizarre things. Like, you know, 90, 99% of the time it was fine, but then, like, every once in a while it would say this, like, really bizarre thing. Like, it would just hallucinate about something. Somebody would reference something that they said earlier in a conversation with the bot, and the bot would be like, "I have no idea what you're talking about.

I never said that." And then the person would be like, "Look, you can just scroll up in the chat, and it's like literally right there." And the bot would be like, "No, you're lying."

Puix2:49

Oh, context windows.

Noam Brown2:52

And when it does these kinds of things, like, people just kinda like shrugged it off as like, "Oh, that's just, you know, the person's tired, or they're drunk or whatever, or they're just, like, trolling me." But I think, like, that's because people weren't looking for a bot.

They weren't expecting a bot to be in the games. We were actually really scared because we were afraid that people would, would figure out at one point that, that there's a bot in these games, and then they would just, like, always be on the lookout for it, and they would always be...

And if you're, if you're looking for it, you're able to spot it. That's the thing. So I think now that it's announced and that people know to look for it, I think they would have an easier time spotting it.

Now, that said, the language models have also gotten a lot better since 2022.

Puix3:27

It's adversarial.

Noam Brown3:28

Yeah. So at this point, like, you know, the truth is, you know, GPT-4o and, like, o3, these models are, like, passing the Turing test, so I don't think they can really ask that many Turing-complete questions that would actually make a difference.

Alessio3:39

And Cicero was very small, like 2.7B, right?

Noam Brown3:42

It was a very small, uh, language model, yeah. It was one of the things we realized over the course of the project that, like, oh, yeah, you, you really benefit a lot from just having, like, larger language models.

Alessio3:51

Right. Yep. How do you think about today's perception of AI and a lot of, like, maybe the safety discourse of, like, you know, you're gonna build a bot that is really good at persuading people into, like, helping them win a game.

And I think maybe today labs wanna say they don't work on that type of problem. How do you think about that dichotomy, so to speak, between the two?

Noam Brown4:09

You know, honestly, like, after we released Cicero, a lot of the AI safety community was really happy with the research and, like, the way it worked because it was a very controllable system.

Alessio4:19

Mm.

Noam Brown4:19

Like, we conditioned Cicero on certain concrete actions, and that gave it a lot of steerability to say like, "Okay, well, it's going to pursue a behavior that we can, like, very clearly interpret and, and very clearly define." It's not just like, oh, it's a language model, like, running loose and doing whatever it feels like.

No, it's actually, like, pretty steerable and there's this whole reasoning system that steers the way the language model interacts with the human. Actually, a lot of researchers reached out to, like, reached out to me and said, like, "We think this is, like, potentially a really good way to, uh, achieve safety with these systems."

Puix4:51

I guess the last Diplomacy-related questions that, that we might have is have you updated or tested, like, O-Series models on Diplomacy? And would you, would you expect a lot more difference?

Reasoning5:02

Noam Brown5:02

I have not. I think I said this on, on Twitter at one point that I think this would be a great benchmark. I would love to see all the leading bots play a game of Diplomacy with each other and see-

Puix5:10

Yeah

Noam Brown5:10

... like, who does best. And I think a couple people have, like, taken inspiration from that and are actually, like, building out these benchmarks and, like, evaling the models. My understanding is that they don't do very well right now.

Um, but I think, I think it really is a fascinating benchmark, and I think it would be, um... Yeah, I think it'd be a really cool thing to, to try out.

Puix5:25

Well, we're gonna go a little bit into O-Series now. I think the last time you did a, a lot of publicity you were just launching o1. You did your TED Talk and everything. How has the vibe... How have the vibes changed just in general?

You said you were very excited to learn from domain experts, like, in chemistry, like, how they review the O-Series models. Like, how have you updated since, let's say, end of last year?

Noam Brown5:49

I think the trajectory was pretty clear pretty early on in the development cycle, and I think that everything that's unfolded since then has been pretty on track for what I expected. So I wouldn't say that my perception of where, where things are going or has honestly changed that much.

Um, I think that we're gonna continue to see... I s- I said before that we're gonna see this paradigm continue to progress rapidly, and I think that that's true even today. That we, we saw that with, like, going from o1-preview to o1 to o3, consistent Progress.

And we're going to continue to see that going forward, and I think that we're going to see a broadening of what these models can do as well. You know, like, we're gonna start seeing agentic behavior. Um, we're already starting to see agentic behavior.

Like honestly, for me, o3, I've been using it a ton in my day-to-day life. I just find it so useful, es- especially th- the fact that it can now browse the web and like, you know, do meaningful research on my behalf.

Like, it's kinda like a mini Deep Research that you can just get a response in three minutes. So yeah, I think it's just gonna continue to become more and more useful and, uh, more powerful as time goes on, and pretty quickly.

Guest6:52

Yeah, and, uh, talking about Deep Research, you tweeted about if you need proof that we can s- do this in non-verifiable domains, Deep Research is kinda like a great example. Can you maybe talk about if there's something that people are missing, you know?

I- I feel like I hear that repeated a lot, is like, you know, it's easy to do in coding and math but, like, not in these other domains.

Noam Brown7:08

I frequently get this question, including from pretty established AI researchers, that, okay, we're, we're seeing these, like, reasoning models exceed in math and coding and these, these easily verifiable domains, but are they ever going to succeed in domains where success is less well-defined?

I'm surprised that this is such a common perception because we've released Deep Research and people can try it out. People do use it. It's very popular.

Guest7:32

Mm-hmm.

Noam Brown7:32

And that is very clearly a domain where you don't have an easily verifiable metric for success. It's very-- like what, what is the best research report that you could generate? And yet these models are doing extremely well at this, at this domain.

So I think that's like an existence proof that these models can succeed in tasks that don't have as easily verifiable rewards.

Guest7:52

Is it because there's also not necessarily, like, a wrong answer? Like, there's a spectrum of Deep Research quality, right? You can have, like, a report that, like, looks good but the information is kinda so-and-so, and then you have a great report.

Do you think people have a hard time understanding the difference when they get the result?

Noam Brown8:08

My impression is that people do understand the difference when they get a result and, and I think that they're surprised at how good the Deep Research results are. There's certainly-- it's not, not 100%. It could be better, and we're gonna make it better.

But, uh, I think people can tell the difference between a good report and a bad report and, and certainly and, and a good report and a mediocre report.

Guest8:25

And that's enough to kinda feed the, the loop later to build the product and improve the model performance?

Noam Brown8:30

I mean, I think if you're in a situation where people can't tell the difference between the outputs, then it doesn't-

Guest8:34

Right

Noam Brown8:34

... really matter if you're, like, you know, hill climbing on, on progress. Uh, these models are gonna get better at domains where there is a measure of success. Now, I think this idea that it has to be, like, easily verifiable or something like that, I don't think that's true.

Guest8:47

Mm-hmm.

Noam Brown8:47

I think that you can have, you can have these models do well even in domains where success is a very difficult to define thing. Could sometimes even be subjective.

Guest8:57

People lean on a lot, you've, you've done as well, is the thinking fast and slow analogy for just, uh, thinking models, and, uh, I think it's re- reasonably well diffused now, the idea of, uh, that, that this is kind of the next scaling paradigm.

All analogies are imperfect. What is one way in which thinking fast and slow or system one, system two kinda doesn't transfer to how we actually scale these things?

Noam Brown9:23

One thing that I think is underappreciated is that the models, the pre-trained models need a certain level of capability in order to really benefit from this, like, extra thinking. This is kind of why you, you've seen the reasoning paradigm emerge around the time that it did.

I think it could've happened earlier, but if you try to do the reasoning paradigm on top of GPT-2, I don't think it would've gotten you almost anything.

Guest9:44

Is this emergence?

Noam Brown9:45

Hard, hard to say, uh, if it's emergence necessarily but, like, I haven't done the, um, you know, the measurements to really define that clearly. Um, but I think it's pretty clear. You know, people try chain of thought with GPT, like really small models, and they saw that it just, it didn't really do anything.

Then you go to bigger models and it starts to, to give a lift. I think there's a lot of debate about, like, the extent to which this kind of behavior is emergent, but clearly there is a difference. So it's not like there are these two independent paradigms.

I think that they are related in the sense that you need a certain level of system one capability in your models in order to have system two, to be able to benefit from system two.

Guest10:22

Yeah. I have play-- tried to play amateur neuroscientist before and try to compare it to the evolution of the brain-

Noam Brown10:28

Mm-hmm

Guest10:29

... and how you have to evolve the cortex first before you evolve the other parts of the brain, and perhaps that is what we're doing here.

Noam Brown10:36

Yeah. And you could argue that actually this is not that different from like, I guess, the, um, the system one, system two paradigm because, you know, if you ask, like, a pigeon to think really hard about playing chess, you know, it's not gonna get- ...

that far. It's, you know, it doesn't matter if it, like, thinks for a thousand years, it's, like, not gonna be able to be better at playing chess. So maybe you do still also, also in, like, with animals and humans, that you need a certain level of intellectual ability just in terms of system one in order to benefit from system two as well.

Guest11:00

Yeah. Just a s- ta- side tangent, does this also apply to visual reasoning? So let's say we have, now we have the 4o, like, natively omni model type of thing, then that also makes o3 really good at GeoGuessr.

Does that apply to other modalities too?

Noam Brown11:18

I, I think the evidence is yes. It depends on exactly the kinds of questions that you're asking. Like, there are some questions that I think don't really benefit from system two.

Guest11:26

Like what?

Noam Brown11:26

I think GeoGuessr is certainly one, uh, where you, where you do benefit. I think image recognition, if, if I had to guess, it's like one of those things where you probably benefit less from system two thinking.

Guest11:35

'Cause you know it or you don't.

Noam Brown11:36

Yeah, exactly.

Guest11:37

That's-

Noam Brown11:37

Yeah

Guest11:37

... there's no way.

Noam Brown11:38

Yeah, and the, the thing, the thing I typically point to is just, like, information, like, retrieval. If somebody asks you, like, "When was this person born?" and you don't have access to the web, then you either know it or you don't, and you can sit there and you can think about it for a long time.

Maybe you can make an educated guess and you can say like, "Well, this person was like probably lived around this time, and so this is, like, a rough date." But you're not gonna be able to, like, get the date unless you've actually just, just know it.

Guest12:02

But, like, spatial reasoning, like Tic-Tac-Toe might be better 'cause you have all the information there.

Noam Brown12:06

Yeah, and I think it's true that, like, with Tic-Tac-Toe, we see that, like, GPT-4.5 falls over. You know, it plays decently well. I shouldn't say falls over. It, it does reasonably well. It can draw the board. It can make legal moves.

But it will make mistakes sometimes, and if you really need that system two to enable it to play perfectly. Now, it's possible that if you got to GPT-6 and you just did system one, it would also play perfectly.

You know, I guess we'll, we'll know one day. But I think right now you would need a system two to really, like, do well.

Alessio12:35

What do you think are, like, the things that you need in system one? So obviously general understanding of, like, game rules. Do you also need to understand some sort of, like, meta game of, like, you know, usually this is, like, how you value pieces in different games, even though it's a...

You know, how do you generalize in system one so that then in system two you can kinda get to the gameplay, so to speak?

Noam Brown12:54

I think the more that you have in your s- in, in the system one, like, this is the same thing with humans. You know, like, humans are when they're playing for the first time, uh, a game like chess, they can apply a lot of system two thinking to it.

Alessio13:05

Mm-hmm.

Noam Brown13:05

And if you, if you apply a ton of system two thinking to it, like if, if you just present a really smart person with a completely novel game and you tell them, like, "Okay, you're gonna play this game against, like, an AI or, like, a human that's, like, mastered this game," and you tell them to, like, sit there and y- and think about it for, like, three weeks about how to play this game, my guess is they could actually do pretty well.

But it certainly helps to build up that system one thinking, like build up intuition about, about the game because it will just make you so much, yeah, so much faster.

Alessio13:36

I think the Pokemon example is a good one of, like, the system one kinda has maybe all this information about games, and then once you put in the game, it still needs a lot of harnesses to work, and I'm trying to figure out how much of can we take things from the harness and have them in system one, so that then system two is as harness-free as possible.

But I guess that's like the question about generalizing games and, and AI.

Noam Brown13:57

Yeah. I, I guess I view that as a different question. I think the question about, like, harnesses in my view is that the ideal harness is no harness. You know?

Alessio14:05

Right.

Noam Brown14:06

I think harnesses are a cr- are like a crutch that eventually we're gonna be able to move beyond.

Puix14:11

So only two calls.

Noam Brown14:12

And you could ask... Yeah, you could just ask o3. And actually, you know, it's interesting 'cause, like, when this, uh, playing Pokemon thing kind of like emerged as, as this like, you know, benchmark- ... I was actually, like, pretty opposed to evaling this with our, with our, like, OpenAI models because my feeling is, like, okay, if we're gonna do this eval, let's just do it with o3.

You know? How far does o3 get without any harness? How far does it get playing Pokemon? And the answer is, like, not very far, you know?

Puix14:38

Mm-hmm.

Noam Brown14:38

Um, and that's fine. I think it's fine to have an eval where the models do terribly, and I don't think the answer to that should be like, "Well, let's build a really good harness so that now it can do well on this eval."

I think the answer is like, "Okay, well, let's just, like, improve the capabilities of our models so they can do well at everything, and then they also happen to make progress on this eval."

Alessio14:57

Would you consider things like checking for a valid move a harness, or is this in the, in the model? You know, like chess is like you can either have the model learn in system one what moves are valid and what it can and cannot do versus-

Noam Brown15:10

Search

Alessio15:10

... in system two figuring out whether or not-

Noam Brown15:12

I, I think, I think there's like... A lot of this is design questions. Like, for me, I think you should give the model the ability to check if a move is legal if you want. Like, that, that could be an option in the environment of like, okay, here's a, you know, an action that you can...

Like, a tool call that you can make to see if an action is legal.

Alessio15:27

Mm-hmm.

Noam Brown15:27

If it wants to use that, it, it can. And then there's, like, a design question of like, well, what do you do if the model makes an illegal move? And I think it's totally reasonable to say, like, well, if they make an illegal move, then they lose the game.

Like, I don't know. What, what happens when a human makes an illegal move in a game of chess?

Alessio15:40

Mm-hmm.

Noam Brown15:40

I actually don't know. I don't play chess that much.

Alessio15:41

Yeah, me neither.

Puix15:41

They're just not allowed to?

Noam Brown15:42

Yeah. Like, do you just lose the game?

Puix15:45

I don't know.

Noam Brown15:45

So if that's, if that's the case, then I think it's totally reasonable to say like, "Yeah, we're gonna have an eval where that's also the criteria for-

Puix15:51

Yeah

Noam Brown15:51

... for the AI models."

Puix15:52

Yeah, but I think, like, maybe one way to interpret that in sort of researcher terms is are you allowed to do search? And one of the famous findings from DeepSeek is that MCTS wasn't that useful to them. But I think, like, there are a lot of engineers trying out search and spending a lot of tokens doing that, and maybe it's not worth it.

Noam Brown16:09

Well, I'm making a distinction here between, like, a tool call to check whether a move is legal or illegal is different from actually making that move and then seeing whether it ended up being legal or illegal, right?

Puix16:23

Mm-hmm.

Noam Brown16:23

So if that tool call is available, I think it's totally fine to make that tool call and check whether a move is legal or illegal. I think it's different to have the model say, "Oh, I'm making this move."

Puix16:31

Yeah.

Noam Brown16:32

And then, you know, it gets feedback that like, "Oh, you made an illegal move." And so then it's like, "Oh, just kidding," like, "I'm gonna do something else now." So, so that's, that's the distinction I'm, I'm drawing.

Puix16:41

Some people have tried to classify that second type of thinking things out as test-time compute. You would not classify that as test-time compute?

Noam Brown16:49

There's a lot of reasons why you would not want to rely on that paradigm when you're going to... Like, imagine you have a robot, you know, and your robot, like, takes some action in the world-

Puix16:57

Yeah

Noam Brown16:57

... and it, like, breaks something-

Puix16:58

You can't afford to do it

Noam Brown16:58

... and you're just like... Oh, you can't say like, "Oh, just kidding. I didn't mean to do that. I'm gonna undo that action." Like, the thing is broken. So if you want to simulate what would happen if I move the robot in this way, and then you, you, in your simulation, you saw that this thing broke, and then you decide not to do that action, that's totally fine.

But you can't just, like, undo actions that you've taken in the world.

Puix17:15

There's a couple more things I wanted to cover in this rough area. I actually had an answer on the, on the Thinking, Fast and Slow si-side, which maybe I, I'm curious what you think about. Like, a lot of people are trying to put in effectively model router layers, let's say between, like, the, the fast response model and the, the long thinking model.

Anthropic is explicitly doing that. And I think there is a question about always do you need a smart judge to route or do you need a dumb jarge-- judge to route because it's fast? So when you have a model router, let's say, let's say you're passing requests between system one side and system two side, does the router need to be as smart as the smart model or dumb to be fast?

Noam Brown17:54

I think it's possible for a dumb model to recognize that a problem is really hard, and that it won't be able to solve it and then route it to a, a more capable model.

Puix18:02

But it's also possible for a dumb model b- to be fooled or to be overconfident.

Noam Brown18:06

I don't know. I think there's a real trade-off there. Uh, but I, I will say, like, I think, I think there are a lot of things that people are building right now that will eventually be washed away by scale.

So I think harnesses are a good example, where I think eventually the models are going to be... And I think this actually happened with the reasoning models. Like, before the reasoning models emerged, there was, like, all of this work that went into engineering these, like, agentic systems that, like, made a lot of calls to GPT-4o or, like, the- these non-reasoning models to get reasoning behavior.

And then it turns out like, oh, we just like created reasoning models, and they-- you don't need this like complex behavior. In fact, in, in many ways it makes it worse. Like, you just give the reasoning model the, the same question without any sort of scaffolding, and it just does it.

Now, the-- you can still... And, and so people are building scaffolding on top of the reasoning models right now, but I think in many ways, like, those scaffolds will also just be replaced by the reasoning models and models in general becoming more capable.

And similarly, I think things like model, uh, like these routers. You know, we've said pretty openly that we want to move to a world where there is a single unified model, and in that world, you shouldn't need a router on top of the model.

So I think that the router issue will eventually be solved also by scale.

Puix19:19

Like you're building the router into the model kind of weights itself.

Noam Brown19:23

I don't think there will be a, a, a benefit for... Like I don't-- I shouldn't say it, 'cause it's-- I could be wrong about this. Like, you know, and certainly maybe there's, um, you know, reasons to route to different model providers or whatever.

But I think that routers are going to, um, eventually go away. And I can understand why it's worth doing it in the short term because, like, the fact is, it is beneficial right now, and if you're building a product and, and you're getting a lift from it, then it's, it's worth doing right now.

One of the tricky things I'd, I'd imagine that a lot of developers are facing is that, like, you kind of have to plan for where these models are going to be in six months, in twelve months, and that's like very hard to do because things are progressing very quickly.

You know, you don't want to spend six months building something and then just have it be totally washed away by scale. But I, I think I would encourage developers, like when they, when they're, you know, building these kinds of things like scaffolds and, and routers, keep in mind that the field is evolving very rapidly.

You know, things are gonna change in three months, uh, let alone six months, and that might require radically changing these things around or, or tossing them out completely. So don't spend six months building something that might get tossed out in six months.

Puix20:30

It's so hard though. Everyone says this- -and then, like, no one has concrete suggestions on how.

Alessio20:37

What about reinforcement fine-tuning? Is it something that... Obviously you just released it a month ago at OpenAI. Is this something people should spend time on right now or maybe wait until the next jump up in scale?

RFT20:37

Noam Brown20:47

I, I think re- I think reinforcement fine-tuning is, is pretty cool, and I, I think it's like worth looking into because it's really about specializing the models for the data that you have, and I think that, um, something that's like worth, worth looking into for, for developers.

Like we're not, we're not suddenly gonna like have that data baked into the raw model a lot of times.

Alessio21:09

Yeah.

Noam Brown21:09

So I, I think that's kinda like a separate question.

Alessio21:11

Yeah. So creating the environment and the reward model is the best thing people can do right now. I think the question that people have is like, should I rush to fine-tune the, the model using RFT or should I build the harness to then RFT the models-

Noam Brown21:25

Well-

Alessio21:25

-as they get better?

Noam Brown21:26

I think the difference is that like for reinforcement fine-tuning, you're collecting data that's gonna be useful as the models improve as well.

Alessio21:35

Mm-hmm.

Noam Brown21:35

So if we come out with like future models that are even more capable, you could still fine-tune them on your data. That's, I think, actually a good example where you're building something that's going to complement the model scaling and becoming more capable rather than necessarily getting washed away by the scale.

Alessio21:51

Mm-hmm. Yep.

Puix21:51

One last question on Ilya. You mentioned on, I think, the Sara and Elad podcast where you had this conversation with Ilya a few years ago about more RL and reasoning and language models. Just any speculation or thoughts on why his attempt when he tried it, it didn't work or the, the timing wasn't right and why the time is right now?

Noam Brown22:15

I, I don't think I would, I would frame it that way, that like his, his attempt didn't work. In, in, in many ways it did. Um, so Ilya, for me, I saw that in all of these domains that I'd worked on, in poker and Hanabi and Diplomacy, having the models think before acting made a huge difference in performance, like orders of magnitude difference.

Puix22:36

Like ten thousand times or something.

Noam Brown22:37

Yeah, like, you know, a thousand to a hundred thousand times. Like it's the, the equivalent of a model that's like a thousand to a hundred thousand times bigger. And in language models, you weren't really seeing that, that the model, the models would just respond instantly.

Some people in the field, in, in the LLM field were like convinced that like, "Okay, we just keep scaling pre-training, we're gonna get to super intelligence." And I was kinda skeptical of that perspective. In late twenty twenty-one I was having a meal with Ilya.

He asked me what my AGI timelines are, a very standard SF question. And I told him like, "Look, I think it's actually quite far away because we're gonna need to figure out this reasoning paradigm in a very general way."

And with things like LLMs, LLMs are very general, but they don't have a reasoning paradigm that's very general. And until they do, they're gonna be limited in what they can do. You know, like we're gonna scale it, sure.

We're gonna scale these things up by a few more orders of magnitude. They're gonna become more capable, but we're not gonna see super intelligence from just that. And like, yes, if we had a quadrillion dollars to train these models, then maybe we would, but like you're gonna hit the limits of what's economically feasible before you get to super intelligence unless you have a reasoning paradigm.

And I was convinced incorrectly that the reasoning paradigm would take a long time to figure out because it's like this big unanswered research question. And, you know, Ilya agreed with me. And he said like, "Yeah, you know, I think we need this like additional paradigm."

But his take was that like maybe it's not that hard. I, I didn't know it at the time, but like he and others at OpenAI had also been thinking about this. They'd also been thinking about RL. They'd been working on it, and I think they had some success.

But like, you know, with most research, like it does-- you have to iterate on things. You have to try out different ideas. You have to, yeah, try different things. And then also as the models become more capable, as they become faster, it becomes easier to iterate on experiments.

And I think that the work that they did, even though it didn't like result in a reasoning paradigm, it all builds on top of previous work, right? So they built a lot of things that over, over time led to this reasoning paradigm.

Puix24:32

For listeners, Noam can talk about this, but the rumor is that that thing was code named GPT Zero if you wanna search for that, that line of work. I think there was a time where like basically RL kind of went through a dark age when everyone like- Went all in on it, and then nothing happens, and they gave up.

And like, now it's like sort of the golden age again. So that's what I'm like trying to identify. Like, why? What is it? And it could just be that we have smarter base models and better data.

Noam Brown24:57

I don't think it's just that we have smarter- ... base models. I, I think it's that, yeah. So I-- We did end up getting a big success with, with reasoning and I-- but I think it was, in many ways, a gradual thing.

Uh, a l- To some extent it was gradual. You know, like there were signs, there were signs of life, and then we like, you know, iterated and tried out some more things. We got like better signs of life.

I think it was around like No-November twenty twenty-three or October twenty twenty-three when I think I was convinced that we had like very conclusive signs of life that like, "Oh, this is, this is going to be a-- This is the paradigm, and it's gonna be a big deal."

That was in many ways a g-a gradual thing. I think what OpenAI did well is like when we got those signs of life, they recognized it for what it was and invested heavily in, in scaling it up. And I think that's, that's ultimately what, what led to reasoning models arriving when they did.

Alessio25:46

Was there any disagreement internally, especially because like, you know, OpenAI kind of pioneer pre-training scaling, you know, and kinda like compute is all you need, and then you're kinda saying maybe that's not how we get there. Was it clear to everybody that like, okay, this is gonna work or was it controversial?

Noam Brown26:02

There's always different opinions about this stuff. I think there were some people that felt that pre-training was all we need, that we scaled it up to infinity and we're there. I think a lot of the leadership actually at OpenAI recognized that there was another paradigm that was needed, and that was why they were investing all of this like research effort into this like RL, um, stuff.

And I think that's also to the credit of OpenAI that like, okay, yes, they figured out the pre-training paradigm, and they were very focused on scaling it up. In fact, the vast majority of resources were focused on-

Alessio26:24

Mm-hmm

Noam Brown26:24

... scaling it up. But they also recognized the value that, that something else was gonna be needed, and it was worth researching, putting researcher effort into, into other directions to figure out what that extra paradigm was going to be.

There was a lot of debate about, first of all, like what is that extra paradigm?

Alessio26:40

Mm.

Noam Brown26:40

So I think a lot of the researchers looked at reasoning and, and RL was not really about scaling test-time compute. It was more about data efficiency because, you know, the feeling was that, well, we have tons and tons of compute, but we actually are more limited by data, so there's, there's the data wall.

And we're gonna hit that before we hit limits on, on the compute. So how do we make these algorithms-

Alessio27:02

Mm

Noam Brown27:02

... more data efficient? They are more data efficient, but I think that also like, um, they are also just like the equivalent of scaling up compute also by a ton. That was interesting. There was like a lot of debate around like, okay, well, what exactly are we doing here?

And then I think also, even when we got the signs of life, I think there was a lot of debate about the significance of it. That like, okay, how much should we invest in scaling up this paradigm? I think especially when you're, when you're in a small company like, you know, OpenAI, like in twenty twenty-three was not as big as it is today, and compute was more constrained than it is today.

Alessio27:31

Mm-hmm.

Noam Brown27:32

And if you're investing resources in, in a direction, that's coming at the expense of something else. And so if you look at these signs of life on reasoning and you're saying like, "Okay, well, this looks promising. We're gonna scale this up by a ton and invest a lot more resources into it," where are those resources coming from?

You have to make that tough call about where to, where to draw the resources from, and that is a very controversial, very difficult call to make, um, that makes some people unhappy. And I think there was debate about whether we're focusing too much on this paradigm, whether it's really a big deal, whether we would see it generalize and do various things.

And I remember it was interesting that I, I talked to somebody who left OpenAI after we had discovered the reasoning paradigm but before we announced o1.

Alessio28:15

Mm-hmm.

Noam Brown28:16

And they ended up going to a competing lab. I saw them afterwards, after we announced, um, o1, and they told me that like at the time, they really didn't think this like reasoning thing, like this, these O-Series, the Strawberry models, were like that, that big of a deal.

It was like they thought we were making a bigger deal of it than it really deserved to be.

Alessio28:32

Mm-hmm.

Noam Brown28:33

And then when we announced o1 and they saw the reaction of their coworkers at this competing lab about how everybody was like, "Oh, crap," like, "This is a big deal," and they like pivoted the whole research agenda-

Alessio28:45

Oh my God.

Noam Brown28:46

... um, to, to focus on this, that then they realized like, "Oh, actually, like this maybe is a big deal." You know, a lot of this seems obvious in retrospect, but at the time it's actually not so obvious a-and can be quite difficult to recognize something for what it is.

Alessio29:01

I mean, OpenAI has like a great history of just making the right bet. I feel like GPT models are kinda similar, right? Where like it started with games and RL, and then it's like maybe we can just scale these language models instead and, um, I'm just impressed by the leadership and obviously the, the research team that keeps coming out with these insights.

Noam Brown29:20

Looking back on it today, it, it might seem obvious that like, oh, of course, like these models get better with scale, so you should just scale them up a ton and it'll get better. But it, it really is...

The best research is obvious in retrospect, and at the time it's, it's not as obvious as it might seem today.

Data Efficiency29:33

Puix29:36

Follow questions on data efficiency. Uh, this is, this is a pet topic of mine. It seems that our current methods of learning are so inefficient still, right? Like compared to the existence proof of humans, we take five samples and we learn, learn something.

Machines, kinda two hundred maybe, you know, per, per like whatever data point you, you might need. Anyone doing anything interesting in data efficiency or do you think like there's just a fundamental inefficiency that machine learning has that will just always be there compared to humans?

Noam Brown30:06

I think it's a good point that if you look at the amount of data these models were trained on and you compare it to like the amount of data that a human observes to get the same performance, I guess pre-training, it's a little hard to make an apples to apples comparison because like I don't know how many, how many tokens does a baby actually absorb when they're developing.

But I think it's a fair statement to say that these models are, are less data efficient than humans, and I think that that's an unsolved research question and probably one of the most important unsolved research questions.

Puix30:33

Maybe more important than algorithmic improvements 'cause you can just, you can, y-we, we can increase the supply of data out of the existing set of the world and humans.

Noam Brown30:43

I guess there's, okay, so a couple thoughts on that. Like one is that the answer might be an algorithmic improvement, like maybe, maybe algorithmic improvements do lead to greater data efficiency. And the second thing is that like- It's not like humans learn from just reading the internet.

So, um, I think it's certainly easiest to learn from just, like, data that's on the internet, um, but I don't think that's, like, the limit of what data you could collect.

Puix31:06

The last follow-up before we change t-topics to coding, any other just anecdotes or insights from Ilya, just in general? 'Cause, like, you've worked with him, so there's n-not that many people that we can talk to that have worked with him .

Noam Brown31:18

I think I've just been very, very impressed with his vision. That I think, like, especially when I joined and, and I saw, you know, the internal documents at, at OpenAI of, like, what he had been thinking about back in like 2021, 2022, even earlier.

Uh, I was very impressed that he had a clear vision of, like, where this was all going and what was needed.

Puix31:37

Some of his emails from 2016, '17 when they were founding OpenAI was published, and even then he was talking about how he thinks, like, one big experiment is much more valuable than 100 small ones. That was, like, a core insight that differentiated them from Brain, for example.

It just seems very insightful that he just sees things much more clearly than others, and I, I would-- I just wonder what his production function is. Like, how do you make a human like that, and how do you improve your own thinking to better model it?

Noam Brown32:05

I, I mean, I think it is true that, I mean, one of OpenAI's big success was betting on the scaling paradigm. It is just kind of odd because, you know, they were not the biggest lab. You know, it was, like, difficult for them to scale.

Back then, it was much more common to do, like, a lot of small experiments, more academic style. People were trying to figure out these, um, various, like, algorithmic improvements, and OpenAI bet pretty early on, like, large scale.

Puix32:28

We had David Luan on, who I think was VP Eng at the time of GPT-1 and 2, and he talked about how the differences between Brain and OpenAI was basically the cause of the-- Google's inability to come out with a scaled model.

Like, just structurally, the-- everyone had allocated compute, and you had to pool resources together to make bets, and you just couldn't.

Noam Brown32:48

I think that's true, that OpenAI was structured differently, and I think that really helped them. Like, OpenAI functions a lot like a startup, and other places tended to function more like universities or, or, you know, research labs as they traditionally existed.

The way that OpenAI operates, more like as a startup with this mission of building AGI and, and superintelligence, that helped them organize, collaborate, pool resources together, make hard choices about, like, how to allocate resources. And I think a lot of the other labs, like, have now been trying to adopt paradigms more like that, like, setups more like that.

Coding33:21

Alessio33:23

Let's talk about maybe the killer use case, at least in my mind, of these models, which is coding.

Noam Brown33:27

Mm.

Alessio33:28

You released Codex recently, but I would love to talk through the Noam Brown coding stack. What models you use, how you interact with them.

Noam Brown33:35

Cursor-

Alessio33:35

Uh

Noam Brown33:35

... Windsurf. Uh, lately I've been using, uh, Windsurf and Codex. Like, actually a lot of Codex.

Alessio33:40

Yeah.

Noam Brown33:40

I, I've been having a lot of fun. Like, you just give it a task, and it just goes off and does it and comes back five minutes later with, like, a, you know, pull request.

Puix33:46

And is it core research task or, like, side stuff that you don't super care about?

Noam Brown33:50

I wouldn't say it's, like, side stuff. I would say basically anything that I would normally try to code up, I try to do it with Codex first. For, for-

Puix34:03

Well, for you it's free, but yeah. For everybody, it's free right now.

Noam Brown34:05

And I think that's partly because it's the, it's the most effective way for me to do it, and also it's for-- good for me to get experience working with this technology and then also, like, seeing the shortcomings of it.

It just helps me, like, better understand, like, okay, this is the, the limits of these models and, like, what we need to push on next.

Puix34:21

Have you felt the AGI?

Noam Brown34:22

I've felt the AGI multiple times, yes.

Puix34:25

Like, like, um, how should people push Codex in ways that you've, you've done and, you know, I think you, you see it before others 'cause obviously you, you work closer to it.

Noam Brown34:34

I think anybody can use Codex and feel the AGI. It's kind of funny how, like, you feel the AGI and then you get used to it very quickly. You know, so, so it's really, like-

Puix34:45

Dissatisfied with, like, where it's lacking.

Noam Brown34:47

Yeah, no, it, it, you know, it's, it's magical one day. I was actually looking back at the old, uh, Sora videos when they were announced.

Puix34:53

Yeah.

Noam Brown34:53

'Cause, like, you remember when Sora came out, it was just like-

Puix34:55

The biggest news ever

Noam Brown34:56

... it was just, it was just magical. You look at that and you're like, "Oh, it's like it's really here. Like, this is AGI." But if you look at it now and it's kinda like, oh, you know, the people, like, don't move, like, very organically and it's like there's, like, a lack of consistency in some ways.

And you see all these flaws in it now, uh, that you just didn't really notice when it was first came out. And yeah, you get used to this technology very quickly and... But I think what's cool about it is that because it's developing so quickly, you get those feel the AGI moments, like, every few months.

So something else is gonna come out and just, like, it's magical to you and, uh, and then you get used to it very quickly, yeah.

Puix35:29

What are your, uh, Windsurf pro tips now that you've immersed in it?

Noam Brown35:34

I think one thing I was surprised by is how few people... I mean, maybe your audience is gonna be more comfortable with reasoning models and, like, use reasoning models more, but I'm surprised at how many people don't even know that o3 exists.

Like, I've been using it day to day. It's basically replaced Google Search for me. Like, I just use it all the time. Like, and also for things like coding, like I, I tend to just use the reasoning models.

My suggestion is, like, if people are not-- have not tried the reasoning models l-yet, 'cause, like, honestly, like, we do-- Like, people love them, people that use it love them. Obviously, a lot more people use GPT-4o, um, and just, like, the default on what-- on ChatGPT and that kind of stuff.

I think it's worth trying the reasoning models. Like, I think people would be surprised at what they can do.

Puix36:16

I use Windsurf daily, and they still haven't actually enabled it as, like, a default in Windsurf. Like, I always have to dig up, like, type in o3 and then it f- and then it's like, oh yeah, that, that exists.

Noam Brown36:26

Mm-hmm.

Puix36:27

It's, it's, uh, it's weird. I would say, like, my struggle with it has been that it's takes so long to reason that I actually s- break out of flow.

Noam Brown36:34

I think that is true, yes.

Puix36:35

And that's-

Noam Brown36:36

And I think this, this is one of the advantages of Codex, that, like, okay, you can give it a task that's kinda self-contained and, like, it can go off and do its thing and come back ten minutes later.

And I can see that if you're doing-- if you're using this thing as, like, more like a, like a pair programmer kinda thing, then yeah, you wanna use GPT 4.1 or something like that.

Alessio36:52

What do you think are the most broken part of the development cycle with AI?

Noam Brown36:56

Ooh.

Alessio36:56

Like, in my mind it's like, um, pull request review. Like, for me, like, I use Codex all the time, and then I got all these pull requests, and-

Puix37:02

And you need to review

Alessio37:03

... it's kinda hard to, like, go through all of them. What other thing would you like people to build to make this even more scalable?

Noam Brown37:10

I think it's really on us to build a lot more stuff. These models are very limited in w- in, in some ways. I think I find it frustrating that, you know, you ask them to do something, and then they spend 10 minutes-

Guest37:20

Mm-hmm

Noam Brown37:20

... doing it, and then you ask them to do something pretty similar, and then they go spend 10 minutes doing it.

Guest37:26

Right.

Noam Brown37:26

And like you know, it's... I, I think I describe them as like they're geniuses, but it's their first day on the job, you know? And that's, like, kind of annoying. Like, even the, the smartest person on Earth when th- when it's their first day on the job, you know, they're not gonna be, like, as useful as you would like them to be.

So I think being able to get more experience and, like, act like somebody that's actually been on the job for, like, six months instead of o- one day, I think would make them a lot more useful. But that's really on us to build, to build that capability.

Guest37:52

Do you think a lot of it is, like, GPU constrained for you? Like, if I think about Codex, why is it asking me to set up the environment myself when, like, the model... If I ask o3 to, like, create an environment setup script for a repo-

Noam Brown38:03

Mm

Guest38:03

... I'm sure it'll be able to do it. But today in the product, I have to do it. So I'm wondering, in your mind, could these be a lot more if we just, again, put more test time compute on them, or do you think there's, like, a fundamental model capability limitation today that we still need a lot of, like, human harnesses around it?

Noam Brown38:20

I think that we're in an awkward state right now where, like, progress is very fast, and there's things that are like, clearly we could do this, and the models would be better. We're gonna get to it. It's, um...

You're just limited by how many hours there are in the day.

Guest38:31

Right.

Noam Brown38:31

You know? So progress can only proceed so quickly. We're trying to get to everything as fast as we can, and, and I think that, uh, o3 is not where the technology will be in six months.

Guest38:42

I like that question overall in, like, there's a software development life cycle, not just generation of the code. Like, from issue to PR, basically is, is, like, the, the typical commentary of that. And then there's the Windsurf side, which is inside your ID.

Like, what else, right? Pull request review is, like, something that people don't really... There are startups that are built around it. It's not something that Codex does, and it could. And so, like, then there's, like, what else is there, you know, that is sort of rate limiting the amount of software you could be iterating on.

It's an open question. I don't, I don't, I don't know if there's an answer. Anything else on, on ASE in general? Like, where do you think this goes just in form factors, or what will we be looking at this time next year in terms of how things are...

how... what models are able to do that they're not able to today?

Noam Brown39:30

I don't think it's gonna be limited to ASE, you know? I think-- I don't think it's gonna be limited to software engineering. I think it's gonna be able to do a lot of remote work kinda tasks.

Guest39:38

Yeah, like freelancer-type Upwork.

Noam Brown39:40

Yeah, or just, like, even things that are not necessarily software engineering.

Guest39:43

Okay.

Noam Brown39:44

You know? So the way that I think about it is like anybody that's doing a remote work kinda job, I think it's valuable to become familiar with the technology and, like, kinda get a sense of, like, what it can do, what it can't do, what it's good at, what it's not good at because I think the, the breadth of things that it's gonna be able to do is gonna expand over time as well.

Guest40:00

I feel like virtual assistants might be the next thing after ASE then because they're the most easily... Like, y- you know virtual asssist-- like, hire someone in the Philippines, someone s- uh, who, who just look through your email and all that.

Because that is entirely... You can intercept all the inputs and all the outputs and train on that. And maybe OpenAI just buys a virtual assistant company.

Noam Brown40:20

Yeah, I think what I'm looking forward to is that, um, for things like virtual assistants, the, the models, like, if they're aligned well, they could end up being, like, really preferable for the, for that kind of work, you know?

If... There's always this, like, principal-agent problem, where if you delegate a task to somebody, then, like, are they really aligned with, like, doing it as-

Guest40:40

Mm-hmm

Noam Brown40:40

... you would want it to be done and, um-

Guest40:42

Or just as cheaply, as quickly as they can.

Noam Brown40:44

Yeah.

Guest40:44

Yeah.

Noam Brown40:44

Yeah. And so if you have an AI model that's, like, actually really aligned to you and your preferences, then that could end up doing a way better job than a human could.

Guest40:53

Yeah.

Noam Brown40:53

Well, not, not it's doing a better job than a human could, but, like, it's doing a better job than a human would.

Guest40:57

That word, alignment, by the way, I think there's, like, an interesting overriding or, uh, homomorphism between safety alignment and instruction-following alignment, and I wonder where they diverge.

Noam Brown41:10

Okay, so I think where it diverges is, like, what do you want to align the models to? Like, that-

Guest41:14

Yeah

Noam Brown41:14

... that's, I think, a difficult question, you know? Like, you could say, like, you want it to align it to the user. Okay, well, what happens if the user wants to build a novel virus that's gonna wipe out half of humanity?

You know?

Guest41:23

That safety alignment.

Noam Brown41:23

Yeah.

Guest41:24

Yeah.

Noam Brown41:24

So there's a question of like... I think alignment... I think they're related, you know? And I think the, the, the big question is, like, what are you aligning towards?

Guest41:31

Yeah. There's, like, humanity goals, and then there's your personal goals and everything in between.

Noam Brown41:35

Mm-hmm.

Multi-Agent41:35

Guest41:35

So that's kinda, uh, I guess the individual agent. And you announced the... you're leading the multi-agent team at OpenAI. I haven't really seen many announcements. Maybe I missed them on what you've been working on, but what can you share about interesting research directions or, um, anything from the-

Noam Brown41:51

Yeah, there's... uh, hasn't really been announcements on this. I think we're working on cool stuff, and I think we'll get to announce some cool stuff, uh, at some point. I think the team i- in many ways is actually a misnomer 'cause we're working on more, more than just multi-agent.

Multi-agent is one of the things we're working on. Um, some other things we're working on is just, like, being able to scale up test time compute by, by a ton. So how... you know, we get these models thinking for 15 minutes now.

How do we get them to think for hours? How do we get them to think for days, even longer, and be able to solve incredibly difficult problems? So that's one direction that we're pursuing. Uh, multi-agent is another direction, and here I think there's a few different motivations.

Uh, we're interested in, like, both the collaborative and the competitive aspect of multi-agent. I think the way that I describe it is people often say in AI circles that humans occupy this very narrow band of intelligence, and AIs are just gonna, like, quickly catch up and then surpass, like, this band of intelligence.

And I actually don't think that the band of intell- of human intelligence is that narrow. I think it's actually quite broad. Because if you compare anatomically identical humans from, you know, caveman times, they didn't get that far in terms of, like-

Guest43:00

Mm-hmm

Noam Brown43:00

... you know, we, we would consider intelligence today, right? Like, they're not putting a man on the moon. You know, they're not, like, building semiconductors or nuclear reactors or anything like that. And, and we have those today, even though we as humans are not anatomically different.

And so what's the difference? Well, I think the difference is that you have thousands of years- A lot of humans, billions of humans cooperating and competing with each other, building up civilization over time. The technology that we're seeing is the product of this civilization, and I think similarly, the AIs that we have today are kinda like the cavemen of AI.

And, and I think that if you're able to, um, have them cooperate and compete with billions of AIs over a long period of time and build up a civilization, essentially, the things that they would be able to produce and answer will be far beyond what is possible today with g- with the AIs that we have today.

Alessio43:57

Do you see that being similar to maybe like Jim Fan's Voyager skill library idea, resaving these things, or is it just the models then being retrained on this new knowledge? Because the humans then have it, a lot of it, in the brain as they grow.

Noam Brown44:11

I think I'm gonna be evasive here and say that, like we're-

Alessio44:13

Okay. Yeah

Noam Brown44:14

... we're, we're not gonna... Yeah. We're not gonna-

Alessio44:15

Ooh.

Noam Brown44:16

We're- we... Until we have something to announce, which I think that we-

Alessio44:19

Yeah, yeah, yeah

Noam Brown44:19

... I think that we will in the not too distant future, I think I'm going to, to, uh, be a bit vague about like exactly-

Alessio44:25

Yeah

Noam Brown44:25

... what we're doing. Uh, but I will say that the way that we are approaching multi-agent in the details and the, the way we're actually going about it, is, I think, very different from how it's been done historically and how it's being done today by, by other places.

Um, I've been in the multi-agent field for a long time. I've kinda felt like the multi-agent field has been a bit misguided in some ways in the, the things that, in the, the approaches that the field has taken and, like, the way that it's been approached.

And so I think we're trying to take a very principled approach to multi-agent.

Alessio44:53

Sorry, I gotta add. Like, so you, you can't talk about what you're doing, but you can say what's misguided. What's misguided?

Noam Brown44:58

I think that a lot of the approaches that have been taken have been very heuristic-

Alessio45:02

Okay

Noam Brown45:02

... and haven't really been following, like, the bitter lesson approach to scaling and research.

Alessio45:08

Okay. I think maybe this might be a good, a good spot. So obviously, you've done a lot of amazing work in, in poker, and I think as the reasoning model got better, I was talking to one of my friends who used to be a, a hardcore poker grinder, and I told them I was gonna interview you, and, uh, their question was: at the table, you can get a lot of information from a small sample size about how a person plays.

But today, GTO is, like, so prevalent that sometimes people forget that you can play exploitatively. What do you think is the state as you think about multi-agent and kinda, like, competition? Is it always gonna be trying to find the optimal thing, or is a lot of it trying to think more in the moment, like how to exploit somebody?

Noam Brown45:47

I'm guessing your audience is probably not super familiar with poker terminology, so I'll just, like, explain this a bit. Uh- A lot of people think that poker is just, like, a luck game, and that's not true. It's actually-

Alessio45:57

Right

Noam Brown45:57

... there's a lot of strategy in poker, so you can win consistently in poker if you're playing the right strategy. So there's different approaches to poker. One is game theory optimal. This is like you're playing an unbeatable strategy and expectation, like you're just unexploitable.

It's kinda like in rock, paper, scissors. You can be unbeatable in rock, paper, scissors if you just randomly choose between rock, paper, and scissors with equal probability because no matter what the other guy does, you know, they're not gonna be able to exploit you-

Alessio46:22

Mm-hmm

Noam Brown46:22

... or, uh, you're gonna win. You're gonna, like, not lose in expectation. Now, a lot of people hear that and they think, like, well, that also means that you're not going to win in expectation because you're just playing totally randomly.

Alessio46:31

Mm-hmm.

Noam Brown46:32

But in poker, if you play the equilibrium strategy, it's actually really difficult for the opponents to figure out how to tie you, and they're gonna end up making mistakes that will lead you to win over the long run.

It might not be a massive win, but it is going to be a win. If you play enough hands for a long enough period of time-

Alessio46:48

Mm-hmm

Noam Brown46:49

... you're, you're going to win in expectation. Now, there's also exploitative poker, and the idea here is that you're trying to spot weaknesses in how the opponent plays. You know, maybe they're, maybe they're not bluffing enough or maybe they fold too easily to a bluff.

And so you start adapting from the game theory optimal balance strategy of, like, you bluff sometimes, you, you don't bluff sometimes, to then playing a very unbalanced strategy that's like, "Oh, I'm just gonna, like, bluff a ton against this person because they always fold whenever I bluff."

Now, the key is that there's a trade-off here because if you're taking this exploitative approach, then you're opening yourself up to exploitation as well.

Alessio47:22

Right.

Noam Brown47:22

And so you have to choose this balance between playing a defensive game theory optimal policy that guarantees you're not going to lose but might not make you as much money as you potentially could versus playing an exploitative strategy that can be much more profitable, but also it creates weaknesses that the opponents can take advantage of and, and trick you.

And there's no way to, to perfectly balance the two. It's kinda like in rock, paper, scissors. If you notice somebody is, like, playing paper for five times in a row, you might think like, "Oh, they're-- they, they have a weakness in their strategy.

I should just be throwing scissors, and I'm gonna take advantage of them." And so on the sixth time, you throw scissors, but actually, that's the time when they-

Alessio47:55

Where they do rock. Yeah

Noam Brown47:55

... throw rock, you know? So y- and you never really know. So you always have this trade-off. The poker AIs that have been extremely successful, and, like, my background is, like, I worked on AI for poker for several years during grad school and made the first superhuman no-limit poker AIs.

The approach that we took was this game theory optimal approach, where the AIs would play this unbeatable strategy and they would play against the world's best and beat them. Now, that also means they, they beat the world's worst.

Like, they would just beat anybody. But if they were up against a, a weak opponent, they might not beat them as severely as a human expert might because the human expert would know how to adapt from the game theory optimal policy to be able to exploit these weak players.

And so there's this kind of unanswered question of like, how do you make an exploitative poker AI? And a lot of people had pursued this research direction. I had, like, dabbled in it a little bit during grad school.

And I think fundamentally it just comes down to AIs not being as sample efficient as humans, you know, we discussed earlier. If a human's playing poker, they're able to get a really good sense of, of the strengths and weaknesses of a player within a dozen hands.

It's, like, honestly really impressive. And back when we were working on AI for poker in like the, you know, mid-twenty tens, you'd have to... These AIs would have to play like 10,000 hands of poker-

Alessio49:06

Mm-hmm

Noam Brown49:06

... to like get a good profile of like who this player is, like how they're playing, where their weaknesses are. Now, I think with more recent technology, that has come down, um, but still the sample efficiency has been a big challenge.

Now, what's interesting is that after working on poker, I worked on Diplomacy. I think we talked about this earlier. And Diplomacy is this, you know, it's a seven-player negotiation game, and- When we started working on it, I took a very game theory approach to the problem.

I, I felt like, okay, we're... It's kinda like poker. You have to compute this game theory optimal policy, and you just play this. You're gonna not lose an expectation, you're gonna win in practice. But that actually doesn't work in Diplomacy, and it doesn't work, again, for it's a question about like how, how much of a, of a rabbit hole do we wanna go down on this?

But, like, basically, when you're playing like the zero-sum games like, like poker, game theory optimal works really well. When you're playing a game like Diplomacy, where there's like you need to collaborate and compete, and you need... There's, there's room for collaboration, then game theory optimal actually doesn't work that well, and you have to understand the players and adapt to them much better.

So this ends up being very similar to the problem in poker of like, how do you adapt to your opponents? In poker, it's about adapting to their weaknesses and take advantage of that. In Diplomacy, it's about adapting to their play styles.

It's kinda like if you're at a table and everybody's speaking French, you don't wanna just keep talking in English. You wanna adapt to them and speak in French as well. That's the realization that I have with Diplomacy, that we need to shift away from this game theory optimal paradigm towards modeling the other players, understanding who they are, and then r-responding accordingly.

And so in many ways, the techniques that we developed in Diplomacy are exploitative. Like, they're not exploitative, they're, they're really, you know, just adapting to the, to the opponents, to the other players at, at the table. Um, but I think the same techniques could be used in AI for poker to make exploitative poker AIs.

If I didn't get, you know, AGI pilled by the incredible progress that we were seeing with language models and, like, shifting my whole research agenda to focusing on, like, general reasoning, probably what I would've worked on next was making these, like, exploitative poker AIs.

It, it would be a really fun research direction to go down. I think it's still there for anybody that wants to do it, and I think the key would be taking the techniques that we use in, in Diplomacy and applying them to things like poker.

Alessio51:17

I think to me the core piece is when you play online, you have a HUD which tells you, you know, all the stats about the other player, like, you know, how much they participate pre-flop, blah, blah, blah. And to me, it's like a lot of these models, from my understanding, are not really leveraging the behavior of the other players at the table.

They're just kind of looking at the board state and kinda working from there.

Noam Brown51:36

That's correct. The, the way the, the way the poker AIs work today, they're just kind of like sticking to their precomputed-

Alessio51:43

GTO

Noam Brown51:43

... GTO-

Alessio51:44

Yep

Noam Brown51:44

... strategy, and they're not adapting to the other players, um, at the table. And, like, you can do various, like, kinda hacky things to get them to adapt, but, you know, they're not, they're not very principled. They're not, they don't work super well.

Alessio51:56

Yep.

Puix51:56

Okay, any grad students listening-

Alessio51:58

Yeah.

Puix51:58

... uh, if you want to work on that, I, I think that is a very, very reasonable research direction that'll at least, uh, get in front of you and, y- you know, get some attention at least.

Noam Brown52:07

Yeah.

Puix52:07

The other thing that this conversation brings up for me is, yeah, well, one of the hypothesis for, like, what is the next step after test-time compute is world models. Is world modeling-

Noam Brown52:17

Uh-huh

Puix52:17

... importance or worthwhile research direction? Like Yann LeCun has been talking about this nonstop, but, like, basically no LLMs have... Like, they have internal world models, but, like, not explicitly a world model.

Noam Brown52:30

I think it's pretty clear that as these models get bigger, they have a world model, and that world model becomes better, uh, with scale. So they are implicitly developing a world model, and I don't think it's something that you need to explicitly model.

Um, I could be wrong about that. You know, there's-

Puix52:48

When, when dealing with people or multi-agents, it might be because you have entities that are not the world, and you're resolving hypotheses of what, which of the many types of entities you could be dealing with.

Noam Brown53:01

You know, there was this, like, long debate in the multi-agent AI community for a long time about, and it's, it's still going on, about whether you need to explicitly model other agents, like other, like other people, or if they can be implicitly modeled as part of the environment.

For a long time, I was like on the, on the, uh, took the perspective of like, of course you have to, like, explicitly model these other agents because they're, they're behaving differently from the environment. Like, they, they take actions, they're unpredictable, you know, they, they have agency.

But I think I've actually shifted over time to thinking that, like, actually, if these models become smart enough, they develop things like theory of mind. They develop an understanding that there are other agents that, like, can take actions and, and have motives and all this stuff, and these models just develop that implicitly with scale and, and more capable behavior broadly.

Puix53:46

Cool.

Noam Brown53:46

So that's the perspective I take these days.

Puix53:48

So, like, what I just said was an example of a heuristic that is not bitter lesson pilled, and you just, it just goes away. Okay.

Noam Brown53:53

Yeah. It, it's really all comes back to the bitter lesson.

Puix53:57

Gotta cite them every, every AI podcast. So one of the interesting findings and most consistent findings, you know, I, I think you were at ICLR, and, uh, one of the hit talks there was about open-endedness, and this guy Tim, who gave that talk, has been doing a b- b- bunch of research about, uh, multi-agent systems too.

One of the most consistent findings is always that, um, it's better for AIs to self-play and improve competitively as opposed to sort of humans training and guiding them, and you find that with like, you know, AlphaZero and R10, whatever that was.

Do you think this will hold for multi-agents, like self-play to improve better than humans?

Noam Brown54:34

Yeah. So okay. So th- this is a great question, and I, I think this is, like, worth, uh, expanding on. So I think a lot of people today see self-play as, like, the next step and maybe the last step that we need for superintelligence, and I think if you're following...

You know, you look at something like AlphaZ- AlphaGo and AlphaZero, we seem to be following a very similar trend, right? Like, the first step in AlphaGo was you do large scale pre-training. Uh, in that case it was on human Go games.

Uh, with LLMs it's pre-training on, you know, tons of, like, internet data. But that gets you a strong model, but it doesn't get you, uh, you know, an extremely strong model. You know, it doesn't get you superhuman model.

And then the next step in the AlphaGo paradigm is you do large scale test-time compute or, like, large scale inference compute, in, in that case with, um, MCTS, and now we have, like, reasoning models that also do, like, this large scale inference compute.

And again, that, like, boosts the capabilities a ton. Finally, with AlphaGo and AlphaZero, you have self-play, where the model plays against itself, learns from those games, gets better and better and better, and just, like, goes from something that's, like, around human level performance to, like, way beyond Human capability.

It's like these Go policies now are so strong that it's just, like, incomprehensible. Like, what they're doing is incomprehensible to humans. Same thing with chess, and we don't have that right now with language models. And so it's like it's really tempting to look at that and say like, "Oh, well, we just need these, like, AI models to now interact with each other and learn from each other, and they're just gonna, like, get to superintelligence."

The challenge, and I kind of mentioned this, like, a little bit when I was talking about, uh, Diplomacy, the challenge is that Go is this two-player zero-sum game, and two-player zero-sum games have this very nice property where when you do self-play, you are converging to a minimax equilibrium.

And I, I guess I should take a step back and say, like, in two-player zero-sum games, two-player zero-sum games are, are chess, Go, even two-player poker, all two-player zero-sum, what you typically want is what's called a minimax equilibrium.

This is that, that GTO policy, this policy that you play where you're guaranteeing that you're not going to lose to any opponent in expectation. I think in chess and Go, that's, like, pretty clearly what you want. Interestingly, in...

when you look at poker, it's not as obvious. In a two-player zero-sum version of poker, you could play the GTO minimax policy, and that guarantees that you won't lose to any opponent on Earth. But again, I mentioned, uh, there's-- you're not going to beat a weak player.

You're not gonna make as much money off of them as you could if you instead played an exploitative policy. So there's this question of, like, do-- what do you want? Do you want to make as much money as possible, or do you want to guarantee that you're not gonna lose to any human alive?

What all the bots have decided is like, well, we're-- what, what all the, like, AI developers in these games have decided is like, well, we're gonna choose the minimax policy. And conveniently, that's exactly what self-play converges to. If you have these AIs play against each other, learn from their mistakes, they converge over time to this minimax policy guaranteed.

But once you go outside of two-player zero-sum games, like in the case of Diplomacy, that's actually not a useful policy anymore. You don't want to just, like, have this very defensive policy, and you're gonna end up in-- with really weird behavior if you start doing the same kind of self-play in things like math.

Test-Time Limits57:30

Noam Brown57:41

So, for example, what does it mean to do self-play in math? You could fall into this trap of like, well, I just want one model to pose really difficult questions and the other model to solve those questions. You know, that's like a two-player zero-sum game.

The problem is that, like, well, you could just, like, pose really difficult questions that are not interesting. You know, you could just, like, get-- ask it to do, like, 30-digit multiplication. It's a very difficult problem for the AI models.

Is that really making progress in the dimension that we want? Like, not really. So self-play outside of these two-player zero-sum games becomes, like, a much more difficult, nuanced question. So I think, and, and Tim, Tim kind of like basically said something similar in his talk, that there's a lot of challenges in really deciding what you're optimizing for when you start to talk about self-play outside of two-player zero-sum games.

My point is that, like, this is where the AlphaGo analogy breaks down and... Uh, not necessarily breaks down, but, like, it's not gonna be as easy as self-play was in AlphaGo.

Puix58:40

What is the objective function then for that? What is the new objective function?

Noam Brown58:44

Yeah, it's a g- it's a good, it's a good question. Yeah, and I think that that's something that, um, you know, a lot of people are thinking about.

Puix58:49

Yeah. Um, I'm sure you are. One of the last podcasts that you did, you mentioned that you were very impressed by Sora. You don't, you don't work directly on Sora, but obviously it's part of OpenAI. I think the, the most recent, uh, new updates or in that sort of generative media space is autoregressive Image Gen.

Is that interesting or surprising in any way that you wanna comment about?

Noam Brown59:11

I don't work on Image Gen, so, like, I've-- my ability to comment on this is kinda limited, but I will say, like, I, I love it. Like, I think it's super impressive. It's like one of those things where, you know, you work on these reasoning models and you think, like, "Wow, we're gonna, like, be able to do all sorts of crazy stuff like advanced science and, um, you know, solve agentic tasks and, and software engineering."

And then there's, like, this whole other- ... like, dimension of progress where you're like, "Oh, you're able to, like, make images and videos now," and it's, like, so much fun. And that's getting a lot more of the attention, to be honest, especially in the general public, and it's probably driving a lot more of the, like, you know, subscription plans for ChatGPT, which is, is great, but I think it's just kinda funny that like...

Yeah, we're also, I promise, we're also working on super intelligence.

Puix59:51

But you can make everything Ghibli. Uh, I, I think the, the delta for me was, um, I was actually harboring this thesis that diffusion was over because of autoregressive emission. Like, there were rumors about this end of last year, and obviously now it's come, now it's come out.

Then Gemini comes out with text diffusion and, like, diffusion is so back and, like, there's this two directions, and it's very relevant for inference of autoregressive versus, um, diffusion.

Noam Brown1:00:17

Mm.

Puix1:00:17

Do we have both? Does one win?

Noam Brown1:00:19

The beauty of research is- ... like, you know, you gotta pursue different, different directions, and it's, it's not, it's not always gonna be clear, like, what is, um, you know, the promising path. Like... And I think it's great that people are looking into different directions and trying different things.

I, I think that there's a lot of value in that exploration, and I think we all benefit from seeing what works.

Puix1:00:39

Any potential in diffusion reasoning, let's say your train of thought-

Noam Brown1:00:43

Probably can't answer that.

Puix1:00:43

Okay.

Alessio1:00:44

So you did a master's in robotics too. Would love to get your thoughts on, one, you know, OpenAI kinda started with the pen spinning trick and, like, the robotic arm they wanted to build. Is it right to work on this humanoid likes?

Do you think that's kinda like the wrong embodiment of AI? Uh, outside of the usual, you know, how long until we get robots, blah, blah, blah, is there something that you think is, like, fundamentally not being explored right now that people should really be doing in robotics?

Noam Brown1:01:09

I did a master's in robotics years ago, and my takeaway from that experience... Uh, first of all, I didn't actually work with robots that much. I was, like, technically in a robotics program. I played around with some Lego robots my, my first week of the program.

But then honestly, I just, like, pretty quickly shifted just working on AI for poker and, um, was kinda nominally in the robotics master's. But my takeaway from, like, interacting with all these roboticists and seeing their research was that I did not wanna work on robots because the research cycle is so much slower and so much more painful when you're dealing with, like, physical hardware.

Like, software goes so much more quickly, and I think that's why we're seeing so much progress with language models and, like, all these, like, virtual coworker kinda tasks but haven't seen as much progress in robotics that, like Physical hardware just is much more painful to iterate on.

On the question of humanoids, I don't have very strong opinions here because this isn't what I'm working on, but I think there's a lot of value in non-humanoid robotics as well. I think drones are a perfect example where, like, there's clearly a lot of value in that.

Is that a humanoid? No. But in many ways, that's great. You know? Like, you don't want a humanoid for, for that kind of technology. I think weekly, um, I think that non-humanoids provide a lot of value.

Alessio1:02:24

I was reading, um, Richard Hamming's The Art of Doing Science and Engineering, and he talks about how when you have a new technological shift, people try and take the old workloads and, like, replicate them just in the new technology versus you actually have to change the way you do it.

And, you know, when I see this video of, like, you know, your humanoid in the house, it's like, well, the human shape has kind of- has a lot of limitations that could actually be improved. But I think people want what's familiar, you know?

It's like, would you put a robot with, like, 10 arms and, like, you know, five legs in your house, or would that be eerie at night when you get up and you see that thing walking around? And is that why we use humanoids?

So I, I think to me, there's almost, like, this local maximum of, like, you know, we gotta make it look like a human, but I think, like, what's like the, the best shape, uh, in-house would be.

Noam Brown1:03:08

I- I am terrible at product design. So I, I am not the person to ask on this. I think there is a question of like, is it better to make humanoids because they're more familiar to us, or is it worse to make humanoids because they're more similar to us, but not quite-

Alessio1:03:22

Right

Noam Brown1:03:22

... identical? Like I, I don't know which one I would actually find creepier.

Alessio1:03:26

Hmm.

Puix1:03:26

Yeah.

Alessio1:03:27

Yeah.

Puix1:03:27

The thing that got me humanoid-pilled a little bit was just the argument that most of the world is made for humans anyway, so if you want to replace human labor, you have to make a humanoid. I don't know if that's convincing.

Noam Brown1:03:39

Again, I don't have very strong opinions in this field-

Puix1:03:41

Yeah. Yeah

Noam Brown1:03:41

... because, like, I don't work in it. Um, I was like weakly in favor of humanoids, and I think what really persuaded me to be weakly in favor of, like, non-humanoids was listening to, um, the Physical Intelligence CEO and, like, some of his pitches about, like, why they're not pursuing-- why they're pursuing, like, non-humanoid robotics.

Puix1:03:57

Okay.

Noam Brown1:03:57

And conveniently, their office is actually, like, very close to here, so if you wanted to-

Puix1:04:01

They're speaking at the, the conference I'm running.

Noam Brown1:04:03

Okay, perfect. Yeah.

Puix1:04:03

So I'm looking forward to that.

Noam Brown1:04:04

So, like, you know, I- I'd say, like, listen to his pitch, and maybe he can convince you- ... that non-humanoid is the way to go.

Puix1:04:09

Awesome. The other one I would refer people to is Jim Fan recently did a talk on the physical Turing test, which, uh, which he did at the Sequoia conference, which, um, was very, very good. Um, he's such a great educator and explainer of things.

Um, it's very hard, especially in that field. Um, cool. We're, we're, we're done asking you about things that you don't work on. So these are just more rapid fires to, to sort of explore some of your boundaries and get, get some quick hits.

Rapid Fire1:04:33

Puix1:04:33

How do you or top industry labs keep on top of research? Like, what are your tools or practices?

Noam Brown1:04:40

Uh, it's, it's really hard. I think that a lot of people have this perception that, like, academic research is irrelevant, and that's actually not the case. I think that we do-- Uh, we look at academic research. I, I think one of the, um, challenges is, like, a lot of academic research shows promise in their papers but then actually doesn't work at scale or even doesn't replicate.

I think if we find interesting papers, like we're, we're gonna try to reproduce that in-house and see if it, like, still holds up and then also does it scale well. But that is like a big source of inspiration for us.

Puix1:05:10

Whatever hits arXiv, literally you do the same as the rest of us? Or do you have like a special process?

Noam Brown1:05:16

Especially if I get recommendations. Like we have an internal-

Puix1:05:17

It's word of mouth

Noam Brown1:05:18

... channel-

Puix1:05:18

Yeah

Noam Brown1:05:19

... where people will post interesting papers.

Puix1:05:20

Yeah.

Noam Brown1:05:20

And, like, I think that's a good source of like, okay, well, this person, uh, that is more familiar with this area thinks that this paper is interesting-

Puix1:05:26

Yeah

Noam Brown1:05:26

... so therefore I should read it.

Puix1:05:27

Yeah.

Noam Brown1:05:27

Um, and I, and similarly, like, I'll keep track of things that are happening in, like, in my space that I think are interesting and, like, if I think it's really interesting, maybe I'll share it.

Puix1:05:34

For me it's like WhatsApp and Signal group chats with researchers, and that's it.

Noam Brown1:05:38

Yeah.

Puix1:05:38

Like

Noam Brown1:05:39

I think it is like... I mean, a lot of people look at things like Twitter, and I think it's really unfortunate that we've reached this point where things need to get a lot of attention on social media for it to be paid attention to.

Um-

Puix1:05:51

That's what the grad students are trained. They're taking classes to do this.

Noam Brown1:05:54

I, I do recommend to like... You know, I've worked with grad students. I work with fewer now because we don't publish as much. But when I was at FAIR publishing papers, like I would tell the grad students I was working with that like you need to post it on Twitter-

Puix1:06:06

Yeah

Noam Brown1:06:06

... and you need to... And we go over like the Twitter thread about like how to present the work and everything and, um, there's a real art to it and it does matter, and it's kind of the sad truth.

Alessio1:06:16

I know when you were doing the ACPC, like the AI poker competition, you mentioned that people were not doing search because they were limited to like two CPUs at inference.

Noam Brown1:06:25

Mm-hmm.

Alessio1:06:26

Do you see similar things today that are like keeping interesting research from being done that might be it's not as popular, it doesn't get you into the top conferences? Like, uh, are there some environmental limiters?

Noam Brown1:06:37

Absolutely. And I, I think one example is for benchmarks that you look at things like humanity's last exam, like you have these incredibly difficult problems, but then are still very easily gradable.

Alessio1:06:48

Mm-hmm.

Noam Brown1:06:48

And I think that actually limits the scope of what you can evaluate these models on if you, if you stick to that paradigm. It's very convenient because, you know, it's very easy to like then score the models. But actually a lot of the things that we wanna, you know, evaluate these models on are kinda like more fuzzy tasks that are not multiple choice questions.

Alessio1:07:03

Mm-hmm.

Noam Brown1:07:04

And making benchmarks for those, for those kinds of things is so much harder, and probably also like a lot more expensive to evaluate. But I think that those are really valuable things to work on.

Alessio1:07:14

And that would fit the Sam Altman GBD 4.5 as like a high taste model in a way. There's kinda like all these like non-measurable things about a model that are really good that maybe people are not.

Noam Brown1:07:26

Well, I think there are things that are measurable, but they're just like much more difficult to measure. And I think that a lot of benchmarks have kinda stuck to this paradigm of posing really difficult problems that are really easy to measure.

Alessio1:07:37

To measure, yep.

Puix1:07:38

So let's say the pre-training scaling paradigm took about five years from like discovery of GPT to scaling it up to GPT 4, and then we give you- you- we give test-time compute five years as well.

Noam Brown1:07:49

Mm-hmm.

Puix1:07:50

So, um, if test-time compute hit a wall by 2030, what would be the probable cause?

Noam Brown1:07:55

It's very similar to pretraining where, like, you can push pretraining a lot further, it just becomes more expensive with each iteration. I think we're gonna see something similar with test-time compute, where like, okay, we're gonna m- get them thinking instead of three minutes, they're for three hours, and then three days, and then three weeks.

Um-

Puix1:08:10

Well, you run out of human life.

Noam Brown1:08:11

Well, I mean, so there's, there's two, there's two, there's two concerns. One is that- ... it becomes much more expensive to get the models to, like, think for that long or, like, scale up test-time compute. Like, as you scale up test-time compute, you're spending more on test-time compute, which means that, like, there's a limit to how much you can spend.

That's one potential ceiling. Now, obviously... Well, not obviously, but, like, I should say that we're also becoming more efficient. These models are becoming more efficient in the way they're thinking, is they're able to do more with the same amount of test-time compute.

And I think that's a very underappreciated point, that it's not just that we're getting these models to think for longer. In fact, if you look at o3, it's thinking for longer than o1-preview for some questions, but it's not, like, a radical difference, but it's way better.

Why? Because it's just, like, becoming better at thinking. Anyway, yeah, these models, um, you're gonna scale up test-time compute, you can only scale it up so much. Like, that becomes a soft barrier, in the same way that pretraining, it's becoming more and more expensive to train better and better pretrained models, uh, or bigger pretrained models.

The second point is that, like, as you have these models think for longer, you kinda get bottlenecked by wall clock time. Like, if you want to-

Puix1:09:07

Yeah

Noam Brown1:09:07

... iterate on experiments, it is really easy to iterate on experiments when these models would respond instantly. It's actually much harder when they take three hours to respond. And what happens when they have three weeks? It takes you at least three weeks to do those evaluations and to then iterate on that.

And, and a lot of this, you can paralyze experiments to some extent, but a lot of it you have to run the experiment, complete it, and then see the results in order to decide on the next set of experiments.

I think this is actually the strongest case for, for long timelines, that the models, because they just have to, like, do so much in serial time, we can only iterate so quickly.

Puix1:09:42

How would you overcome that wall?

Noam Brown1:09:44

It's, it's a challenge, and I think, I think it depends on the domain. So drug discovery, I think, is one domain where this could be a real bottleneck. I mean, if you want to see if something, like, extends human life, it's gonna take you a long time-

Puix1:09:54

Mm

Noam Brown1:09:54

... to figure out if, like, this new d- drug that you developed, like, actually extends human life and doesn't have, like, terrible side effects along the way.

Puix1:10:00

Side note, do we not have perfect models of human chemistry and biology by now?

Noam Brown1:10:04

Well, so this, this is, I think, the thing. And, and again, I, I wanna be cautious here because I'm not actually a biologist or a chemist.

Puix1:10:09

Sure. Yeah.

Noam Brown1:10:09

Like, I know, I know very little about, about these fields. I... Last time I took a biology class was 10th grade in high school. I don't think that there's a perfect, uh, simulator of human biology right now-

Puix1:10:19

Weird

Noam Brown1:10:19

... and I think that that's something that could potentially help address this problem.

Puix1:10:22

Right, that's like the number one thing that we should all work on.

Noam Brown1:10:25

Well, that's one of the things that we're hoping that these reasoning models will help us with.

Puix1:10:28

Yeah. How would you classify midtraining versus posttraining today?

Noam Brown1:10:33

It's, it's such... The, all these definitions are so fuzzy. So I, I don't have, I don't have a great answer there. Uh-

Puix1:10:39

It's a question people have, and you're... and, like, OpenAI is, like, now explicitly hiring for midtraining, and everyone is like, "What the hell is midtraining?"

Noam Brown1:10:46

I think midtraining is between pretraining and posttraining.

Puix1:10:51

Come on.

Noam Brown1:10:51

It's like, it's like, uh... It's, it's not, it's not posttraining. It's not pretraining. It's, like, adding more to the models, but, like, after pretraining. Like, I don't, I don't know if that's-

Puix1:11:00

In interesting ways.

Noam Brown1:11:01

Yeah.

Puix1:11:01

Okay. All right. Well, you know, I'm, I was trying to get some clarity.

Alessio1:11:07

Is the pretrained model now basically like a, just a artifact that then spawns other models, and it's almost like the core pretraining model is never really exposed anymore, and it's the midtraining the new pretraining, and then there's the posttraining once you have the models branched out?

Noam Brown1:11:23

You never interact with an actual just, like, raw pretrained model. Like, if you're gonna interact with the model, it's gonna go through midtraining and posttraining. So, um, so you're seeing the final product.

Puix1:11:32

Well, you don't let us do it, but, you know-

Noam Brown1:11:33

Yeah.

Puix1:11:33

We used to.

Noam Brown1:11:34

Well, yeah. I mean, I guess you, you, you know, there's open source models where you can just, like, interact with the raw pretrained model. Um, but for, for OpenAI models, like, you, they go through a midtraining step and then they go through a posttraining step, and then, and then they're released.

And they're a lot more useful. Like, frankly, if you interacted with a only pretrained model, it would be super difficult to work with, and it would-

Alessio1:11:49

Yeah

Noam Brown1:11:50

... it would seem kinda dumb.

Alessio1:11:51

Yeah. But so-

Puix1:11:52

It'd be, it'd be useful in weird ways, you know? Because there's a mode collapse when you, when you posttrain for it for, like, chat.

Noam Brown1:11:59

Yeah. And in some ways you want that mode collapse.

Puix1:12:02

Yes.

Noam Brown1:12:02

Like, you want, well, you want that collapse of, like-

Puix1:12:03

Yes

Noam Brown1:12:03

... distribution.

Puix1:12:03

To be useful.

Noam Brown1:12:04

Yeah.

Puix1:12:04

I, I get it.

Noam Brown1:12:05

Yeah.

Puix1:12:05

We're interviewing Greg Brockman next.

Noam Brown1:12:07

Exciting.

Puix1:12:07

Uh, you've talked to him a lot. What would you ask him?

Noam Brown1:12:11

What would I ask Greg? I mean, I mean, I get to ask Greg all the time. What, what should you ask Greg? Um-

Puix1:12:15

Like, to, to evoke an interesting response that, like, uh, not, he doesn't get asked enough about, but you know, like, this is something that he's passionate about or you just want his thoughts.

Noam Brown1:12:26

I think in general it's worth asking where this goes, you know? Like, what does the world actually look like in five years? What does the world look like in 10 years? What does that distribution of outcomes look like?

And what could the world or individuals do to help steer things towards, like, the good outcomes instead of the negative outcomes?

Puix1:12:46

Okay. Like an alignment question.

Noam Brown1:12:49

I think people get very focused on what's going to happen in, like, one or two years, and I think it's also worth spending some time thinking about, like, well, what happens in five or 10 years, and what, what does that world look like?

Um-

Puix1:13:00

I mean, does, he doesn't have a crystal ball. Like...

Noam Brown1:13:02

But he, he certainly has, he certainly has thoughts. Yeah, so I think that's worth exploring.

Puix1:13:08

Yeah, okay.

Alessio1:13:09

What are games that you recommend to people, uh, especially socially?

Noam Brown1:13:13

Oh, what are games that I recommend to people? Uh, I've been playing a lot of this game called Blood on the Clocktower lately. Um-

Puix1:13:19

Hmm. What is it?

Noam Brown1:13:19

It's kinda like Mafia or Werewolf. It's become very popular in San Francisco as, uh-

Puix1:13:25

Oh, that's the one we played in your house.

Noam Brown1:13:26

Yeah.

Puix1:13:27

Okay.

Noam Brown1:13:27

Yeah.

Puix1:13:27

Got it. Yeah, yeah.

Noam Brown1:13:27

It's kinda funny 'cause, like, I, I was talking to a couple people now that had told me that it used to be that poker was the, like, way that, like, the VCs and tech founders and stuff would socialize with each other, and actually now it's shifting more towards Blood on the Clocktower.

Like, that's the-

Puix1:13:42

Huh

Noam Brown1:13:42

... the, the thing that people use to, like, um, you know, connect in the Bay Area. And I was actually told that a, a startup held a recruiting event that was a Blood on the Clocktower game.

Puix1:13:55

Wow.

Noam Brown1:13:55

Yeah. So, uh, I guess it's, like, it's really catching on, but- It's a fun game, and I guess you lose less money playing it than you do-

Puix1:14:02

Right

Noam Brown1:14:02

... playing poker. So it's, like, better for people that are not very good at these things. Um, I, I think it's kind of, like, a weird recruiting event, but-

Puix1:14:08

Mm-hmm

Noam Brown1:14:08

... certainly a fun game.

Puix1:14:09

What qualities make a winner here that i- is interesting to hire for?

Noam Brown1:14:13

That's the thing is, like, okay, I guess you get good at-

Puix1:14:16

Ability to lie

Noam Brown1:14:17

... at deception- ... and, like, picking up on deception. Like, is that the best employee? Like, I don't know.

Alessio1:14:24

So my slight final pet topic is Magic: The Gathering.

Puix1:14:27

Ooh.

Alessio1:14:27

So you have... We talked about some of these games, Chess, Go, and they have perfect information. Then you have poker, which is imperfect information in a pretty limited universe.

Noam Brown1:14:36

Mm-hmm.

Alessio1:14:36

You only have a 52-card deck. And then you have these other games that have imperfect information, like a huge pool of possible options. Do you have any idea of, like, how much harder that is? Like, how does the difficulty of this problem scale?

Noam Brown1:14:49

I love that you asked that because, like, I have this, like, huge store of knowledge on AI for imperfect information games.

Alessio1:14:55

Mm-hmm.

Noam Brown1:14:55

Like, you know, this is my, my area of research for so long, and I know all these things, but I don't get to talk about it very often. We've made superhuman poker AIs for No-Limit Texas Hold'em. One of the interesting things about that is that, like, the amount of in- hidden information is actually pretty limited.

Alessio1:15:11

Mm-hmm.

Noam Brown1:15:11

Because you have two hidden cards when you're playing Texas Hold'em, and so the number of possible states that you could be in is 1,326, when you're playing heads-up at least, and, you know, that's multiplied by the number of other players that there-

Alessio1:15:24

Mm-hmm

Noam Brown1:15:24

... are at the table, but it's still, like, not a, a massive number. And so the way these AI models work is they enumerate all the different states that you could be in. So if you're playing, like, six-handed poker, there's five other players, 5 times 1,326, that's the number of states that you can be in, and then you assign a probability to each one, and then you feed those probabilities into your neural net, and you get actions back for each of those states.

The problem is that as you scale the number of hidden possibilities, like the number of sta- of, of possible states you could be in, that approach breaks down, and there's still this very interesting unanswered question of what do you do when the number of hidden states becomes extremely large?

Alessio1:16:02

Mm-hmm.

Noam Brown1:16:02

You know, so if you go to Omaha Poker, where you have four hidden cards, there are things you could do that's kind of, like, that are kind of heuristic that you could do to reduce the number of states, but actually it's still a very difficult question.

And then if you go to a game like Stratego, where you have 40 pieces, so there's, like, close to 40 factorial different states you could be in, then-

Alessio1:16:20

Right

Noam Brown1:16:20

... all these, like, existing approaches that we used for poker kind of break down, and you do need different approaches, and there's a lot of active research going on about, like, how do, how do you cope with that?

So for something like Magic: The Gathering, the techniques that we used in poker would not out of the box work, and it's still an interesting research question of, like, what do you do? Now, I should say this becomes a problem when you're doing the kinds of search techniques that we used in poker.

Alessio1:16:44

Mm-hmm.

Noam Brown1:16:44

If you're just doing model free RL, it's not a problem, and my guess is that if somebody put in the effort, they could probably make a superhuman bot for Magic: The Gathering now. Yeah, there's still some un, uh, unanswered research questions in that space.

Now, are they the most important unanswered research questions? Like-

Alessio1:16:58

Right

Noam Brown1:16:59

... I'm inclined to say no. I think there's, like... The problem is that, like, the techniques that we used in poker to do this kind of search stuff were pretty limited, and, like, if you expand, if you expand those techniques, maybe you get them to work on things like Stratego and Magic: The Gathering, but they're still gonna be limited.

They're not gonna get you, like, superhuman in Codeforces with language models. So I think it's more valuable to just focus on the very general reasoning techniques.

Alessio1:17:21

Mm-hmm.

Noam Brown1:17:21

And one day, as we improve those, I think we'll have a model that just out of the box one day plays Magic: The Gathering at a superhuman level, and I think that's the more important and more impressive-

Alessio1:17:30

Mm-hmm

Noam Brown1:17:30

... research direction.

Puix1:17:31

Cool. Amazing.

Alessio1:17:32

Yeah. Thanks so much for coming on, Noam.

Noam Brown1:17:34

Mm-hmm.

Puix1:17:34

Yeah. Thanks for your time.

Noam Brown1:17:35

Yeah. Thanks. Thanks for having me.