LALatent SpaceDec 31, 2025· 27:34

[State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI

Josh McGrath, an OpenAI post-training researcher, states that real post-training innovation lies in data quality and signal trust, not optimization methods, with RLVR and token efficiency central. He moved from pre-training (3% compute gains) to post-training (40% behavior change), describes the infrastructure chaos of RL runs, and cites GRPO from DeepSeek Math as underappreciated for providing verifiable reward signals. GPT-5 to GPT-5.1 bumped evals while slashing tokens, emphasizing token efficiency over wall-clock time. The shopping model features interruptibility and chain-of-thought transparency, and personality toggles (Anton vs Clippy) are a key differentiator. For long context, he argues agents with graph walks may be more important than 10M-token windows. He concludes that the education system fails to produce people skilled in both distributed systems and ML research, a critical combination as bottlenecks shift.

  1. 0:00Intro
  2. 1:00Post-Training Shift
  3. 3:04Codex Traps
  4. 4:28Shopping Model
  5. 7:09Personality Toggles
  6. 8:25Data Signal Spectrum
  7. 13:12Token Efficiency
  8. 17:41Context Horizons
  9. 21:24Hiring & Skills
  10. 24:14Team & Balance
  11. 27:11Outro

Powered by PodHood

Transcript

Intro0:00

Host0:03

Light in space, light 2025. Break ups. Light in space, light- Well, here with Josh from OpenAI. Welcome. How would you introduce yourself? How, what, what else?

Josh McGrath0:16

Uh, yeah, I work on a bunch of the thinking models at OpenAI. Um, and like recently I've been sort of focused on doing search-related stuff, but yeah, just a post-training researcher at OpenAI.

Host0:26

Yeah, yeah. And you were on with us for GPT-4.1. We were talking with Michelle, who's on maternity leave. I, I didn't know that. Um, and, uh, now we're at 5.1.

Josh McGrath0:37

Yeah.

Host0:37

It's been a-- it's been a whole generation.

Josh McGrath0:39

Yeah, it's been wild. And like, you know, 4.1 was a non-thinking model, and then since then I, you know, we sort of switched into do-

Host0:45

Was that your last? Was your last?

Josh McGrath0:47

Uh, no. We're still-- we still are releasing non-thinking models-

Host0:50

Oh

Josh McGrath0:50

... um, but that one was the one that we did that was like API specific non-thinking. Um, so, you know, focus has shifted a little.

Host0:58

Yeah. How'd you get into post-training?

Post-Training Shift1:00

Josh McGrath1:00

Um, so previously before OpenAI, I was doing like pre-training, data curation stuff, and I think what I was seeing from like the news and looking at papers is like, oh, it seems like a lot of-

Host1:12

Pre-training is dead.

Josh McGrath1:13

Not pre-training is dead, but I was like-

Host1:15

In those days

Josh McGrath1:15

... oh, there's gonna be so much interesting stuff in post-training, and at that point I was like, I really wanna like make some contributions there. And I mean, it's not even necessarily that like pre-training was dead, but it was definitely changing and like, you know, do I wanna make compute efficiency wins of like 3%, or do I wanna like change the behavior by 40%?

And honestly, it just seemed more, more exciting to go to post-training and many late nights later. Uh, that's definitely true.

Host1:41

It's a different kind of data and engineering discipline too. It's very strange, like the, the, the kind of work that you need, uh, in especially RL, s- like scaling it.

Josh McGrath1:52

Yeah, definitely. I think like, uh, for example, the number of moving parts in an RL run is just a lot higher. Like in some ways-

Host2:00

Easy of order of magnitude or?

Josh McGrath2:02

I don't know if we can do order of magnitude, but if you think about like pre-training, you know, you're moving tokens to many machines, and then you're getting like a s- basically a scaler from them, and then you're back propping.

Host2:12

Yeah.

Josh McGrath2:12

The issue with RL is like you're doing tasks and each task could have like a, a different grading setup and each one of those different grading setups, that's like more infrastructure. And so, you know, when I'm staying up late trying to figure out what's going on with a run, it could be in way more things than there is in a pre-training run generally.

Host2:33

Yeah. And does it matter if you own the code of the task or is it an outsourced third party person or, um... You know, my sense of it and the external sense of it, obviously I don't see it up close, is that you work a lot with external partners and I'm sure also some internal stuff, but which is better?

Josh McGrath2:51

Honestly, I don't think I'll comment like too much on like how many external partners-

Host2:56

Well, there, there are some and there, there's some internal. There, you know.

Josh McGrath2:58

Yeah. There, there's-- we do like stuff but I think-

Host3:00

It's like the technical trade-off of like, well, shit, like I don't own the code.

Codex Traps3:04

Josh McGrath3:04

Oh, okay.

Host3:05

You know.

Josh McGrath3:05

So, well, when it comes to I don't own this code, actually like when, you know, when I'm babysitting a run or something, it doesn't really matter if it's like internal, external, whatever, like do I understand the system that's going underneath?

And I think you end up having to like jump into a lot more code that you're like, "I actually don't know what this does." Because like I'll be watching the-- you know, I, I work on my pieces of a run, um, and then there's also, you know, other people working on it and like, do I understand what their code is doing?

So at-- that way at like twelve thirty in the morning when I'm like, "Something looks wrong," and it's-- I'm like looking at this code, can I like get context fast-

Host3:42

Right

Josh McGrath3:43

... enough to understand-

Host3:43

Do you throw Codex at it?

Josh McGrath3:44

... what's going wrong. Oh, I use Codex so much. It's really changed how I work. I feel like there's a degree to which like sometimes I feel trapped by Codex because if I spend like, you know, thirty, forty minutes writing something that looks like a design doc or something, Codex can do more work than I could do in a few hours in like fifteen minutes.

But then like what do I do during those fifteen minutes after? And like the-- it's actually just like really changed how the flow of my day goes because I have to somehow now manage these like forty-minute sessions with like fifteen minutes where like I could do something, but it's actually not nearly as effective as like this new flow to the day.

Um, so I think I'm still getting used to that, honestly.

Shopping Model4:28

Host4:28

Yeah. Yeah. Uh, yeah, I think it should be interesting for like also just code-based understanding when you're encountering unfamiliar code.

Josh McGrath4:36

Uh, absolutely.

Host4:37

So you briefly before we started talked a little bit about the shopping model, which is like the, the latest, hottest thing, and obviously we're just recording this right after Black Friday, Cyber Monday. First of all, any interesting findings from basically releasing shopping in ChatGPT, uh, right into that period?

Josh McGrath4:53

Okay. Well, I think the first thing is, I don't know like why I would say in a meeting in, you know, August or so like, "Oh, hey, Black Friday's coming up. Like maybe we could- ... maybe we could do a release by them."

In hindsight like, wait, why would I say something like that?

Host5:06

You're like, "Yes."

Josh McGrath5:07

Yeah.

Host5:08

Now you own it.

Josh McGrath5:08

Yeah. Yeah. Ex-exactly. Um, I guess the most interesting thing to me is the new interruptibility and like the, the sort of qualitative experience of using it. Um, and the same thing happens with Codex, right? Like you, you write a prompt and you can like press escape and say like, "Oh, I like-- I messed something up."

Um, and we actually did the same thing in the shopping model. So it shows you its chain of thought with like what products it's looking at, and you can write it new messages saying like, "Oh, you know, I actually wanted to make sure-

Host5:35

Didn't need this. Yeah

Josh McGrath5:36

... yeah, like I wanted USB-C on this," or whatever it is. And like I think that's a really new inter-interesting like interaction paradigm that we have in a couple of our different services, and I'm excited to see, uh, how people use it and if they enjoy it.

Host5:48

Yeah. Why did it have to be its own model and not just like a, a new tool?

Josh McGrath5:53

Stay tuned. I think like there's no reason that we couldn't do it in the same model eventually.

Host5:57

Okay.

Josh McGrath5:57

But I think, you know, if we wanna try out new things, sometimes it makes sense to, to make a new model and I think it just made sense to this time say like, can we do a deep research style sh- model but like- For shopping where it's gonna look really hard, uh, all across the internet for different things.

You know, I think if you look at like Deep Research, the original one and GPT-5 thinking on like high reasoning today, I think you'll see that like eventually the models all sort of converge in their, their capabilities.

Host6:25

Yeah. Would you say that, uh, this is a discussion that also a little spicy that I've kicked off in the community. Uh, there's still maybe thirty percent of the community is still using Deep Research. A lot of them have moved over to just using 5 Thinking as Deep Research.

Is that the spiritual successor? Are they direct replacements? Are there things that we lose in the original Deep Research model if we, if we do that?

Josh McGrath6:46

Um, I mean, I think if you look at our published evals, they're, they're, they look like basically on par-

Host6:51

Yeah

Josh McGrath6:51

... if it's not better. So like, I mean that's personally what I do. I, I use like think, uh, Thinking on High, uh, versus using the Deep Research model. But like, you know, I think every... as we've learned over the past, uh, few months there, sometimes people prefer the quirks of like one model over another.

Host7:06

Mm-hmm.

Josh McGrath7:06

And so people like the Deep Research model, you know, more power to them.

Personality Toggles7:09

Host7:09

People like 4o? Uh, anything special in the 4o post-training that like-- Or, or are people like really responding to personality? Is that like a differentiator that people really care about and, and you-

Josh McGrath7:22

Yeah

Host7:22

... like it's a part of your job to care about personality.

Josh McGrath7:24

Yeah, I mean definitely people like care quite a bit about personality and I think like over the past few months we've been working a lot on giving users more choice over what personality they want. Uh-

Host7:33

Right, which is the, the toggles.

Josh McGrath7:34

Yeah, yeah. So now we, we have those toggles.

Host7:36

Tell us what's your favorite toggle?

Josh McGrath7:37

Uh, honestly custom instruction for like I want-- I personally want my model to like be a tool and so like I don't, I don't necessarily like want the, the warmth or anything. I just want some answers 'cause I'm, you know, I'm mostly using it at work.

Host7:49

Yeah. So I call this the Anton versus Clippy divide. So Anton is the Silicon Valley HBO-

Josh McGrath7:54

Okay

Host7:54

... uh, sort of machine and it's, it's-- it only does work, uh, and doesn't, doesn't try to be helpful or friendly or anything. Uh, I mean, it tries to be helpful but like doesn't try to be cheery whereas Clippy tries to be cheery and I'm like, "Well, stop smiling at me.

I'm like having a problem." It's like... Or-

Josh McGrath8:09

So it sounds like you also come down on the side of like using it, using-

Host8:12

Anton.

Josh McGrath8:12

Yeah.

Host8:13

Yeah. I think a lot of developers want Anton.

Josh McGrath8:15

Yeah.

Host8:15

They, they're just like it just quietly does its work and when it's done it shuts up and-

Josh McGrath8:18

Yeah. Yeah. Well, I think like we're, we're doing a lot of work to provide both like-

Host8:22

Yeah

Josh McGrath8:23

... people Antons and Clippys and I, I hope they all like it.

Data Signal Spectrum8:25

Host8:25

Yeah, yeah. So s- just generally I was thinking about like, well, what can we update people on post-training? You know, what, what do we know today in, at, in NeurIPS 2025 that we didn't know in NeurIPS 2024? I would say like, uh, a lot of people at the time there's, there's still like this whole PPO versus DPO discussion that was there.

That was a whole era.

Josh McGrath8:48

Yeah.

Host8:49

And since then we've moved on to RLVR and, uh, I think a lot of like agents', um, specific RL, uh, training. I guess like am I missing any large chunks of the post-training debates that are going on?

Josh McGrath9:02

Yeah, I mean, so not necessarily debates internal but like my read personally from like looking at different papers that are coming out, when you look at like an RLVR paper or like a RLHF paper, they read more like an optimization paper and to me like the, the sort of interesting thing that's going on is we have this like spectrum of how high quality a signal is.

So like really at the end of the day like RLHF, RLVR, they're both policy gradient methods but the-- what's different is just like the input data and it's always interesting to me that we call RLHF non-verifiable because we've trained this model to be good at like predicting human feedback so in some sense that's like verification.

But obviously it's-

Host9:45

It's human preference-

Josh McGrath9:46

Yeah

Host9:46

... rather than truth.

Josh McGrath9:47

Yeah, yeah. But like if the-- if like your value of truth is like does the user like this more? Like see there's, there's something strange that I think we haven't like looked at that axis of, okay, well how like sort of clean is this signal?

How much do I trust it? And like I totally agree that, you know, you don't necessarily trust the RLHF signal as much as like is this the solution to this polynomial? But I think there's a whole spectrum of like how high quality is this signal?

What's gonna happen when I like do a lot of optimization against it? And that's very different than I think worrying about like the variance of different gradients, which I think is what you end up seeing in a lot of the, the papers that are currently coming out.

Um, rather than being like very data centric, they're pretty optimization centric even though I think the, the innovation really is, is where the data's coming from.

Host10:33

Yeah. And before I wanna go broad before I go deep.

Josh McGrath10:36

Yeah.

Host10:36

Um, any other discussions that maybe you're having at NeurIPS or, or sort of, uh, round about this time on post-training debates? Like what are, what are-- You meet your, your peer at Anthropic and DeepMind and what, what are you talking about?

Josh McGrath10:48

Uh, well, Anthropic and DeepMind, we're all saying and working on stuff and things, you know.

Uh, and I think like it's more so, uh, talking a lot more broadly with my, my friends there, or, or we're just talking about, man, the, the infra is so hard to keep up. We're not, not necessarily talking too much about methods directly.

Um-

Host11:09

Because on, on one level it kind of doesn't matter.

Josh McGrath11:11

Yeah. And I think also like there's, there's something that's very different about academic work where like the-- what really matters is how narrativizable it is and I think that's one of the reasons you see like a lot of optimization papers come out is a lot of the data work there's a less clear narrative around it.

Host11:28

The, I think my, my s- the data and the scaling is actually more important than the specific algorithm.

Josh McGrath11:33

Yeah. But it doesn't have like necessarily the same narrative that you get out of like some of the papers that you see here.

Host11:39

Yeah.

Josh McGrath11:39

And so like there becomes more of a like given a, a specific vertical how do I like understand that? Um, and I, I wish there was actually more papers on it here but I think it can sometimes be harder to wrap up into a, a clean story.

Host11:53

Yeah. That's also something that like we're, we're actually having a lot of conversations about with other folks as well like what's, what's next, right? Like what, what, where do you go from here now that we, we have like some kind of roadmap?

I think what's interesting also for me is I guess the innovations that are exposed by the Chinese models are Maybe copies or, like, discussions of what's going on in, in the labs. I think obviously GRPO, uh, when you mention, like, a lot of these RL optimizations, uh, they come out as they, they, they present themselves as optimizations.

Um, GRPO came out in the DeepSeek Math paper, which, uh, when it came out, I read it and I was like, "Okay, this is kind of cool. It's, like, a little bit cheaper," but, like, it does seem to have a more broad impact, I think, on the industry as a whole than was in- initially appreciated.

And I just wanna... I, I don't feel like we've processed that enough.

Josh McGrath12:41

Yeah, definitely. I mean, like, yeah, the-- as you said, it came out in the DeepSeek Math paper and, like, it's an interesting optimization method, but it's, like, the more interesting thing that they have a new reward signal that they sort of like re- that we can really, really trust.

Like, when, you know, you find the answer to a math problem, it's a lot less debatable than like, "Oh, well, is this thing that the human preferred actually what we want to do?"

Host13:03

Yeah.

Josh McGrath13:03

Like, you wanna be right at math.

Host13:05

Yeah, yeah.

Josh McGrath13:06

And so I think in some ways it's underappreciated in, I would say, what's getting published. Yeah.

Host13:12

Yeah. Let's talk about, I guess, long horizon.

Token Efficiency13:12

Josh McGrath13:15

Yeah.

Host13:15

What do people consider in terms of, like, very long horizon? Like, we're talking like 30 hours, s- you know, more than, more than a day of autonomy. Does, does it-- is it just more of the same, or is there anything, like, sort of qualitatively different?

Josh McGrath13:28

Okay, so first off, what I would first say is I tend to think more in terms of, like, actual number of tokens-

Host13:34

Yeah

Josh McGrath13:34

... than, than time because-

Host13:35

Well, I think-

Josh McGrath13:35

Yeah

Host13:36

... the human-

Josh McGrath13:36

Yeah

Host13:36

... human in the loop can take a while.

Josh McGrath13:38

Yeah. Well, and also, like, it, it gives you a different, uh, measure to optimize against, right? Like, as I was saying earlier with, um, when I use Codex, it does something that would take me much longer, uh, you know, it would take me like four hours in, you know, 10 minutes.

What we can actually push on there is token efficiency. So like-

Host13:55

Yeah

Josh McGrath13:56

... and the, it-

Host13:56

That is a huge, huge research area.

Josh McGrath13:58

Yeah. And so you can see like from 5 to 5.1 our, our overall evals, you know, we, we bumped some, but if you look at a 2D plot of how many tokens it takes for us to get that, it went way down.

Um, and so I think that's like an act-

Host14:13

Did you guys cheer when you, when you had that? Like, that was such a great chart.

Josh McGrath14:16

Dude, I live by those charts. Like that- There was... That-

Host14:19

Was, was your chart? Okay.

Josh McGrath14:20

Not necessarily that, but, like, that shape of chart.

Host14:22

Yeah, yeah.

Josh McGrath14:23

Like, I think that's something that we think about a lot just because it contributes so much to your experience. Like, how long does it take to, to do this task?

Host14:30

Yeah.

Josh McGrath14:31

And I think the other thing is, as you're pushing that token efficiency, it changes, you know, how many tool calls can I make, and like how many different things can the agent do in a reasonable number of tokens that we can actually serve.

Host14:44

Yeah.

Josh McGrath14:44

Um, and so I personally think in terms of tokens, yeah.

Host14:47

I think the interesting thing or the, the hard to understand thing from the outside is having explicit router in GPT-5, and but then also basically having an implicit router in terms of the thinking, uh, spending thing. That conflates a little bit, right?

Like, at some point you do kinda need to merge them or else you're just gonna get these, like, weird bumps where sometimes the router at the top decides something and it's wrong, and actually if you just handed it to GPT-5 it would have figured it out.

Josh McGrath15:15

Yeah. And I think, you know, we'll figure out the correct abstractions over time. I think, like, there's a-

Host15:20

Is the intention is still to merge? Because it-- that's what it was said in the paper.

Josh McGrath15:23

Yeah. I think, like, eventually, you know, we'll have, uh, AGI and, like, you're not gonna have to worry too much about how hard to, to think directly. It'll just, you know... We'll have a one tool that you always go to, and it knows how long to think for and things like that.

I think the, the abstractions and the way that we drive these things today, it, it'll change and, like, you know, I think even the amount that we've changed from, you know, having a non-thinking model to you can choose between two and like, you know, now we can sort of route and how hard do you wanna think?

We're adding lots of knobs and, you know, o- eventually it'll, it'll probably simplify.

Host15:55

Yeah. Another super interesting knob that everyone is doing is context compaction or memory compaction. What's going on there?

Josh McGrath16:02

Uh, nothing to share at the moment.

Host16:04

Nothing to share.

Josh McGrath16:04

Yeah.

Host16:04

Okay. Clearly an important feature, clearly inspired by Codex usage as well, obviously. Um, but I think, like, from the engineer's point of view, it feels like I used to do that as part of my harness and now it's now the model's doing it for me, and I don't know how to th- think about that, like, in terms of, I guess, I used-- I'm used to having more control, and now I have less.

Josh McGrath16:27

Yeah. Is there, is there a specific, like-

Host16:28

There, there's no specific question. I just-- I'm just getting, like, feedback on, like, well, is this a trend that, like, we need-- will you-- it's basically a permanent fact of life for- from here on out.

Josh McGrath16:37

Oh, uh, I see. Uh, you know, I don't know. I worked on long context. That was why I was on last was for 4.1-

Host16:43

Yeah

Josh McGrath16:44

... where we, you know, I think 10X'd the- ... the effective context window for 4.1, and so there'll always be some dance of like, well, if we wanna push as much as what we can do, not only should we increase the length of the context window, but, like, we should also have strategies for keeping that context window available for as long as possible.

Um, I'm guessing that both things will, will sort of happen just because we wanna put as much power into the models as possible.

Host17:09

Yeah.

Josh McGrath17:09

Yeah. I think we're still in a period where we should all be expecting changes in the, the interfaces that all of the models give to us. Um, that way we can improve the models. 'Cause if we lock the interface, I think what would be sad from, from my perspective is if we lock the interface, if we discover something new about models, we might sort of trap that improvement under a, an interface that needs to change.

Host17:29

Got it. Talking about long context as well, there is some discussion about, I guess, context rot or, like, the utilization of the context. Even if you gave us like a million token context, probably wouldn't use all of it.

Context Horizons17:41

Host17:41

What's the recommendation there? Where are things going? Are we gonna have, I guess, perfect context by next year? Is that, is that an impossible dream? I don't know.

Josh McGrath17:49

Um, no, it's not an impossible dream. Uh, I think I'll give a shout-out to some of the evals that we did for 4.1 with, uh, called Graph Walks where-

Host17:57

I love Graph Walks. We covered this in an-- in a podcast.

Josh McGrath18:01

Yeah, yeah. Yeah, we did. You know, I think if you look over time, all of those, uh, all of those evals are, are still climbing.

Host18:07

So many.

Josh McGrath18:07

And I think one of the interesting things about that is you have to do complicated transformations across the entire context window. Like, that's sort of the issue with, um- Those heat map plots of the, those different-

Host18:18

A needle in a haystack.

Josh McGrath18:19

Yeah, but the problem is if you only have to sample from one point in the context window, it's like sort of easy whereas with those graph walks problems, you're having to do multiple transformations across the entire context window.

Um, and so I think keep watching those. I think they've, they've been climbing. They'll continue to climb. I would say that that's definitely like a temporary issue that we are climbing on over time.

Host18:41

Yeah. So and then, like, is 10 million tokens realistic? Is 100 million? Like, where does ... Does the ... Is there a natural end or there's no end and we just are going as far as the eye can see?

Josh McGrath18:52

Oh, gosh, I, I don't know.

Host18:53

Yeah.

Josh McGrath18:53

Like, what, what do you think? Yeah.

Host18:55

I, I feel like, okay, there are use cases that require billions, and there are use cases that require many, many billions, maybe trillions.

Josh McGrath19:02

Yeah. Out, out of curiosity, like, what, what would be billions of tokens?

Host19:05

Uh, we just had, uh, a context engineering discussion about, uh, like a RAG code base over support issues for a, a company, and it was 100,000 documents totaling about 8 billion tokens. You can't stick that in a context window for now.

Josh McGrath19:19

That's fair. I guess the ... So I would still say, like, I don't know, but I think I've been, like, really surprised. It reminds me of when I was doing, like, more information retrieval stuff and, like, uh, BM25 and these, like, very simple, like, Ngram indexes were, like, just super hard to beat.

I think the agents with Grep are, like, they feel really similar to me where it's, like, just unreasonably effective.

Host19:40

Infinitely effective.

Josh McGrath19:40

Yeah. Uh-

Host19:41

So, so that, that ... But at then I will not use your 10 million token context window even if you gave it.

Josh McGrath19:45

Maybe but, like, what if we're using that context window in service of, like, some larger goal that just has a lot of, uh, sub-search calls? Which is why I'm saying, like, I, I just don't know, and I think that's what makes it so exciting.

Host19:59

Yeah, yeah. I would say also, like, the other, other modalities like video, um, would eat up a lot and, like, uh, then obviously the hard sciences have proteins and all, all that which, uh, uh, a lot, a lot of information just encoded in, uh, in, in physics.

Um, so, so I mean, yeah, I, I, I'm mixed feelings about it just because I'm like, well, this will never scale, uh, not with, like, full attention and, uh, we, we probably just need to invest in systems anyway, which means we- we're good with what we have.

I mean, like, get your, get your graph walks up.

Josh McGrath20:31

Yeah.

Host20:31

But, like, I don't know if we need, like, 10, 100x when actually maybe we need to figure out ways to 1,000, 1 million X.

Josh McGrath20:38

Yeah.

Host20:39

Right? Like, like these, these are just different slopes.

Josh McGrath20:41

I mean, I'm definitely ... I'm, I'm glad that you're happy with the, the current context windows. I think my dream would be to push it and see what happens anyway but I think there-

Host20:49

Right. The engineer's-

Josh McGrath20:50

Yeah

Host20:50

... the eng- engineer's incentive is always to say, "Well, the systems matter more than the models."

Josh McGrath20:53

Yeah.

Host20:53

And the researcher's incentive is say, "Well, screw your systems or we'll just build the models."

Josh McGrath20:58

Oh, no, I mean, it's so differently. Yeah.

Host21:00

Yeah.

Josh McGrath21:00

And I think that's one of the most, like, sorta beautiful things about post-training at OpenAI is everyone-

Host21:06

Co-design. Yeah

Josh McGrath21:07

... yeah, it's, it's also co-design. Like, you know, I, I spend a lot of time just doing our system stuff, and I also do lots of stuff like where I'm making graph walks, and I'm, like, doing a lot more, like, things on the learning side, and I think it's a great culture to have a place where people just move seamlessly between the two.

Hiring & Skills21:24

Host21:24

Yeah. What are you guys hiring for? Presumably you're hiring. What are you guys hiring for that is hard to hire? What is the skill set that is like we really need this, can't find it? Please, everyone, go skill up on this.

Josh McGrath21:36

As my, my definitely personal opinion here, I think we're still having trouble at, not at OpenAI, but I think as a whole producing lots of people that do lo- want to do lots of both systems work and ML work.

And I think if, if you're trying to push the frontier, you don't know which place is currently bottlenecking the frontier and it, and it changes all the time. I mean, even within one project it might change multiple times where the, the current bottleneck is, but I think the education system we have right now isn't really optimized for that.

So, like, I personally ... I studied math, and then I was very, very lucky to have some, like, great mentors after school that, like, taught me to be a, a good software engineer, but it seems like if we're gonna be in this place for a while, and I th- I think we will be, we should probably be producing more students that are great at doing both, you know, distributed systems and, like, a lot of core engineering as well as the statistics and other, like, things that are required to be a good machine learning researcher.

Host22:34

If we were to throw Codex at it, obviously we can't do Codex at everything. That's why it's still ... Let's say, like, which will progress faster? Which is more solvable by LLM?

Josh McGrath22:44

That is a ... That's a spicy question. Uh-

Host22:47

You can't say they're both equally hard. I don't know. Maybe, maybe they are. I mean, they're differently hard.

Josh McGrath22:51

I think-

Host22:52

But, like, one is more hill climbable than the other. Which is it? 'Cause then we can go, go do it.

Josh McGrath22:56

Okay. I think, I think one thing that's slightly simpler about some of the ML research, like or, you know ... ML research is also distributed systems to be clear, but, like, some of the things that I would say, like, get traditionally called ML research are things that you can treat a bit more of as a black box whereas, like, you know, the, the environment to train on, you know, building these, these different systems is actually just, like, a complicated engineering problem.

And so theoretically the ... I would say that they're, like, probably roughly equal. Um, but I think the ... There's some, there's some amount of effort I feel like to making the, the environments for it.

Host23:39

Yeah.

Josh McGrath23:39

Yeah. But-

Host23:39

Let's say they require-

Josh McGrath23:40

Oh, I guess. Yeah

Host23:41

... let's say they require GPUs in themselves as well.

Josh McGrath23:44

Yeah, I ... Yeah. I, I guess they both would, but yeah, that's, that would be my guess. It ... But I, I don't have high confidence in it.

Host23:51

Well, so, so a lot of people are building this, like, AI scientist, right? That, that automates A- A-

Josh McGrath23:55

Yeah

Host23:55

... research. You guys have a, your own benchmark on, on Paperbench, and though that's the one area that, um, like for example at Cognition we've just decided to not do 'cause it's so hard. Okay. Any other people on the post-training team that you wanna shout out have done, like, uh, interesting work this year, they should get more attention but they're, they're not getting credit?

Team & Balance24:14

Josh McGrath24:14

Well, uh, okay, for sure everyone on the shopping team that I was just working with.

Host24:18

Right.

Josh McGrath24:18

So, like, Andrew Hoyos, uh, Manuka Strata, John Holman, all, all, all great people. Yeah, Isa Fulford, obviously the, the manager for it.

Host24:26

And she was the original deep-

Josh McGrath24:28

Deep research

Host24:28

... research, uh-

Josh McGrath24:29

Yeah, yeah

Host24:29

... person?

Josh McGrath24:30

Yep. Yeah.

Host24:30

Yeah. There was, like, three of them. Yeah.

Josh McGrath24:31

Yeah, yeah. And so definitely that part of the team. But, I mean, everyone, everyone is so great. Like, I think it's hard to, to give out a list. It's a, it's a really fun time on, on post-training right now.

It's exciting every day. Yeah, it feels like, uh, we're all enjoying our Diet Cokes together in the office late at night.

Host24:48

Yeah. I ... Oh, I, I did want to squeeze this in before we end. Nobody actually serious is saying that pre-training is dead. It's just a meme. There's a lot of work going on in pre-training. And in fact, actually, a lot of my researcher friends are saying too much money is going to post-training.

That's also spicy. I don't know. One of the charts I hold in memory from this year is the Grok 4 chart. I don't, uh, I don't know if you've seen it, but, uh, it's basically saying, "Well, we scaled pre-training to here and about the same level of, uh, about this level of compute, and now we're spending the same level of compute on post-training as well."

That's very controversial, I guess, to me because, like, we're all used to post-training taker- taking orders of magnitude less data, compute, whatever, and obviously we're scaling that up now. Do we get to a point where they're equal? I don't know, but that's a topic for conversation, I, I think.

How much do we invest in this versus more, like, different po- pre-training?

Josh McGrath25:40

Yeah.

Host25:40

Like you said both.

Josh McGrath25:40

Yeah. Yeah. So first off, neither, neither one of those is dead. I think it's really interesting to sort of be living through something that I ... You know, all of my other, like, historic or technological revolutions are things that I read about in, in history books and, like-

Host25:56

This one's live. This is happening.

Josh McGrath25:57

Yeah, this one, this one's live.

Host25:57

We don't know the end yet.

Josh McGrath25:59

Yeah. And so there's this almost, like, fog of war where I'm like, oh, did people think that, like, we got, like, the steam, like, uh, the steam engine and they would have, you know, the factories ... I don't know if you know this, but, like-

Host26:09

Yeah

Josh McGrath26:09

... the factories, they used to be, like, very linear because you had to drive, like, one motor across an, an entire room and it made it so when electricity got developed they just tried to do the same thing and they're like, "Ah, this isn't all that useful," and it took, I think, like, a couple of decades before they realized, wait, if we have electricity we can move the little, like, stations in what's- whatever is most ergonomic and then, you know, manufacturing was transformed by electricity.

And I think, like, it really gives me no confidence in being like, "Oh, this thing is dead."

Host26:39

Yeah. Our timelines are so short.

Josh McGrath26:41

Yeah. Yeah.

Host26:41

But usually the way, like, good ideas get experimented and funded and propagated, actually that's, that's still on the human timeline.

Josh McGrath26:47

Yeah.

Host26:47

That's not on AI timeline.

Josh McGrath26:49

Yeah. Yeah. And so I think, like, things will maybe be, like, dormant but it'll be spiky. Like, there'll be all of a sudden, you know, some whoop-

Host26:55

Yeah.

Josh McGrath26:56

Yeah, yeah. And, and then we'll all feel different. It's like we're ... Uh, what's, what's the meme? Uh, it's so over. We're so back.

Host27:00

Oh, yeah.

Josh McGrath27:01

It's gonna be that many times and I think having, like, a, some, some emotional, uh, stabilizing to it is probably gonna be good for, for everyone's sanity.

Host27:11

Yeah. More sanity. Well, thank you so much for joining. Thanks for, uh, all the great post-training this year.

Outro27:11

Josh McGrath27:16

Yeah. Thank you.

Host27:17

Yeah.

Josh McGrath27:17

And yeah, th- continue giving feedback. I love to hear what you think.

Host27:20

Yeah. Awesome.