# [State of RL/Reasoning] IMO/IOI Gold, OpenAI o3/GPT-5, and Cursor Composer — Ashvin Nair, Cursor

Latent Space · 2025-12-30

<https://addtry.com/e1454be6-7b35-4e9e-a9d0-5a009ed30a74>

Ashvin Nair, now ML lead at Cursor, traces his path from Berkeley robotics and an OpenAI Dota-era internship to OpenAI's reasoning team (which grew from a dozen to 300+ people) and explains why IOI Gold in 2022 felt like solving AI but didn't change the world—because RL doesn't generalize beyond training distribution. He argues most RL research from 2017-2022 overfit to benchmarks, rewarding complex ideas over simple ones that scale. At Cursor, he sees a unique opportunity for continual learning with policy updates every two hours and product-model co-design, keeping engineers in the loop instead of context-switching. His bet is that the next paradigm shift is continual learning with infinite memory: models experience something once and never forget it, storing millions of deployment tokens in weights without overloading capacity.

## Questions this episode answers

### Why does achieving IOI Gold not mean AI is solved?

Ashvin Nair says if told in 2022 they could achieve IOI Gold, he'd assume AI was solved and there was no point working anymore, but now that it's happened, life is unchanged. He finds it interesting that beating such benchmarks doesn't directly automate real-world jobs, partly because RL doesn't generalize beyond training distribution, and we keep shifting the goalposts.

[7:07](https://addtry.com/e1454be6-7b35-4e9e-a9d0-5a009ed30a74?t=427000)

### How does Cursor use continual learning to improve its code AI?

Ashvin Nair explains that Cursor does policy updates for its models every two hours, enabling continual learning at scale. He envisions a paradigm where models experience something once—like a bug or user pattern—and store it in their weights without forgetting, leveraging the fact that deployment tokens are a drop in the bucket compared to pre-training data, so capacity isn't overloaded.

[34:06](https://addtry.com/e1454be6-7b35-4e9e-a9d0-5a009ed30a74?t=2046000)

### What did Ashvin Nair learn from the RL research era (2017-2022) that didn't pan out?

Ashvin Nair, reflecting on his PhD with Sergey Levine, notes that the community overfitted to benchmarks, introducing many implicit knobs to tune, which didn't generalize. He says academia rewarded mathier, complex ideas over simple ones that actually work. As a result, much of that RL research hasn't been widely used, and the era collapsed into an "RL winter" where startups founded on those premises failed or pivoted.

[9:12](https://addtry.com/e1454be6-7b35-4e9e-a9d0-5a009ed30a74?t=552000)

## Key moments

- **[0:00] Intro & Robotics**
  - [0:28] Ashvin Nair interned at OpenAI in 2017 when it was just robotics, Dota, and 15 interns.
  - [2:02] Lex Fridman: 'Robotics people are the most grounded at NeurIPS; simulation people are the most unhinged.'
  - [2:44] Ashvin Nair: 'Robotics is where you feel AGI the least because it's so far away from working.'
  - [3:50] Ashvin Nair predicts LLM agents will be a trillion-dollar market before robotics is even a $10 billion market.
  - [4:54] Ashvin Nair: Robotics is in the GPT-1 to GPT-2 era right now.
- **[5:54] IOI Paradox**
  - [6:00] Ashvin Nair joined OpenAI in September 2022, right before ChatGPT launched.
  - [7:04] Ashvin Nair on IOI Gold: 'If you told me we could get IOI gold in 2022, I'd assume AI is solved. But life is still the same.'
- **[9:12] RL Winter**
  - [9:12] Ashvin Nair on 2017-2022 RL research: 'We overfit to benchmarks, and most of that research didn't pan out.'
  - [11:04] Ashvin Nair: 'Academia doesn't reward simple ideas that work, rewards mathier ideas with implicit knobs to tune.'
  - [12:38] Ashvin Nair: RL doesn't generalize beyond training distribution, so we must bring economically useful tasks in distribution.
- **[13:33] Co-design**
  - [14:52] Ashvin Nair: Co-design product and model to bring context in distribution — coding is easiest first step because context is codebase.
- **[16:48] One Model?**
- **[18:37] The Blip**
  - [18:44] Ashvin Nair's Blip story: He was at Thanksgiving with OpenAI friends when Sam Altman was fired on Friday.
- **[21:22] o1 Origins**
  - [21:22] Ashvin Nair: OpenAI's reasoning team has grown to around 300 people.
  - [22:11] Ashvin Nair: Ilya Sutskever and Jakub Pachocki had conviction that RL would be the path to AGI.
  - [23:34] Ashvin Nair: 'RL is not about copying the internet; you can do RL to get much better intelligence.'
  - [24:59] Ashvin Nair: The o1 prototype demo was RL on a small model producing interesting reasoning traces and surprisingly good math scores.
  - [26:00] Ashvin Nair: Internally at OpenAI, research progress feels smooth, not the wild swings seen in media.
  - [27:03] Ashvin Nair: Competitive pressure has reduced internal-to-external lead time to 1-2 months.
- **[32:28] Cursor Move**
  - [32:28] Ashvin Nair on joining Cursor: Unique opportunity to co-design product and model because RL doesn't generalize well.
  - [33:34] Ashvin Nair: Cursor can do policy updates every 2 hours for Tab, enabling continual learning at scale.
- **[34:56] Continual Learning**
  - [35:40] Ashvin Nair predicts continual learning with infinite memory — experiencing something once and never forgetting — will be the next paradigm shift within a year.
  - [37:18] Ashvin Nair on Cursor Composer: 'Other smart models are slow, you context switch and it sucks, gives you ADHD.'
  - [38:34] Ashvin Nair: Cursor aims to automate software engineering as a process — write code, check Datadog, hypothesize, re-run.
- **[40:39] RL Future**
- **[44:00] Hiring Cursor**
  - [44:01] Q: What is a good RL interview question? A: Ashvin Nair asks: 'Why is off-policy RL unstable?'

## Speakers

- **Ashvin Nair** (guest)

## Topics

Reasoning, Reinforcement Learning

## Mentioned

Anthropic (company), Cursor (company), DeepSeek (company), NVIDIA (company), OpenAI (company), ChatGPT (product), Codex (product), Composer (product), GPT-4 (product), GPT-5 (product), Tab (product), o1 (product), o3 (product)

## Transcript

### Intro & Robotics

**Host** [0:00]
Okay, we're here at NeurIPS, and we're recording a special LaneSpace coverage of just the folks at NeurIPS, and we're here with Ashvin from Cursor. Welcome.

**Ashvin Nair** [0:08]
Hi. Yeah, thanks for having me.

**Host** [0:10]
So I guess the, like, Ashvin from Cursor is, like, a new identity. I didn't even know if I should say that 'cause you only joined Cursor for three months.

**Ashvin Nair** [0:18]
Mm-hmm.

**Host** [0:18]
Uh, before that, you were OpenAI working on o1-o3.

**Ashvin Nair** [0:20]
Mm-hmm.

**Host** [0:20]
Before that, Berkeley-

**Ashvin Nair** [0:21]
Mm-hmm

**Host** [0:22]
... uh, PhD in RL, just but focused on robotics.

**Ashvin Nair** [0:24]
Robotics, yeah.

**Host** [0:25]
Is it weird switching from robotics to language models?

**Ashvin Nair** [0:28]
Okay, this is kind of interesting 'cause a lot of people have been kind of doing this, um-

**Host** [0:32]
I mean, OpenAI-

**Ashvin Nair** [0:33]
Yes

**Host** [0:33]
... started from robotics.

**Ashvin Nair** [0:34]
Yeah, exactly right. So actually, no, uh, I, I actually was, um, at OpenAI in 2017 also, uh, working on robotics.

**Host** [0:40]
Is it?

**Ashvin Nair** [0:40]
Uh, yeah, I was interning, like, right before my PhD where I worked on robotics there.

**Host** [0:43]
2017, is that-

**Ashvin Nair** [0:44]
Yeah

**Host** [0:44]
... like, Jivan and-

**Ashvin Nair** [0:46]
Uh-

**Host** [0:46]
... that is Jivan over there?

**Ashvin Nair** [0:47]
I think. Uh-

**Host** [0:47]
He, he was o- famously OpenAI's first intern.

**Ashvin Nair** [0:50]
Oh, really? Okay. Then he might have been before. Um, but yeah, uh, there was, like, uh, like, 15 interns. It was a very different company. It was just, like, um, robotics, Dota, and, like, 15 interns that summer all having, like, pretty exciting individual projects.

Like, uh, yeah, that set of interns, if you look over there now, it's kind of cool.

**Host** [1:06]
Yeah.

**Ashvin Nair** [1:06]
Um, but yeah, uh, yeah.

**Host** [1:07]
Anyone, anyone from that class that, like, you would shout out that-

**Ashvin Nair** [1:09]
Um, like, uh, th- there's just, like, a lot of cool papers that came out, like, Léo Pinto now is at, uh, NYU. Um, uh-

**Host** [1:19]
Yeah, in favor of him

**Ashvin Nair** [1:19]
... the, uh, the person who leads, um, reasoning at xAI, I forgot his name.

**Host** [1:22]
Eric?

**Ashvin Nair** [1:23]
Uh-

**Host** [1:24]
Well, he left.

**Ashvin Nair** [1:24]
Not Eric, but yeah, I forgot his name, but, um, he, he, he worked on, like, CAFAC and stuff, I think. Um, yeah.

**Host** [1:30]
Vision dude, Greg?

**Ashvin Nair** [1:31]
Uh, not Greg, but yeah. Um, but yeah, it was, it was, like, an exciting time to be there. Um, but yeah, I, I think, I think robotics is a pretty good fit for LLMs because, like, the switch ends up being pretty, um, like, you know, you kind of do similar things.

Like, you wanna look at a lot of data. It's, like, kind of, um, hard to get s- like, stuff get-- hard to get stuff working in robotics world. I think, you know, it kind of builds, like, very gritty people who, like, look at data a lot, that kind of thing.

So, um, yeah.

**Host** [1:58]
Yeah.

**Ashvin Nair** [1:58]
For whatever reason, I think, like, that transfer is, like, yeah, it's happening a lot, and I think it makes a lot of sense.

**Host** [2:02]
One of my NeurIPS highlights so far, I had dinner with-- It's, like, a small group dinner with Lex Fridman yes, uh, yesterday, and Lex used to be in robotics.

**Ashvin Nair** [2:09]
Mm-hmm.

**Host** [2:10]
And he was like, "My assessment of robotics people, robotics people are the best to talk to at a NeurIPS-"

**Ashvin Nair** [2:15]
Mm-hmm

**Host** [2:15]
... because they're most grounded," he says.

**Ashvin Nair** [2:17]
Mm.

**Host** [2:17]
Because they don't have a choice. They work with the real world, so-

**Ashvin Nair** [2:20]
Yeah

**Host** [2:20]
... looking data, and then, like, the, the most-

**Ashvin Nair** [2:21]
It's hard

**Host** [2:22]
... the most unhinged, the most detached from reality are, like, the, the simulation people.

**Ashvin Nair** [2:26]
I see. Yeah, yeah, yeah. Yeah. I think I agree, yeah. Um, yeah, and I kind of-- I actually did a little bit, bit of both during my PhD. Like, I work in kind of like, you know, s- like, prototype ideas in sim and then get them working on, uh, real world robotics.

And yeah, I mean, probably, like, robotics is where you kind of feel AGI the least.

**Host** [2:44]
Mm-hmm.

**Ashvin Nair** [2:44]
Right? 'Cause it's just so far away from working. Um, now I think, like, uh, I think over the last year maybe there's been demos that have been super interesting from, like, physical intelligence and, like, Sunday and stuff that, yeah, I'm, I'm starting to be like, okay, like, this, this kind of feels like-

**Host** [2:57]
Have you seen Sunday robots themselves?

**Ashvin Nair** [2:58]
Uh, I haven't seen them, no.

**Host** [2:59]
Well, apparently they've been doing demos. Uh-

**Ashvin Nair** [3:01]
Mm-hmm

**Host** [3:01]
... and I'm pretty keen to seeing them.

**Ashvin Nair** [3:02]
Yeah. I've, I've seen the, uh, physical intelligence ones live, and yeah, it's pretty impressive, like, just on, like, a, uh, in, like, someone's living room-

**Host** [3:08]
Folding laundry

**Ashvin Nair** [3:09]
... like, folding laundry and stuff. Like, you know, you can just, like, toss it in there and it works.

**Host** [3:11]
Mark, your buddy must dematerialize.

**Ashvin Nair** [3:14]
Yeah, yeah, yeah, yeah.

**Host** [3:15]
Okay, any, and, and, uh, n- last thing on robotics, and we can kind of pivot to o1-o3. Uh, just OpenAI has, like, is, like, restarting a robotics te- team.

**Ashvin Nair** [3:24]
Mm-hmm.

**Host** [3:25]
Is that, uh, serious? Is that-

**Ashvin Nair** [3:28]
I actually know very little about it, um, 'cause yeah, I was in, like, a pretty different part of the org.

**Host** [3:31]
Yeah.

**Ashvin Nair** [3:31]
So, um, yeah, I mean, I, I think it's serious. Um, I think there's a ton of excitement around robotics right now. I'm actually kind of curious what drives it, 'cause I don't think I fully understand, um, you know, like, uh, like, there's been, like, crazy raises and stuff recently, right, um, uh, for robotics companies.

I guess my own view on it, and so when I, um, left robotics in 2022, I thought I would actually come back to robotics. But I think my view on it now is that it feels like LLM agents are gonna be, like, a trillion-dollar market before robotics is maybe even, like, a $10 billion market.

And this is just because-- So I mean, you know, LLM agents already create value out in the world.

**Host** [4:11]
Yeah.

**Ashvin Nair** [4:11]
Um, robotics, it's, like, kinda hard to make the case that, like, you know, kind of AI robotics, like, uh, does anything that useful yet. And then once it does something useful, um, then you have to make the unit economics work out, and I think that's also quite hard.

Like, reliability, you know. I mean, these robots have to be, like, fixed and, uh-

**Host** [4:30]
Yeah

**Ashvin Nair** [4:30]
... these kind of things. So I think it's kind of hard.

**Host** [4:32]
I would say the market is kind of efficient in that the software, uh, LLM companies are raising tens of billions.

**Ashvin Nair** [4:39]
Yeah.

**Host** [4:39]
And then the robotics companies are raising hundreds of millions, so-

**Ashvin Nair** [4:43]
I think s- this-- very recently it's been, like, single digit billions.

**Host** [4:46]
Oh, really? Okay.

**Ashvin Nair** [4:47]
Yeah. So I, I think, I think that's, I think that's, like, the maybe surprising thing to me is that, like, um, it feels like-

**Host** [4:53]
It's ahead of where it's actually at.

**Ashvin Nair** [4:54]
Yeah. Like, I would say that robotics is in kind of, like, the GPT-1 to GPT-2 era right now.

**Host** [5:00]
Okay.

**Ashvin Nair** [5:00]
Uh, and I haven't worked on robotics closely.

**Host** [5:02]
What, what task would qualify as like, oh, that's, that's the inflection?

**Ashvin Nair** [5:07]
It, it's, it's a little bit like you know it when you see it. I thought the, like, Sunday demos-

**Host** [5:11]
Yeah

**Ashvin Nair** [5:11]
... were kinda cool.

**Host** [5:11]
Additional stuff.

**Ashvin Nair** [5:12]
Like, maybe it's, like, starting to get there where-- And the details matter a lot, where it, it's kind of like it, it can't be-- It has to be in, like, a new scenario, like, in one that you haven't seen before, and maybe on, like, objects you haven't seen before.

**Host** [5:23]
So, like, generalized OD.

**Ashvin Nair** [5:24]
Yeah, exactly.

**Host** [5:25]
Okay.

**Ashvin Nair** [5:25]
And I think that was kind of what GPT-2 was too, right? Is kind of like you start to see hints of, like, cool generalization.

**Host** [5:31]
But, yeah.

**Ashvin Nair** [5:32]
Um, but, like, uh, and I think that's fine. Like, you know, it doesn't have to, like, work out of the box. But, um, yeah, I think at this point especially, it still feels like in robotics you're not exactly investing in a technology probably.

You're just investing in a team.

**Host** [5:44]
Yeah.

**Ashvin Nair** [5:44]
Uh, yeah. I, I'm not in the space whatsoever.

**Host** [5:46]
Sure, sure.

**Ashvin Nair** [5:46]
But, like, that's kind of my impression of it, yeah.

**Host** [5:47]
It, it's actually nice when you're not in there-

**Ashvin Nair** [5:49]
Yeah

**Host** [5:49]
... 'cause you're like, you know as much as, like, most basically everyone else.

**Ashvin Nair** [5:52]
Yeah, yeah.

**Host** [5:52]
So we just kinda speculate.

**Ashvin Nair** [5:53]
Exactly, yeah.

**Host** [5:54]
Remind people there's a robotics team at OpenAI.

### IOI Paradox

**Ashvin Nair** [5:56]
Yeah.

**Host** [5:56]
Uh, so coming back to language models, uh, did you join 4o1 or were you like-

**Ashvin Nair** [6:00]
Um, so I joined right before ChatGBT in, like, um, I think September of 20- Two?

**Host** [6:07]
Yeah.

**Ashvin Nair** [6:07]
Um, so yeah, actually, yeah, I was, like, pretty, uh, burnt out from my PhD, and I was like, "Okay, I'm gonna go to this, like, chill research lab," and, um, and then, like, um, yeah-

**Host** [6:17]
Ooh.

**Ashvin Nair** [6:17]
Yeah, like, ChatGPT happens and, uh, like, you know, everything kind of blew up, and, like, a lot of stuff got kind of, like, refocused.

**Host** [6:23]
But what do-- I guess, what-- So OpenAI-- Obviously, ChatGPT surprised OpenAI.

**Ashvin Nair** [6:27]
Mm-hmm.

**Host** [6:28]
What did they tell you they were looking for you to do? And then obviously it changed, but-

**Ashvin Nair** [6:33]
Um, yeah, I mean, so I joined on the code gen team.

**Host** [6:35]
The Codex.

**Ashvin Nair** [6:36]
Yeah, like-

**Host** [6:37]
O3 Codex.

**Ashvin Nair** [6:37]
Exactly, yeah. It was, like, the team that shipped Codex, but by the time we were working on, like, uh, by the time I joined, we were kind of more so working on the model doing tool use and these kind of things.

**Host** [6:46]
Yeah.

**Ashvin Nair** [6:47]
Um, and so, like, very related to the ChatGPT-- Like, we're kind of like a sister team to the team that made, um, like, ChatGPT.

**Host** [6:51]
Instruction ChatGPT, yeah.

**Ashvin Nair** [6:53]
Yeah, exactly. So yeah, so we were just kind of working on making the models, like, smarter, uh, like, kind of, um, uh, programming competitions, like, um, yeah, how to do, like, SFT for that, that kind of stuff.

**Host** [7:04]
Would an IOI gold have felt reachable in that title?

**Ashvin Nair** [7:07]
Oh, yeah, no, like, crazy. Like, I think, um- I think-- And this is something I've, like, like, repeat to people again and again these days. Like- ... if you told me that, uh, we could have gotten IOI gold then, I would have just assumed that we could all just go on vacation, like, you know, it's all over, like-

**Host** [7:23]
Right

**Ashvin Nair** [7:23]
... AI is solved.

**Host** [7:25]
It-

**Ashvin Nair** [7:25]
Like, no point in working anymore.

**Host** [7:26]
We got it.

**Ashvin Nair** [7:27]
Yeah.

**Host** [7:27]
It feels like nothing's-

**Ashvin Nair** [7:28]
Nothing that much has changed, right? Like, life is still the same.

**Host** [7:31]
Yeah.

**Ashvin Nair** [7:31]
Yeah, so I think that's, like, super interesting. Yeah, I, I, I don't have a great way to explain it, but I think that's, that's actually, like, what I spend a lot of time thinking about, is like, you know, why is that the case?

**Host** [7:40]
Yeah.

**Ashvin Nair** [7:40]
Um, 'cause yeah, I mean, you kind of see this again and again in AI, right? With, like, solving chess, and then, like, it doesn't really matter, and solving Go and, uh, yeah. So you keep seeing it, but, um, yeah, I think it, like, surprises you every single time.

**Host** [7:53]
Yeah. I think maybe-- I think one is we keep moving the m- the goalposts.

**Ashvin Nair** [7:57]
Yeah.

**Host** [7:57]
We're very good at that. And then two is I think actually our just, our definitions of what constitutes AGI is bad.

**Ashvin Nair** [8:03]
Mm-hmm.

**Host** [8:04]
And we don't actually mean what we say when we say, "Oh, when we have achieved this, then we have AGI."

**Ashvin Nair** [8:09]
Yeah.

**Host** [8:09]
So, like, clearly when we have achieved IOI gold with a language model, we have AGI. It's wrong.

**Ashvin Nair** [8:14]
Yeah, and I think, I think, uh, shifting the goalposts to some extent is correct. Like, we keep goodhearting whatever goalposts we have.

**Host** [8:20]
Yeah. Yeah, yeah.

**Ashvin Nair** [8:21]
And, uh, I think it's kind of hard to, like, uh-

**Host** [8:23]
To me, goodheart is, like, too negative.

**Ashvin Nair** [8:25]
Mm.

**Host** [8:25]
It's like, I will cheat to do what I-

**Ashvin Nair** [8:27]
Mm-hmm

**Host** [8:27]
... what you ask me to do.

**Ashvin Nair** [8:28]
Mm-hmm.

**Host** [8:28]
But I don't think it was cheating. It was just-

**Ashvin Nair** [8:30]
Yeah

**Host** [8:30]
... it was just scaling test time compute.

**Ashvin Nair** [8:32]
At a meta level, I think the community, not cheating, but, like, makes a lot of, like, implicit decisions to go after the, you know, evals and benchmarks that matter the most. Um, and so-

**Host** [8:43]
So Supa is verified for sure.

**Ashvin Nair** [8:44]
Yeah, exactly.

**Host** [8:44]
But yeah, IOI gold-

**Ashvin Nair** [8:46]
Yeah

**Host** [8:46]
... hopefully not, not that goodhearted.

**Ashvin Nair** [8:48]
Well, but, but, like, it, it kind of clearly is to some extent, right? Because, uh, like, you know, most programmers in the world cannot do IOI at any decent level, but, like, we're still struggling to, like, automate most programming jobs or, like, you know, there's a lot left to do.

**Host** [9:03]
So it's like, it's like, like we're-- like language models are here, like junior, senior dev, and then suddenly for IOI, you're, like, spiking people.

**Ashvin Nair** [9:09]
Exactly, and, like, there's something so switched about that.

**Host** [9:11]
Yeah, okay.

**Ashvin Nair** [9:12]
Um, I kinda saw this at a meta level also with RL research. Um, so yeah, I did my PhD with Sergey Levine at Berkeley from, like, twenty seventeen to twenty two, and that era of RL research was, like, super interesting because IOI was, like, super hyped, right?

### RL Winter

**Ashvin Nair** [9:28]
Like, starting from about DQN in, like, twenty fifteen. And a lot of the methods that people were really excited about, uh, is like, um, you know, off-policy learning, like value functions, like these kind of things. And, um, somehow that, that stuff hasn't really panned out, I would say, and it's not exactly clear why.

But in the academic literature, we thought we were making a ton of progress. And I think in retrospect, I have to say that, um, we probably kind of overfit to the benchmarks pretty heavily, and, you know, how I see this in retrospect is that we gave ourselves a lot of, like, new knobs to tune and then implicitly kind of tuned those to fit the benchmarks.

Everyone knew that we were doing that at some level, but I think it's hard to appreciate, like, that it's not just happening for a single paper. At kind of like a meta level for the whole community, that's happening too.

**Host** [10:16]
Yeah.

**Ashvin Nair** [10:16]
And I think the result is that, like, I, I don't know, like, uh, a lot of the RL research that came out of that era I don't think is, like, that used, you know? And I think it's kind of for a similar reason that basically we were kind of like benchmark maxing.

**Host** [10:28]
I will full out say there was RL winter, right?

**Ashvin Nair** [10:31]
Mm.

**Host** [10:31]
Like entire startups that were founded based on, uh, premise and at the time-

**Ashvin Nair** [10:35]
Mm

**Host** [10:35]
... basically gave up.

**Ashvin Nair** [10:36]
Mm-mm-mm.

**Host** [10:36]
Some of them-

**Ashvin Nair** [10:37]
Yeah

**Host** [10:37]
... died, some of them pivoted, whatever.

**Ashvin Nair** [10:39]
Yeah, yeah. Yeah, I think, uh, in-- So because I was, like, in, in academia, there was still quite a lot of excitement over it. But yeah, it still felt quite academic. And yeah, I, I in that era was a little bit frustrated because, um, I felt like, you know, one of the pitfalls of academia is that, um, it doesn't really reward, like, simple ideas that work, and instead kind of tends to reward, like, kind of mathier ideas.

Uh, those mathier ideas also give you these, like, kind of implicit knobs to tune, uh, that allow you to, like, overfit. While-

**Host** [11:11]
Mm

**Ashvin Nair** [11:11]
... you know, the, the things that actually work tend to be kind of simple ones that have less knobs and just generalize to, like, many things without-

**Host** [11:17]
There's just, like, less secret sauce to it-

**Ashvin Nair** [11:20]
Exactly

**Host** [11:20]
... apart from just throw a lot of compute in.

**Ashvin Nair** [11:22]
Exactly, exactly. But th- those are things that tend to, like-

**Host** [11:24]
It's, like, not intellectually interesting.

**Ashvin Nair** [11:26]
Yeah, exactly. And from, yeah, from academic point of view, it's like, oh, like, why am I sitting in school? Like, like, yeah, I think, I think for a lot of people who do PhDs, they're kind of wired in a way that they want to, like, think about interesting new stuff.

**Host** [11:36]
Yeah.

**Ashvin Nair** [11:37]
And yeah, like, the, you know, the, the scaling era kind of like, you know, probably sucked for that.

**Host** [11:41]
Scaling era. Uh, is the scaling era over since we're-

**Ashvin Nair** [11:44]
Ooh

**Host** [11:45]
... brought up?

**Ashvin Nair** [11:45]
Um, well, I, I think I've just been, like, paged into that from, um, Ilya Sutskever's interview. I don't think it's over, but there's definitely something interesting happening, right? Like the, the thing I was saying about, like, IOI and IMO.

Um, I think we'll still continue more or less on the same track. Like, clearly, you know, like, uh, these labs are, like, releasing their new pre-trained models, and they're, like, still doing, like, like much better than before. So I think, um- I think scaling is still happening, but, um, I think it's happening in a different way.

Or, you know, it's, like, worth, like, seriously interrogating why is it that we're not just, like, automating all jobs right now. Um, I think my view it is something like RL, the way it's applied to LLMs right now is kind of a weird, funny tool where it doesn't really generalize beyond the training distribution that much.

Um, it generalizes to some extent, and it generalizes in interesting ways, but it's, like, very picky, right? Like, it, it kind of-- it can, it can kill the training distribution, like, completely. Um, it can be, like, best in the world at it, uh, with, like, not that much effort really.

But yeah, it doesn't really generalize. So I think what we had to do is bring the world of economically useful tasks in distribution for RL-

**Host** [12:58]
Mm-hmm

**Ashvin Nair** [12:58]
... if, if we commit to using RL as a tool.

**Host** [13:00]
Mm-hmm.

**Ashvin Nair** [13:01]
And, you know, it might be the case that maybe there's some, like, cool continual learning thing or something that, like, shifts the paradigm next year or something like that.

**Host** [13:07]
Mm-hmm.

**Ashvin Nair** [13:07]
But, um, it really feels like if RL is a tool, then, um, yeah, a big thing that needs to happen is like it's not-- it doesn't feel like intelligence of the models is the bottleneck. It's more like you just have products that bring the entire context of what someone wants to do into the product so that the LLM can, like, see it, and then use the RL on top of that.

**Host** [13:32]
Yeah.

**Ashvin Nair** [13:32]
Um, yeah.

**Host** [13:33]
Have you seen GDP Eval?

### Co-design

**Ashvin Nair** [13:34]
Uh, yeah, I've seen it. Yeah, yeah.

**Host** [13:36]
Is that basically what you're envisioning?

**Ashvin Nair** [13:37]
Um, yeah. I haven't, I haven't looked at GDP Eval closely. Uh, actually, haven't seen exactly. Like, uh, like-

**Host** [13:43]
So-

**Ashvin Nair** [13:43]
... roughly, yeah

**Host** [13:44]
... to recap-

**Ashvin Nair** [13:44]
Yeah, yeah

**Host** [13:44]
... it's 128 tasks-

**Ashvin Nair** [13:46]
Mm-hmm

**Host** [13:46]
... across, like, any white collar job-

**Ashvin Nair** [13:47]
Mm-hmm

**Host** [13:47]
... that takes, like, more than 5% of GDP.

**Ashvin Nair** [13:49]
Mm-hmm, mm-hmm.

**Host** [13:49]
Right? And, uh, they, they, they basically cr- created all the context-

**Ashvin Nair** [13:54]
Mm-hmm

**Host** [13:54]
... the eval on it and, uh, evaluated every model. Like-

**Ashvin Nair** [13:57]
Yeah

**Host** [13:57]
... famously OpenAI, OpenAI's evals team, whoever runs that one-

**Ashvin Nair** [14:01]
Yeah

**Host** [14:01]
... always finds the Anthropic site the best one.

**Ashvin Nair** [14:03]
Yes. Yeah, yeah. It's, uh, yeah. Props to them for, uh-

**Host** [14:07]
They even publish it, yeah

**Ashvin Nair** [14:07]
... doing that, yeah. Yeah.

**Host** [14:08]
It's actual science.

**Ashvin Nair** [14:09]
I think it's good. Um-

**Host** [14:10]
But, but, like, I think, like-

**Ashvin Nair** [14:10]
Yeah

**Host** [14:11]
... uh, in a t- in a, in a sense of, like, generalizing beyond coding competitions to economically useful tasks-

**Ashvin Nair** [14:17]
Yeah

**Host** [14:18]
... that is it.

**Ashvin Nair** [14:19]
I, I, I think that is a-

**Host** [14:20]
What is more important for GG6?

**Ashvin Nair** [14:21]
Yeah. What I'd like to do is kind of like I haven't-- I just haven't read the, like, GDP Eval traces close- closely. Um, it's not clear to me that, you know, like, what, like, what does the job of an accountant entail, and, like, what kind of context needs to be in the product-

**Host** [14:34]
Yeah

**Ashvin Nair** [14:34]
... so, so you can actually do it.

**Host** [14:35]
They have, like, PDFs.

**Ashvin Nair** [14:36]
I see, I see.

**Host** [14:37]
Like, so they, they try to go as close to source documents as possible.

**Ashvin Nair** [14:39]
I see. I see.

**Host** [14:40]
Yeah.

**Ashvin Nair** [14:40]
Yeah. So yeah, I think, like, roughly operating in this kind of thing is what I envision.

**Host** [14:44]
Yeah.

**Ashvin Nair** [14:44]
But-

**Host** [14:44]
Because I think it can't, it can't be, like, an artificial like, "Oh, let me clean up this data for you-

**Ashvin Nair** [14:48]
Exactly. Yeah

**Host** [14:48]
... to make it easy for the LLM to process." No.

**Ashvin Nair** [14:51]
Yeah.

**Host** [14:51]
PDF in an agent, go.

**Ashvin Nair** [14:52]
Uh, yeah, I think that's sort of, like, roughly the right shape of the thing, and I guess how I imagine this being, um, you know, operationalized is that you'd want to, like, co-design the product and the model so that, like, the product for whatever it is, like, I mean, coding is kind of maybe the easiest first step because most of the context that you care about is just your code base and, like, being able to run stuff in the terminal and that kind of stuff.

And still, like, we're not that close really to automating it necessarily. But, you know, for, for, like, all the other jobs, the context is, like, insane, right? It's, like, all the conversations you had with your coworkers, like, your Slack messages.

You know, like, for my, um, so at OpenAI, I was working on kind of like hyperparameter scaling research, and I actually wrote not that much code.

**Host** [15:32]
Like, like grid search or, uh-

**Ashvin Nair** [15:34]
Um-

**Host** [15:35]
... neural architecture search or-

**Ashvin Nair** [15:36]
Uh, no, more like, um, understanding how diff- like, um, science of deep learning in, like, 2020 where it's like, oh, you have to, like, initialize the layers in a particular way to get good scaling law, kind of the analog for that for RL.

**Host** [15:48]
Okay.

**Ashvin Nair** [15:49]
Um, the thing is, like, I didn't write a ton of code, so the LLM, you know, like, writing code is not the, uh, bottleneck, but it's more like, you know, over the course of a year, I, like, run sweeps, look at, like, the interaction between different hyperparameters, and kind of build up that knowledge for, like, a year of, like, just different graphs.

**Host** [16:07]
Mm-hmm.

**Ashvin Nair** [16:08]
And to do my job, the model would also need all those things in context, you know, to, like, successfully, like, you know, kind of automate my job, and you'd kind of want a product that allows you to, like, bring all that context in.

**Host** [16:21]
Mm. Did you have to build it for yourself, or-

**Ashvin Nair** [16:24]
Uh-

**Host** [16:24]
... was there an existing one?

**Ashvin Nair** [16:25]
Uh, no. I mean, like, uh, no, I mean, I, like I, you know, those graphs are just sitting in my head.

**Host** [16:31]
Yeah.

**Ashvin Nair** [16:31]
Right? So, um, I think it's-- it would be pretty hard to, like, go automate that job. But I think what you need to do is build a product that kind of, yeah, has-- brings that context in, and then you want to RL on top of that to, you know, understand, like, to teach the models to, like, use that, um, that context.

**Host** [16:48]
Yeah. Another conversation that I think has really come to a phase this year is kind of the, the death of one model fits all.

### One Model?

**Ashvin Nair** [16:54]
Mm-hmm.

**Host** [16:56]
I feel like the point of the G in AGI is, like, one model fits all.

**Ashvin Nair** [16:59]
Mm-hmm.

**Host** [17:00]
I think OpenAI has, like, clearly abandoned that this year.

**Ashvin Nair** [17:03]
Oh, where do you see that?

**Host** [17:04]
Fiji Simo writing a blog post with the title that we are no longer doing one model fits all.

**Ashvin Nair** [17:09]
Oh, okay. Interesting. Okay. Uh-

**Host** [17:11]
Uh, and I think Mark Chen or one of the other senior people that are not Sam also saying this in a, as-

**Ashvin Nair** [17:16]
Mm-hmm

**Host** [17:16]
... in a podcast. Uh, so, so basically, like, the idea was you, you started with Codex, someone else was doing Instruct GPT, then we launched GPT-4, 4o-

**Ashvin Nair** [17:27]
Mm-hmm

**Host** [17:27]
... I guess, o1.

**Ashvin Nair** [17:28]
Mm-hmm.

**Host** [17:28]
And o1 was kind of a supposed to be, like, a reasoning one model fits all.

**Ashvin Nair** [17:32]
Mm-hmm.

**Host** [17:33]
And then we merged the 4o and o1, o3 line into 5.

**Ashvin Nair** [17:38]
Mm-hmm, mm-hmm.

**Host** [17:39]
And now we're splitting it out into 5 and 5 Codex again. It's like it's just a weird-

**Ashvin Nair** [17:43]
Well, uh, OpenAI is very guilty-- I mean, you know, I, I don't think you should interpret those as, like, scientific facts about the universe. It's just more like, um, OpenAI has a tendency to shift the org chart, basically.

**Host** [17:55]
Yeah.

**Ashvin Nair** [17:56]
Uh, right?

**Host** [17:57]
The world has the tendency.

**Ashvin Nair** [17:58]
Yeah, exactly. So I think, um, a lot of it is related to that. But yeah, I see what you mean by, like, um, yeah, actually I do, I do wonder if, um, yeah, like the current reasoning paradigm, like the current reasoning paradigm is just kind of fitting itself to this kind of peaky in certain areas thing, right?

**Host** [18:12]
Yeah.

**Ashvin Nair** [18:12]
I don't think it's so much a matter of, like, model capacity, though. It's just more of a, another kind of organizational thing that, like- If you care really, like a lot about coding, you probably don't have the data to do all the other stuff.

I don't think it's so much a matter of like if, if you had all the data, probably you would benefit from just like training on all of it, and you'll get some generalization between these. But it's hard to find like one organization that cares about all these at once.

**Host** [18:36]
Yeah, yeah.

**Ashvin Nair** [18:36]
Yeah.

**Host** [18:37]
So before I double-click on just like the o-series in, in OpenAI-

### The Blip

**Ashvin Nair** [18:41]
Mm-hmm

**Host** [18:42]
... uh, I do like to p- ask OpenAI people-

**Ashvin Nair** [18:43]
Mm-hmm

**Host** [18:44]
... who are there, uh, do you have a favorite Blip story?

**Ashvin Nair** [18:46]
Yeah, the, the Blip was crazy for me. Like, uh- Yeah, I was, um ... It was like Thanksgiving.

**Host** [18:52]
Like everyone remembers where they were, what they were wearing.

**Ashvin Nair** [18:54]
Yeah, yeah, exactly. Uh, I was, um, at Thanksgiving with, um, two OpenAI friends actually, and then one of them on like Friday, uh, afternoon is like, "Oh, like Sam Altman just got fired." We were just like co-working together.

**Host** [19:08]
He's like-

**Ashvin Nair** [19:09]
I'm like, "What? Oh, ha ha, like good joke." And then, yeah, no, it was crazy. And then, yeah, it was, it was just like a crazy weekend of just like ups and downs. Like, you know, we thought, yeah, like, uh, like Sam came back.

**Host** [19:22]
You, you signed, you signed a letter, um-

**Ashvin Nair** [19:23]
Um, yeah, I did

**Host** [19:24]
... where you say 95% of people signed it.

**Ashvin Nair** [19:25]
Yeah, yeah, yeah. Um, yeah, I thought, you know, like I, I-

**Host** [19:29]
Are you going to Microsoft or

**Ashvin Nair** [19:31]
Well, I think maybe I had a slightly more compli- like I actually do think that governance feels really important to me.

**Host** [19:37]
Yeah.

**Ashvin Nair** [19:37]
Uh, because it does feel like no matter if we hit AGI in like two years or 10 or whatever, it's not clear that we have a good structure for the governance of it.

**Host** [19:46]
Okay.

**Ashvin Nair** [19:47]
And so it, it is a question that I think we like probably should spend more time on, and I was like, during that period, just pretty willing to be like, "You know what? Like, let's forget about the like equity and stuff."

Like, you know, I think it's like good and healthy to have a conversation about like how exactly the governance should work.

**Host** [20:03]
Okay. You care about this, yeah?

**Ashvin Nair** [20:04]
Uh-huh.

**Host** [20:05]
Right. So now the OpenAI nonprofit has this like secret shadow board of members that determine when, when we've reached AGI.

**Ashvin Nair** [20:13]
Y- yeah. Yeah.

**Host** [20:13]
Is that better?

**Ashvin Nair** [20:14]
Um, yeah, I, I don't have a like, like maybe ... I, I would say I don't have an answer.

**Host** [20:19]
Yeah.

**Ashvin Nair** [20:19]
Like, you know, like it's just, it's not big-

**Host** [20:20]
It's above my pay grade, but like-

**Ashvin Nair** [20:22]
Yeah, and w- even, even, even back then, I was kind of like, "Well"-

**Host** [20:25]
I don't care

**Ashvin Nair** [20:25]
... I, I do care quite a lot.

**Host** [20:27]
Yeah.

**Ashvin Nair** [20:27]
Um, when the Blip happened, one of my reactions was like, well, you know, this nonprofit board stuff, like actually if it takes such somewhat like surprising, maybe erratic actions, like maybe you'd rather just have like, you know, a thing like the Microsoft board, which is kind of like, you know, like probably like all the pensions of the world.

**Host** [20:47]
Like serious people.

**Ashvin Nair** [20:47]
But wait, like sure, like serious people, but also like, you know, the stakeholders are kind of like the whole world because everyone's kind of, you know, through their pensions or something invested in it. Like maybe that is a bit more of a democratic way to run things than having like seven people, uh, run it.

But yeah, I don't, I don't really know. It feels like we haven't solved governance like at all though, right? Like if forget AI, um, even stuff like unhealthy food or like social media, it kind of feels like, like whatever the kind of like capitalistic incentive is, like doesn't actually like, uh, capture kind of good outcomes for society maybe.

### o1 Origins

**Host** [21:22]
Yeah. So about like the, the transition into reasoning, right?

**Ashvin Nair** [21:25]
Mm-hmm.

**Host** [21:26]
Um, you shocked me by when, by mentioning that the reasoning team is 300 people.

**Ashvin Nair** [21:32]
Uh, I, it's, uh-

**Host** [21:33]
If anywhere you draw the line.

**Ashvin Nair** [21:34]
It's, it's kind of like, you know, now that, um, like, you know, when o3 was kind of showed as a product, like I think it just like kind of gets like larger and larger how many people worked on it, so.

**Host** [21:43]
Yeah.

**Ashvin Nair** [21:44]
Yeah. Uh, I think I've like, like lost track of the numbers, but yeah, like a lot of people contribute to the different aspects of like safety and whatever, eval, those kinds of things.

**Host** [21:49]
Yeah, so like original o1, like I saw the video.

**Ashvin Nair** [21:52]
Yeah.

**Host** [21:52]
It's like a dozen people, you know, like-

**Ashvin Nair** [21:53]
Uh, yeah, e- even then, like if you look at all the contributors, it was probably more like 50 to 100 people.

**Host** [21:58]
Okay.

**Ashvin Nair** [21:59]
Um, yeah.

**Host** [22:00]
So, so I mean, like let's, let's tell that story from your point of view.

**Ashvin Nair** [22:02]
Mm-hmm.

**Host** [22:03]
Uh, figuring out what does RL mean there.

**Ashvin Nair** [22:06]
Mm-hmm.

**Host** [22:07]
And I guess was this a branch of any other prior work that you want to credit?

**Ashvin Nair** [22:11]
Yeah. Um, so I think, um, yeah, like setting the scene, I guess, you know, in like 2023, people were kind of talking about, oh, like, is, uh, are scaling, scaling laws dead, this kind of stuff. Um-

**Host** [22:22]
Every year, every NeurIPS.

**Ashvin Nair** [22:23]
Yeah, yeah. But especially- I think especially that year it f- it felt pretty like serious, you know? Yeah, I think, um, in general, OpenAI is really good about like having conviction in something and just like really like from first principles, like going after it, and I think like the people who are kind of most responsible for that is probably like, like Ilya Sutskever and Jakob Pachocki.

I think e- even like, uh, like, uh, Dota was kind of more or less the same template in some ways, right? And that was, um, 2017. And so a lot of the people there have kind of this like AGI in their bones kind of point of view.

Um, and they've basically been convinced that like RL would be the way to get there. So I think for a long, long time people have been convinced that something like that should work, and it's just that it started to work once the, um-

**Host** [23:13]
Human feedback

**Ashvin Nair** [23:13]
... like kind of pre-training got good enough.

**Host** [23:15]
Okay. Yeah, yeah.

**Ashvin Nair** [23:16]
Yeah. Uh, I think human feedback is kind of like a, a bit of like a side branch because-

**Host** [23:20]
Yeah

**Ashvin Nair** [23:20]
... you can't really pour that much compute into it, right? It's like you, you take the model, and you like elicit it to be a little bit better, uh, in terms of personality. But like the people there are really convinced that like at some point, you know, it's not about copying the internet.

Like you can go, um, yeah, do RL, and, uh, you, like, you know, that's like the path to g- like getting much better intelligence. So I think it, it, there was kind of like a long line of k- kind of like returning to RL in for, in like different ways.

Um, and then it's just that around like, yeah, 2023 is when it started like really clicking, and it was kind of interesting 'cause even, you know, it's, it's not like those initial models performed like, um, way better than the existing models 'cause they're like s- smaller scale.

But people were very good at being like, "Oh, like this is kind of interesting." Like, you know, the, the reasoning trace that you see here is kind of not something that you've really seen be so accurate, um, in other models like this, like this one.

Kind of similar to how I think a lot of people didn't really think of GPT or GPT-2 as something that was like super Compelling probably. I know that I personally didn't like, uh-

**Host** [24:27]
GPT-2?

**Ashvin Nair** [24:27]
... that much of GPT-2.

**Host** [24:29]
Yeah.

**Ashvin Nair** [24:29]
I was like, "Okay, whatever." And then I think-- And then GPT-3 happened, I'm like, "Oh, whoa," like, I feel a lot of FOMO sitting in my PhD. It's kind of that where I think it takes a bit of, like, first principle's conviction to, like, um, yeah, decide that, like, oh, this, this thing, like, there's something here, and we should really go scale it up.

And OpenAI's really good about once you decide that something is good, then you just, like, scale it up all the way.

**Host** [24:48]
Yeah, is-- Was there an internal prototype pre-o1-

**Ashvin Nair** [24:51]
Mm

**Host** [24:52]
... that was like, "Okay, this is the thing. We'll fund it to scale it up," right? Like, there, there usually is.

**Ashvin Nair** [24:57]
Uh, yeah, yeah, exactly. Yeah, like-

**Host** [24:58]
What was the thing?

**Ashvin Nair** [24:59]
Uh-

**Host** [24:59]
What, what was the demo that, like, really sort of sold-

**Ashvin Nair** [25:02]
J-just this, like, you know, like, uh, like, running RL on even, like, a pretty small model producing, like, very interesting reasoning traces and, like, getting, like, uh, surprisingly good scores on math.

**Host** [25:13]
Yeah.

**Ashvin Nair** [25:13]
In a way that we couldn't have done without, like, a bunch more pre-training. Um, and then, you know, once, once that looked good, then, you know, more and more resources went to just, like, scaling up that new, like, uh, law.

**Host** [25:25]
Yeah.

**Ashvin Nair** [25:25]
Um, and, you know, things like adding, um, tool use and this kind of stuff.

**Host** [25:30]
Yeah.

**Ashvin Nair** [25:30]
Um, yeah.

**Host** [25:31]
Did-- I think a lot of people make a lot of headlines on the, the large models.

**Ashvin Nair** [25:35]
Mm-hmm.

**Host** [25:36]
But I think a lot, uh, it's very underappreciated, the minis.

**Ashvin Nair** [25:38]
Mm-hmm.

**Host** [25:39]
Uh, how well-

**Ashvin Nair** [25:40]
Mm

**Host** [25:40]
... this solution works.

**Ashvin Nair** [25:41]
Mm.

**Host** [25:41]
Any comments or just, like, discoveries on...?

**Ashvin Nair** [25:44]
Yeah. No-nothing much to say there.

**Host** [25:45]
Yeah.

**Ashvin Nair** [25:46]
I was also, like, not super involved in-

**Host** [25:47]
Yeah

**Ashvin Nair** [25:47]
... the mini stuff. Um, I think maybe one thing, not, not exactly related to that, but, like, it seems like, um, externally, people are kind of very, like, oh, uh, like, research seems to come in these, like, big leaps.

**Host** [25:59]
Okay.

**Ashvin Nair** [26:00]
Uh, but I think internally at OpenAI, it feels very smooth.

**Host** [26:03]
Like, you have a bunch of experiments.

**Ashvin Nair** [26:04]
Yeah.

**Host** [26:05]
Some of them have inconclusive results, but maybe you stack them.

**Ashvin Nair** [26:08]
Yeah, exactly. You stack them, and just, like, you keep scaling, you keep, like, having, like, different runs that, you know, get a little better each time.

**Host** [26:13]
Okay.

**Ashvin Nair** [26:14]
Um, so I think that's maybe one other aspect that's, like, a little underappreciated is that, like, I don't know. Like, in the media, there's just these wild swings between, like, oh-

**Host** [26:23]
It's so like-

**Ashvin Nair** [26:24]
... Google is wearing it. Yeah, exactly. And I think, like, internally at Big Labs, it's just kind of like, oh, we're just, like, chugging along. Like, maybe this month is a little better than last month or something, but it's, like, not as, um, crazy up and down.

**Host** [26:35]
I think the question is, there used to be more of this, and now I know there's less.

**Ashvin Nair** [26:39]
Mm-hmm.

**Host** [26:39]
Which is, well, the stuff we've released, we're like, you know, internally, we're like six months ahead of-

**Ashvin Nair** [26:45]
Mm.

**Host** [26:45]
Like you say, part of the reason why people, uh-- ChatGPT-- OpenAI wasn't that excited about ChatGPT's launch was 'cause they already had GPT-4.

**Ashvin Nair** [26:53]
Mm-hmm.

**Host** [26:53]
They're like, "Oh, we'll just, like, put this out."

**Ashvin Nair** [26:55]
Mm.

**Host** [26:55]
Like, it's we-we're already way ahead.

**Ashvin Nair** [26:56]
Mm-hmm.

**Host** [26:56]
I think now people are just releasing things as they have them. Like yeah.

**Ashvin Nair** [27:00]
I think, yeah, especially 'cause there's some, like, competitive pressure, right?

**Host** [27:03]
Yeah, yeah.

**Ashvin Nair** [27:03]
Uh, I think people are probably pretty worried that, like, if you, if you let a lead linger for too long, that'll, like, grab a lot of market share. Like, I don't know, like-

**Host** [27:12]
I, I-

**Ashvin Nair** [27:12]
... Nano Banana Pro right now is probably, like, you know, it's like it's pretty good.

**Host** [27:16]
They improve it every month.

**Ashvin Nair** [27:16]
Yeah.

**Host** [27:16]
So I would say, like, now the lead, internal to external lead time is about one to two months.

**Ashvin Nair** [27:20]
Mm. Yeah, yeah.

**Host** [27:21]
Which is-

**Ashvin Nair** [27:21]
Exactly

**Host** [27:21]
... tiny.

**Ashvin Nair** [27:22]
Pretty, pretty sure, yeah.

**Host** [27:23]
Tiny.

**Ashvin Nair** [27:23]
Yeah.

**Host** [27:23]
Anything else on reason- on reasoning side? I guess you can talk about on, on, on, so say the work on coding, um, anything surprise you or, like, is an external misconception on o1, o3 side-

**Ashvin Nair** [27:35]
Um-

**Host** [27:35]
... before we go to Cursor?

**Ashvin Nair** [27:36]
Well, um, not really. Like, uh, yeah, I mean, uh, it's, yeah, pretty cool. Like, uh, I, I think, you know, it felt already by, like, maybe early twenty twenty-four like, "Oh, wow, like, this recipe, like, really works, and we can see how far we take it."

And, um, so I think, um, you know, it was, like, very steady progress and, you know, by that point, it was probably pretty pred-uh, predictable that we could, like, you know, really, like, uh, smash, like, you know, things like IMO or IOI.

Yeah. One, one funny thing that kinda happened is, um, while this was happening, uh, I went to this conference called The Curve-

**Host** [28:08]
Yeah

**Ashvin Nair** [28:08]
... um, which is about, like, kind of AI progress.

**Host** [28:10]
And Joseph Gordon-Levitt, like ran it.

**Ashvin Nair** [28:12]
Yes. I went last year. Um-

**Host** [28:14]
Yeah, yeah

**Ashvin Nair** [28:14]
... uh, this was before the o1 stuff was released.

**Host** [28:16]
Yeah.

**Ashvin Nair** [28:16]
And I, like, went to this thing where people were kind of making bets on, um, where we'd be on, um, epoch AIs, like, um, the, the math, the epoch A math ex-exam and, like, humanities last exam and stuff like that.

And, um, their estimates were like, oh, we'll be at, like, ten, twenty percent in, like, twenty twenty-seven. And I think at the time, there was, like, you know, models internally that were, like, already better than their estimates, so they're like-- it's, like, off by, like, you know, two years or something.

And the interesting thing is, like, those are also people who are kind of, like, you know, predicting that there'd be, like, Dyson spheres by, like, twenty thirty-five or something.

**Host** [28:51]
Okay. So-

**Ashvin Nair** [28:52]
So, so, so, like, their, their current estimate is, is way under-

**Host** [28:55]
Yeah, they're too pessimistic in the short term, too optimistic in the long term.

**Ashvin Nair** [28:58]
Um, yeah, well, I don't know if the-- like, I, I-

**Host** [29:00]
Yeah, yeah

**Ashvin Nair** [29:00]
... I didn't-- There might be Dyson spheres by twenty thirty-five. Like, I, and I don't, I don't really know.

**Host** [29:04]
Yeah.

**Ashvin Nair** [29:04]
Uh, but, uh, I think that, that is, like, one interesting aspect, um, is that, yeah, I think people still seem pretty miscalibrated in different ways. Uh, I, I do really appreciate how that community, uh, makes predictions though. Like, um-

**Host** [29:17]
Yeah

**Ashvin Nair** [29:17]
... 'cause I think most of the rest of the world just kind of, like, cynically says, like, "Oh, I, I saw this the whole time." Like-

**Host** [29:24]
Yeah.

**Ashvin Nair** [29:24]
Um-

**Host** [29:24]
And so-

**Ashvin Nair** [29:25]
... I do appreciate that

**Host** [29:25]
... is this EA adjacent?

**Ashvin Nair** [29:28]
Uh, yeah, I think, I think it's-

**Host** [29:30]
That strong-

**Ashvin Nair** [29:30]
Yeah, exactly. It's like-

**Host** [29:31]
That, that, that group

**Ashvin Nair** [29:31]
... it's like, it's like that group. Yeah, yeah.

**Host** [29:32]
Yeah, yeah.

**Ashvin Nair** [29:33]
Um-

**Host** [29:33]
I like that they re- they like to sort of register their opinions-

**Ashvin Nair** [29:36]
Yeah

**Host** [29:36]
... a-ahead of time, and then they-

**Ashvin Nair** [29:37]
And, and I think, like, broadly, uh, the people who've been, you know, uh, the capabilities predictions in that group have been broadly correct if you look, you know, from, like, twenty fifteen to twenty twenty or something, like where I think a lot of people kind of thought that AI was, like, a sham or, like, you know, not really gonna be that useful for a long time.

And actually, you know, it is-- it's somewhere in the, like, twenty thirty-ish thing that, like, it will probably reach, like, human level intelligence.

**Host** [30:03]
Yeah. It's weird. So, like, I, I, I, I feel like a skeptic when I keep saying, like, everyone always predicts that AGI happens in their lifetime.

**Ashvin Nair** [30:10]
Mm-hmm.

**Host** [30:10]
That's very convenient-

**Ashvin Nair** [30:11]
Mm-hmm

**Host** [30:11]
... for whoever. And, like, we have a consistent view of history where you make-- see, like, people in the eighteen hundreds and nineteen hundreds making predictions.

**Ashvin Nair** [30:18]
Mm-hmm.

**Host** [30:18]
It somehow always lands in their lifetime, whatever the, the thing is.

**Ashvin Nair** [30:21]
Yeah.

**Host** [30:21]
But, like, this time it might happen.

**Ashvin Nair** [30:22]
Almost surely, right? Like-

**Host** [30:24]
But no.

**Ashvin Nair** [30:24]
I'm, I'm pretty sure.

**Host** [30:25]
Yeah.

**Ashvin Nair** [30:26]
Yeah.

**Host** [30:26]
Uh, yeah. So, so i-it's a, it's an interesting observation, like how different are we-

**Ashvin Nair** [30:31]
Mm-hmm

**Host** [30:32]
... from our predecessors-

**Ashvin Nair** [30:34]
Mm

**Host** [30:34]
... in, in terms of, uh, developing our technology.

**Ashvin Nair** [30:36]
Yeah.

**Host** [30:36]
Um, did the DeepSeek moment this year, also this year-

**Ashvin Nair** [30:39]
Mm-hmm

**Host** [30:39]
... crazy Uh, change anything internally?

**Ashvin Nair** [30:42]
Uh, not really. Yeah, I think that was-- I think mo- more so just, like, surprised that, um, it created such a moment. Like, it was kinda confusing, right? It was like DeepSeek shows that Nvidia chips are actually more useful than previously thought, and, like, Nvidia's stock, like, goes down a bunch.

Like, it, it was kind of like a-

**Host** [31:02]
I think it's m- more like, okay, well I'll, I'll, I'll do the steelman-

**Ashvin Nair** [31:05]
Yeah, yeah

**Host** [31:05]
... that side, which is, well, you don't need the top-of-the-line Nvidias. You can just use-

**Ashvin Nair** [31:10]
Mm

**Host** [31:10]
... the, the sort of previous generation or the shackled ones they sell to China to do an equivalent amount of work, uh, for, for a-

**Ashvin Nair** [31:19]
I see

**Host** [31:19]
... recent model.

**Ashvin Nair** [31:20]
I see. Yeah, but then it was als- uh, I guess the feeling at OpenAI is that, like, well, I think we, we had a better model already at the time, right? So, um... And it was quite valuable. Like, uh, like smarter models were clearly quite valuable, so you kind of wanted to be at the frontier.

**Host** [31:37]
Okay. I, so I wasn't quite framing this as like a race-

**Ashvin Nair** [31:40]
Yeah

**Host** [31:40]
... uh, dynamics thing between labs. It was just also more like, well, were they right? Were, were their approaches right?

**Ashvin Nair** [31:46]
Mm.

**Host** [31:46]
They had R10, which is kind of like a really cool branch.

**Ashvin Nair** [31:49]
Mm-hmm.

**Host** [31:50]
So more like commentary on what we learned about RL this year in particular.

**Ashvin Nair** [31:53]
Yeah. Yeah. Well, it does seem like basically, um, a lot of the labs have kind of like converged onto some similar-ish way of doing RL, and they're all kind of back at the same level of, like, frontier again.

Like, even the Anthropic, uh, models, like the, uh, Opus t- 4.5, it has this kind of like, uh-- There's this, like, RKGI-2 plot that looks exactly like the OpenAI ones, right?

**Host** [32:17]
What?

**Ashvin Nair** [32:18]
Like, so I think everyone seems to be converging on a pretty similar, um, form of RL. Um, yeah, it's kind of interesting. I think people basically figured out in one way or another to, like, achieve more or less the same thing.

**Host** [32:27]
Yeah.

**Ashvin Nair** [32:28]
Yeah.

### Cursor Move

**Host** [32:28]
Let's talk about the move to Cursor.

**Ashvin Nair** [32:29]
Yeah.

**Host** [32:30]
Why is Cursor accumulating and drawing so many cool RL people?

**Ashvin Nair** [32:34]
Yeah. Um, yeah. So I'd actually kind of like already, like, uh, talked about this so far, I guess. Like, so yeah, I think from the perspective of Cursor, it's like, you know, nice not to be so like, uh, dependent on, like, external labs for everything.

And like, I think there's also, like, um, unique opportunities to co-design the product with the model in ways that I think we couldn't do unless we actually, you know, built the model ourselves and, like, had access to, yeah, making it good.

**Host** [33:02]
Yeah.

**Ashvin Nair** [33:02]
Um, so yeah, that's kind of like, uh, broadly why Cursor's so excited. Um-

**Host** [33:07]
Okay. I'll, I'll push back a little bit, right?

**Ashvin Nair** [33:09]
Uh-huh.

**Host** [33:09]
Uh, OpenAI is, has infinity resources.

**Ashvin Nair** [33:12]
Mm-hmm.

**Host** [33:12]
Uh, infinity data, has Codex. Uh, you could have just stayed.

**Ashvin Nair** [33:17]
Yeah, yeah. Well, uh, actually right around when I was, um, leaving is when, like, I think people started actually, like, using Codex, uh, a lot. So that was kind of like a-- It like, it like happened right after I left, so that was kind of funny.

Like, yeah.

**Host** [33:29]
So, so mostly people are using Cursor internally.

**Ashvin Nair** [33:31]
Mm.

**Host** [33:31]
Maybe a bit of Windsurf because-

**Ashvin Nair** [33:32]
Yeah

**Host** [33:32]
... it was left over from the previous thing.

**Ashvin Nair** [33:34]
Sure. Yeah, yeah, exactly. So it, it wasn't that obvious. But, um, actually I think more to the point, um, this thing I was saying about, like RL is kind of a tool that doesn't really generalize that well. So what you wanna do is bring the entire like, um, kind of test distribution inside your training distribution.

I saw the opportunity to do that at Cursor kind of like directly, and I think the Cursor folks are also just like really excited about that kind of vision. And it's just like a small place where, you know, like the product people sit like right next to the ML people, and I think there's a lot of potential there.

Um, you can kind of see that, um, recently, uh, Jakob Jackson had this blog post about, um, like online tab where, uh, like, you know, we're doing policy-

**Host** [34:15]
Cursor updates every two hours.

**Ashvin Nair** [34:16]
Exactly. Like, like a policy update every two, two hours or something. And I think that's the type of thing that, you know, I think it's like a little hard to do-- Uh, it's like very hard to imagine that at OpenAI, for example, just 'cause like, you know, it-- the product is this like kind of complicated thing and also like the product people and RL people are pretty like, you know, on like different sides of the org.

**Host** [34:34]
I, I think if you put your mind to it, you would. It's like, you know, tab is an autocomplete. It's a smaller model. It's, you know, it's not-

**Ashvin Nair** [34:40]
Yeah

**Host** [34:41]
... as complex, I guess, as-

**Ashvin Nair** [34:42]
Um

**Host** [34:42]
... below them.

**Ashvin Nair** [34:43]
Yeah, but I don't think that's really this like, you know, I think we-- I don't think that's why Cursor was able to do it. It's actually more about like just the org itself being kind of like smaller and a bit more like focused.

**Host** [34:53]
Yeah. Yeah. Okay. Well, I mean, since you're indulging this-

**Ashvin Nair** [34:56]
Mm

**Host** [34:56]
... uh, I think the question about continual learning, which obviously is a big theme-

### Continual Learning

**Ashvin Nair** [35:00]
Mm

**Host** [35:00]
... it's always been a big theme, is bigger this year, is, well, don't you need to curate your data? You can't just like chuck whatever your users are doing in, straight in-

**Ashvin Nair** [35:08]
Mm

**Host** [35:08]
... because that tends to get you towards the middle of the distribution. They actually want to spike it.

**Ashvin Nair** [35:13]
Mm. I guess it depends how you're thinking about continual learning. I mean, I don't know, like, like humans are quite good about dealing with bad data too, right? Uh, like you can see something-- like you can see someone doing something dumb and decide, like you're not gonna do it.

**Host** [35:26]
Filter it out. Yeah.

**Ashvin Nair** [35:26]
Yeah. But, but like it's not even actually filtered out. Like you have, you know, presumably some kind of value function that like says that if you see someone touch a hot stove, like you're not gonna go, you don't need to-- Like it's not just filtering it out, you're actually not gonna do it, right?

Um-

**Host** [35:38]
You could rediscover hot stoves on purpose.

**Ashvin Nair** [35:40]
Yeah, but like you don't, you don't need to. So, um, I think there's something pretty deep there. Um, yeah, like it seems like we're kind of like a few orders of magnitude of like kind of data efficiency, basically, away from like that kind of like, you know, you, you, you do something once or like you, you make a mistake, uh, like you, yeah, you, you introduce like a bug in your code.

You're not gonna do it again. Uh, but the models will happily just like keep doing it, um, even within the same context, but definitely, you know, of course acro- across context.

**Host** [36:10]
Yeah.

**Ashvin Nair** [36:10]
Um, so I think there's something like interesting and deep there is like maybe, yeah, I suspect that it'll be kind of like paradigm shifting in the next like year or something, but I have no idea, like, you know, what it might be.

Um, yeah.

**Host** [36:20]
So is-- Primarily you worked on Composer-

**Ashvin Nair** [36:24]
Mm-hmm

**Host** [36:24]
... Tab, and maybe Search?

**Ashvin Nair** [36:26]
So I, I have-- I've, I've actually just worked on Composer.

**Host** [36:29]
Yeah.

**Ashvin Nair** [36:29]
Um, and that's kind of like the main focus of the company basically.

**Host** [36:32]
Okay.

**Ashvin Nair** [36:32]
Uh, or like the ML group, um, is-

**Host** [36:34]
Which is-

**Ashvin Nair** [36:35]
... shipping a better-

**Host** [36:36]
Yeah

**Ashvin Nair** [36:36]
... um, shipping a better Composer.

**Host** [36:38]
Can you describe, I guess, the impressive, uh, brag a bit about the ML group?

**Ashvin Nair** [36:42]
Yeah, yeah. I mean, I, I think the ML group is great. Um, it's like, uh- You know, it's just like twenty, twenty-five people and, um, you know, I was like honestly like pleasantly like very, very surprised at like how good Composer is, like given the size of the group and, you know, it's not like a big research lab yet.

And, um, uh, yeah, it's-- I think it's like a really good model. You can kind of see that in the reception, and I think it's kind of the start of hints of like co-design with the product in some ways 'cause I think one of the reasons that people really like it is it's smart enough, um, that peop- that you actually wanna use it.

Um, and it's also fast, so you kind of like stay in the loop with the model while you use it 'cause I think all the other smart models have this kind of-- they're, they're slow that you wanna go kind of context switch away and come back, and that sucks, you know?

Like, uh, just as like a, like programmer, it just sucks to kind of context switch. It kind of like gives you ADHD. Like, it, it's like really terrible.

**Host** [37:39]
Yeah.

**Ashvin Nair** [37:39]
Um...

**Host** [37:40]
I agree.

**Ashvin Nair** [37:40]
And I think, uh, yeah, it's like one step in the direction of like being able to be more sync, and I think that's-- like basically the whole company is just really, you know, full of people who want to, you know, code, even like the co-founders, you know, uh, like actually, uh, the co-founders are often some of the best like, like high taste testers, which also kinda gives you a lot of like reassurance that you're gonna ship good stuff.

So yeah.

**Host** [38:03]
Any example test that like maybe Composer doesn't solve yet but you're really motivated to solve?

**Ashvin Nair** [38:08]
Well, yeah, ironically, uh, I feel like I'm actually like a low taste tester in some ways. Because I don't know, like, you know, I just like write like slow like machine learning code and just like think about, um, algorithms and stuff all day.

**Host** [38:19]
Yeah.

**Ashvin Nair** [38:19]
Um, I think more broadly, I'm super excited about co-designing the product so that you can actually, you know, not just-- Right now we're getting better and better at like answering user prompts, um, and I think that's why Composer One is like quite good.

But, uh, you know, what we're really aiming for is like more like, you know, automate software engineering as a process where you like write code, you go look at Datadog, uh, look at what's like happening, then come back and like, you know, maybe have some hypotheses about what's better, like re-rerun stuff.

I think that's the type of thing that we actually want to make the model do.

**Host** [38:54]
Hmm.

**Ashvin Nair** [38:54]
And I do think that Cursor is kind of like uniquely positioned to do that in the sense of like, you know, if we can kind of-- if, if a lot of what a software engineer does kinda ends up in the product, um, I think we can use that to like get better and better at, you know, not just writing code, but kind of like the whole job.

**Host** [39:09]
Yeah. I think that's very inspiring. Just to double-click on just any sort of, uh, RL insights, uh, y- Sasha and Lee have talked a lot about like the internal tooling that you've had-

**Ashvin Nair** [39:19]
Mm-hmm

**Host** [39:19]
... for all the like the cluster visualizations.

**Ashvin Nair** [39:21]
Mm-hmm.

**Host** [39:21]
Is that helpful? Is that what every lab has?

**Ashvin Nair** [39:24]
Yeah, I think, um, the tooling at Cursor is actually really good, um, I think because, you know, it's just kind of like a-- people are just down to like vibe code stuff. They like do-

**Host** [39:34]
Of course

**Ashvin Nair** [39:34]
... test their own stuff. Like, um, so we just have like a lot of good tooling where you can, you know, like have like a SSH session into like, um, our own like, um, user environment or something and like, you know, see if like, uh, code runs the way that like u-users got it to run, like this kind of thing.

I think that's actually, yeah, quite nice. I think basically one of the big lessons in ML in general is that you wanna be like really close to your data and understand your data well, and, um, yeah, I think Cursor's like kind of, yeah, again, kind of like uniquely positioned to do that well, especially-

**Host** [40:04]
It's all internal tooling. You're not buying anything.

**Ashvin Nair** [40:06]
Yeah, it's just like internal. Um, and part of it is just that we're also working on a product where you can understand it really well because it's a code product. Well, like, you know, if-- I don't know, in, in, in, in OpenAI, if I was like to look at like a biology question, I have no idea, like, you know, what, what this is about.

**Host** [40:22]
Yeah. Yeah. Yeah. Interesting. Okay. So I think that's a good overview of, of everything. I guess other than the-- we covered OpenAI and Cursor, just interesting RL work that other people are doing that you're, that you're like still mulling over, it's influential to your thinking, good papers, anything like that.

### RL Future

**Ashvin Nair** [40:39]
Yeah, you know, um, uh, unfortunately, I've like kind of gotten the habit, especially at OpenAI, of like not reading that much external work and just like reading people's like Slack posts internally -

**Host** [40:49]
Nice

**Ashvin Nair** [40:49]
... as like the main like way to like, you know, um, uh, like learn new stuff. Um, no super inspiring recent things have popped up to me. Um, I do think that this like kind of vibe of like, yeah, continual learning just like does feel like, uh, I think there's something super interesting there, and like it feels like, uh, maybe even in academia people could make like a pr- big crack at it.

**Host** [41:10]
And continual learning specifically meaning kind of what TAB is doing?

**Ashvin Nair** [41:15]
Yeah, maybe what TAB is doing, but also just like kind of like in context learning but with like infinite memory or something so that you don't-- Once you experience something in context, it should just like be in your weights, and you shouldn't have to like-

**Host** [41:28]
Yeah

**Ashvin Nair** [41:28]
... make that same mistake again, that kind of thing.

**Host** [41:30]
Why do you think there's-- Okay, but y- so it sh- it should be in your weights, but there, there's a finite capacity for the weights to remember things.

**Ashvin Nair** [41:37]
Yeah. Yeah.

**Host** [41:37]
You will forget things, uh, if you do that too much, right?

**Ashvin Nair** [41:41]
Not r- I mean, you know, you, you start out by memorizing or, you know, like learning from trillions of tokens.

**Host** [41:47]
Yeah.

**Ashvin Nair** [41:48]
Now you're gonna experience like thousands or maybe millions of tokens, and somehow, like, you know, we can-- and those thou- the million tokens are kind of in deployment.

**Host** [41:57]
And it's on- you only need one epoch.

**Ashvin Nair** [41:58]
Yeah, exactly. So-

**Host** [41:59]
Crazy.

**Ashvin Nair** [41:59]
Yeah, so, uh, it feels like if you, if you could learn enough about those million tokens that you're actually in deployment on, um, I don't think you should need-- like I don't think there's a risk of overloading the capacity of your model, right?

'Cause y- you can train on a trillion tokens, and it's like fine.

**Host** [42:14]
Right. Right. So there's proportionately it's a drop in the water.

**Ashvin Nair** [42:16]
Yeah, exactly.

**Host** [42:17]
Water in a bucket.

**Ashvin Nair** [42:17]
Yeah.

**Host** [42:18]
Unless you run it for years and, you know, at some point it start-

**Ashvin Nair** [42:20]
Maybe, yeah.

**Host** [42:21]
Yeah. So basically, I, I, I find it very curious. I've only had one podcast on information theory of language models.

**Ashvin Nair** [42:27]
Mm-hmm.

**Host** [42:27]
Like, what is the theoretical capacity? How much are we using?

**Ashvin Nair** [42:30]
Mm-hmm.

**Host** [42:31]
And you should probably track that.

**Ashvin Nair** [42:32]
Yeah. Yeah. That's a good idea, yeah.

**Host** [42:35]
Like treat the, like the weights.

**Ashvin Nair** [42:36]
Yeah.

**Host** [42:36]
If you want to store things in weights, okay.

**Ashvin Nair** [42:38]
Yeah.

**Host** [42:38]
Treat it as a hard drive. What's the capacity of the hard drive? How much can be stored in there?

**Ashvin Nair** [42:42]
Yeah.

**Host** [42:42]
We know, we know the capacity. It is the number of bits that, you know, th-this-

**Ashvin Nair** [42:46]
Yeah

**Host** [42:46]
... occupied by the language-- by, by the parameters.

**Ashvin Nair** [42:48]
Yeah. Yeah.

**Host** [42:48]
Physically cannot store more than that.

**Ashvin Nair** [42:50]
Yeah. Yeah. Yeah. And it's-- Yeah. I've, I've heard that there's this kind of like someone recently at Cursor, Jacob, kind of brought up this view. I don't know if it's like a more public view that's like, oh, th- there's kind of like a hard drive view of, um- You know, uh, neural networks and kind of like a CPU view of neural networks where, you know, is, is what's happening the weights?

**Host** [43:07]
Yeah.

**Ashvin Nair** [43:07]
Like, yeah, memorizing stuff or is it like you're, like, having some, like, few circuits that, um, do a lot of work?

**Host** [43:14]
Yeah.

**Ashvin Nair** [43:14]
This kind of thing? And yeah, I don't know. Yeah. You know, I would love to, uh ... Yeah, there's like actually so many of these kind of more science-y questions that I would, like, love to explore sometime, but then it really kind of conflicts with, like, empirical stuff, you know?

Like, uh-

**Host** [43:27]
Mm

**Ashvin Nair** [43:27]
... unfortunately, at any given moment in time, it doesn't seem like the most, um, fruit for like, you know, improving something in the sh- especially in the short run, but even in the next, like, couple years is, like, understanding some of these questions.

Um, yeah, I mean, I guess this is technically supposed to be the role of academia, but it's, like, also hard to explore those ideas there without enough compute. Um, but yeah, actually, I would love to, like, go at some point, um, you know, like return to exploring these kind of, like, fundamental science ideas.

**Host** [43:52]
Okay. This is a ... I'm just kind of springing this on you, so you can take some time. What is a good RL interview question that if somebody can answer, they should join Cursor immediately?

### Hiring Cursor

**Ashvin Nair** [44:01]
Ooh, it's a hard, uh, question. Um-

**Host** [44:04]
I'm assuming you do interviews here.

**Ashvin Nair** [44:06]
Yeah, yeah. Um, well, actually, at Cursor, we do, like, work trials.

**Host** [44:09]
Yeah.

**Ashvin Nair** [44:09]
And it's, like, two-day work trials that I actually think that that's, like, more representative.

**Host** [44:12]
'Cause you plug in and-

**Ashvin Nair** [44:13]
Yeah

**Host** [44:14]
... you see how they behave.

**Ashvin Nair** [44:15]
Exactly. Um, so I actually think it's, like, more valuable. Um, this is honestly less of a thing about how you understand RL and a bit more like were you around in the, like, 2017 to '22 era. But, um, it's like why is off-policy RL unstable i- is kind of, I think, like, a, a good, uh, question to, like, yeah, dive into.

**Host** [44:35]
I don't actually know, so I'm going to have to dig into it.

**Ashvin Nair** [44:36]
Yeah.

**Host** [44:38]
Cool. Thank you. That's, that was great conversation. Uh, do you have any sort of call to action?

**Ashvin Nair** [44:44]
Yeah, I mean, uh, you know, we are definitely hiring at Cursor, so, um, yep, if you're interested in working on especially, like, kind of, uh, data and rewards for code, I think that that's, like, a huge need. Um, yeah, please, like, get in touch.

Um, uh, yeah.

**Host** [44:58]
That's it?

**Ashvin Nair** [44:59]
Yeah.

**Host** [44:59]
Thank you.

**Ashvin Nair** [45:00]
Sweet.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
