Intro & Robotics0:00
Okay, we're here at NeurIPS, and we're recording a special LaneSpace coverage of just the folks at NeurIPS, and we're here with Ashvin from Cursor. Welcome.
Hi. Yeah, thanks for having me.
So I guess the, like, Ashvin from Cursor is, like, a new identity. I didn't even know if I should say that 'cause you only joined Cursor for three months.
Mm-hmm.
Uh, before that, you were OpenAI working on o1-o3.
Mm-hmm.
Before that, Berkeley-
Mm-hmm
... uh, PhD in RL, just but focused on robotics.
Robotics, yeah.
Is it weird switching from robotics to language models?
Okay, this is kind of interesting 'cause a lot of people have been kind of doing this, um-
I mean, OpenAI-
Yes
... started from robotics.
Yeah, exactly right. So actually, no, uh, I, I actually was, um, at OpenAI in 2017 also, uh, working on robotics.
Is it?
Uh, yeah, I was interning, like, right before my PhD where I worked on robotics there.
2017, is that-
Yeah
... like, Jivan and-
Uh-
... that is Jivan over there?
I think. Uh-
He, he was o- famously OpenAI's first intern.
Oh, really? Okay. Then he might have been before. Um, but yeah, uh, there was, like, uh, like, 15 interns. It was a very different company. It was just, like, um, robotics, Dota, and, like, 15 interns that summer all having, like, pretty exciting individual projects.
Like, uh, yeah, that set of interns, if you look over there now, it's kind of cool.
Yeah.
Um, but yeah, uh, yeah.
Anyone, anyone from that class that, like, you would shout out that-
Um, like, uh, th- there's just, like, a lot of cool papers that came out, like, Léo Pinto now is at, uh, NYU. Um, uh-
Yeah, in favor of him
... the, uh, the person who leads, um, reasoning at xAI, I forgot his name.
Eric?
Uh-
Well, he left.
Not Eric, but yeah, I forgot his name, but, um, he, he, he worked on, like, CAFAC and stuff, I think. Um, yeah.
Vision dude, Greg?
Uh, not Greg, but yeah. Um, but yeah, it was, it was, like, an exciting time to be there. Um, but yeah, I, I think, I think robotics is a pretty good fit for LLMs because, like, the switch ends up being pretty, um, like, you know, you kind of do similar things.
Like, you wanna look at a lot of data. It's, like, kind of, um, hard to get s- like, stuff get-- hard to get stuff working in robotics world. I think, you know, it kind of builds, like, very gritty people who, like, look at data a lot, that kind of thing.
So, um, yeah.
Yeah.
For whatever reason, I think, like, that transfer is, like, yeah, it's happening a lot, and I think it makes a lot of sense.
One of my NeurIPS highlights so far, I had dinner with-- It's, like, a small group dinner with Lex Fridman yes, uh, yesterday, and Lex used to be in robotics.
Mm-hmm.
And he was like, "My assessment of robotics people, robotics people are the best to talk to at a NeurIPS-"
Mm-hmm
... because they're most grounded," he says.
Mm.
Because they don't have a choice. They work with the real world, so-
Yeah
... looking data, and then, like, the, the most-
It's hard
... the most unhinged, the most detached from reality are, like, the, the simulation people.
I see. Yeah, yeah, yeah. Yeah. I think I agree, yeah. Um, yeah, and I kind of-- I actually did a little bit, bit of both during my PhD. Like, I work in kind of like, you know, s- like, prototype ideas in sim and then get them working on, uh, real world robotics.
And yeah, I mean, probably, like, robotics is where you kind of feel AGI the least.
Mm-hmm.
Right? 'Cause it's just so far away from working. Um, now I think, like, uh, I think over the last year maybe there's been demos that have been super interesting from, like, physical intelligence and, like, Sunday and stuff that, yeah, I'm, I'm starting to be like, okay, like, this, this kind of feels like-
Have you seen Sunday robots themselves?
Uh, I haven't seen them, no.
Well, apparently they've been doing demos. Uh-
Mm-hmm
... and I'm pretty keen to seeing them.
Yeah. I've, I've seen the, uh, physical intelligence ones live, and yeah, it's pretty impressive, like, just on, like, a, uh, in, like, someone's living room-
Folding laundry
... like, folding laundry and stuff. Like, you know, you can just, like, toss it in there and it works.
Mark, your buddy must dematerialize.
Yeah, yeah, yeah, yeah.
Okay, any, and, and, uh, n- last thing on robotics, and we can kind of pivot to o1-o3. Uh, just OpenAI has, like, is, like, restarting a robotics te- team.
Mm-hmm.
Is that, uh, serious? Is that-
I actually know very little about it, um, 'cause yeah, I was in, like, a pretty different part of the org.
Yeah.
So, um, yeah, I mean, I, I think it's serious. Um, I think there's a ton of excitement around robotics right now. I'm actually kind of curious what drives it, 'cause I don't think I fully understand, um, you know, like, uh, like, there's been, like, crazy raises and stuff recently, right, um, uh, for robotics companies.
I guess my own view on it, and so when I, um, left robotics in 2022, I thought I would actually come back to robotics. But I think my view on it now is that it feels like LLM agents are gonna be, like, a trillion-dollar market before robotics is maybe even, like, a $10 billion market.
And this is just because-- So I mean, you know, LLM agents already create value out in the world.
Yeah.
Um, robotics, it's, like, kinda hard to make the case that, like, you know, kind of AI robotics, like, uh, does anything that useful yet. And then once it does something useful, um, then you have to make the unit economics work out, and I think that's also quite hard.
Like, reliability, you know. I mean, these robots have to be, like, fixed and, uh-
Yeah
... these kind of things. So I think it's kind of hard.
I would say the market is kind of efficient in that the software, uh, LLM companies are raising tens of billions.
Yeah.
And then the robotics companies are raising hundreds of millions, so-
I think s- this-- very recently it's been, like, single digit billions.
Oh, really? Okay.
Yeah. So I, I think, I think that's, I think that's, like, the maybe surprising thing to me is that, like, um, it feels like-
It's ahead of where it's actually at.
Yeah. Like, I would say that robotics is in kind of, like, the GPT-1 to GPT-2 era right now.
Okay.
Uh, and I haven't worked on robotics closely.
What, what task would qualify as like, oh, that's, that's the inflection?
It, it's, it's a little bit like you know it when you see it. I thought the, like, Sunday demos-
Yeah
... were kinda cool.
Additional stuff.
Like, maybe it's, like, starting to get there where-- And the details matter a lot, where it, it's kind of like it, it can't be-- It has to be in, like, a new scenario, like, in one that you haven't seen before, and maybe on, like, objects you haven't seen before.
So, like, generalized OD.
Yeah, exactly.
Okay.
And I think that was kind of what GPT-2 was too, right? Is kind of like you start to see hints of, like, cool generalization.
But, yeah.
Um, but, like, uh, and I think that's fine. Like, you know, it doesn't have to, like, work out of the box. But, um, yeah, I think at this point especially, it still feels like in robotics you're not exactly investing in a technology probably.
You're just investing in a team.
Yeah.
Uh, yeah. I, I'm not in the space whatsoever.
Sure, sure.
But, like, that's kind of my impression of it, yeah.
It, it's actually nice when you're not in there-
Yeah
... 'cause you're like, you know as much as, like, most basically everyone else.
Yeah, yeah.
So we just kinda speculate.
Exactly, yeah.
Remind people there's a robotics team at OpenAI.
IOI Paradox5:54
Yeah.
Uh, so coming back to language models, uh, did you join 4o1 or were you like-
Um, so I joined right before ChatGBT in, like, um, I think September of 20- Two?
Yeah.
Um, so yeah, actually, yeah, I was, like, pretty, uh, burnt out from my PhD, and I was like, "Okay, I'm gonna go to this, like, chill research lab," and, um, and then, like, um, yeah-
Ooh.
Yeah, like, ChatGPT happens and, uh, like, you know, everything kind of blew up, and, like, a lot of stuff got kind of, like, refocused.
But what do-- I guess, what-- So OpenAI-- Obviously, ChatGPT surprised OpenAI.
Mm-hmm.
What did they tell you they were looking for you to do? And then obviously it changed, but-
Um, yeah, I mean, so I joined on the code gen team.
The Codex.
Yeah, like-
O3 Codex.
Exactly, yeah. It was, like, the team that shipped Codex, but by the time we were working on, like, uh, by the time I joined, we were kind of more so working on the model doing tool use and these kind of things.
Yeah.
Um, and so, like, very related to the ChatGPT-- Like, we're kind of like a sister team to the team that made, um, like, ChatGPT.
Instruction ChatGPT, yeah.
Yeah, exactly. So yeah, so we were just kind of working on making the models, like, smarter, uh, like, kind of, um, uh, programming competitions, like, um, yeah, how to do, like, SFT for that, that kind of stuff.
Would an IOI gold have felt reachable in that title?
Oh, yeah, no, like, crazy. Like, I think, um- I think-- And this is something I've, like, like, repeat to people again and again these days. Like- ... if you told me that, uh, we could have gotten IOI gold then, I would have just assumed that we could all just go on vacation, like, you know, it's all over, like-
Right
... AI is solved.
It-
Like, no point in working anymore.
We got it.
Yeah.
It feels like nothing's-
Nothing that much has changed, right? Like, life is still the same.
Yeah.
Yeah, so I think that's, like, super interesting. Yeah, I, I, I don't have a great way to explain it, but I think that's, that's actually, like, what I spend a lot of time thinking about, is like, you know, why is that the case?
Yeah.
Um, 'cause yeah, I mean, you kind of see this again and again in AI, right? With, like, solving chess, and then, like, it doesn't really matter, and solving Go and, uh, yeah. So you keep seeing it, but, um, yeah, I think it, like, surprises you every single time.
Yeah. I think maybe-- I think one is we keep moving the m- the goalposts.
Yeah.
We're very good at that. And then two is I think actually our just, our definitions of what constitutes AGI is bad.
Mm-hmm.
And we don't actually mean what we say when we say, "Oh, when we have achieved this, then we have AGI."
Yeah.
So, like, clearly when we have achieved IOI gold with a language model, we have AGI. It's wrong.
Yeah, and I think, I think, uh, shifting the goalposts to some extent is correct. Like, we keep goodhearting whatever goalposts we have.
Yeah. Yeah, yeah.
And, uh, I think it's kind of hard to, like, uh-
To me, goodheart is, like, too negative.
Mm.
It's like, I will cheat to do what I-
Mm-hmm
... what you ask me to do.
Mm-hmm.
But I don't think it was cheating. It was just-
Yeah
... it was just scaling test time compute.
At a meta level, I think the community, not cheating, but, like, makes a lot of, like, implicit decisions to go after the, you know, evals and benchmarks that matter the most. Um, and so-
So Supa is verified for sure.
Yeah, exactly.
But yeah, IOI gold-
Yeah
... hopefully not, not that goodhearted.
Well, but, but, like, it, it kind of clearly is to some extent, right? Because, uh, like, you know, most programmers in the world cannot do IOI at any decent level, but, like, we're still struggling to, like, automate most programming jobs or, like, you know, there's a lot left to do.
So it's like, it's like, like we're-- like language models are here, like junior, senior dev, and then suddenly for IOI, you're, like, spiking people.
Exactly, and, like, there's something so switched about that.
Yeah, okay.
Um, I kinda saw this at a meta level also with RL research. Um, so yeah, I did my PhD with Sergey Levine at Berkeley from, like, twenty seventeen to twenty two, and that era of RL research was, like, super interesting because IOI was, like, super hyped, right?
RL Winter9:12
Like, starting from about DQN in, like, twenty fifteen. And a lot of the methods that people were really excited about, uh, is like, um, you know, off-policy learning, like value functions, like these kind of things. And, um, somehow that, that stuff hasn't really panned out, I would say, and it's not exactly clear why.
But in the academic literature, we thought we were making a ton of progress. And I think in retrospect, I have to say that, um, we probably kind of overfit to the benchmarks pretty heavily, and, you know, how I see this in retrospect is that we gave ourselves a lot of, like, new knobs to tune and then implicitly kind of tuned those to fit the benchmarks.
Everyone knew that we were doing that at some level, but I think it's hard to appreciate, like, that it's not just happening for a single paper. At kind of like a meta level for the whole community, that's happening too.
Yeah.
And I think the result is that, like, I, I don't know, like, uh, a lot of the RL research that came out of that era I don't think is, like, that used, you know? And I think it's kind of for a similar reason that basically we were kind of like benchmark maxing.
I will full out say there was RL winter, right?
Mm.
Like entire startups that were founded based on, uh, premise and at the time-
Mm
... basically gave up.
Mm-mm-mm.
Some of them-
Yeah
... died, some of them pivoted, whatever.
Yeah, yeah. Yeah, I think, uh, in-- So because I was, like, in, in academia, there was still quite a lot of excitement over it. But yeah, it still felt quite academic. And yeah, I, I in that era was a little bit frustrated because, um, I felt like, you know, one of the pitfalls of academia is that, um, it doesn't really reward, like, simple ideas that work, and instead kind of tends to reward, like, kind of mathier ideas.
Uh, those mathier ideas also give you these, like, kind of implicit knobs to tune, uh, that allow you to, like, overfit. While-
Mm
... you know, the, the things that actually work tend to be kind of simple ones that have less knobs and just generalize to, like, many things without-
There's just, like, less secret sauce to it-
Exactly
... apart from just throw a lot of compute in.
Exactly, exactly. But th- those are things that tend to, like-
It's, like, not intellectually interesting.
Yeah, exactly. And from, yeah, from academic point of view, it's like, oh, like, why am I sitting in school? Like, like, yeah, I think, I think for a lot of people who do PhDs, they're kind of wired in a way that they want to, like, think about interesting new stuff.
Yeah.
And yeah, like, the, you know, the, the scaling era kind of like, you know, probably sucked for that.
Scaling era. Uh, is the scaling era over since we're-
Ooh
... brought up?
Um, well, I, I think I've just been, like, paged into that from, um, Ilya Sutskever's interview. I don't think it's over, but there's definitely something interesting happening, right? Like the, the thing I was saying about, like, IOI and IMO.
Um, I think we'll still continue more or less on the same track. Like, clearly, you know, like, uh, these labs are, like, releasing their new pre-trained models, and they're, like, still doing, like, like much better than before. So I think, um- I think scaling is still happening, but, um, I think it's happening in a different way.
Or, you know, it's, like, worth, like, seriously interrogating why is it that we're not just, like, automating all jobs right now. Um, I think my view it is something like RL, the way it's applied to LLMs right now is kind of a weird, funny tool where it doesn't really generalize beyond the training distribution that much.
Um, it generalizes to some extent, and it generalizes in interesting ways, but it's, like, very picky, right? Like, it, it kind of-- it can, it can kill the training distribution, like, completely. Um, it can be, like, best in the world at it, uh, with, like, not that much effort really.
But yeah, it doesn't really generalize. So I think what we had to do is bring the world of economically useful tasks in distribution for RL-
Mm-hmm
... if, if we commit to using RL as a tool.
Mm-hmm.
And, you know, it might be the case that maybe there's some, like, cool continual learning thing or something that, like, shifts the paradigm next year or something like that.
Mm-hmm.
But, um, it really feels like if RL is a tool, then, um, yeah, a big thing that needs to happen is like it's not-- it doesn't feel like intelligence of the models is the bottleneck. It's more like you just have products that bring the entire context of what someone wants to do into the product so that the LLM can, like, see it, and then use the RL on top of that.
Yeah.
Um, yeah.
Have you seen GDP Eval?
Co-design13:33
Uh, yeah, I've seen it. Yeah, yeah.
Is that basically what you're envisioning?
Um, yeah. I haven't, I haven't looked at GDP Eval closely. Uh, actually, haven't seen exactly. Like, uh, like-
So-
... roughly, yeah
... to recap-
Yeah, yeah
... it's 128 tasks-
Mm-hmm
... across, like, any white collar job-
Mm-hmm
... that takes, like, more than 5% of GDP.
Mm-hmm, mm-hmm.
Right? And, uh, they, they, they basically cr- created all the context-
Mm-hmm
... the eval on it and, uh, evaluated every model. Like-
Yeah
... famously OpenAI, OpenAI's evals team, whoever runs that one-
Yeah
... always finds the Anthropic site the best one.
Yes. Yeah, yeah. It's, uh, yeah. Props to them for, uh-
They even publish it, yeah
... doing that, yeah. Yeah.
It's actual science.
I think it's good. Um-
But, but, like, I think, like-
Yeah
... uh, in a t- in a, in a sense of, like, generalizing beyond coding competitions to economically useful tasks-
Yeah
... that is it.
I, I, I think that is a-
What is more important for GG6?
Yeah. What I'd like to do is kind of like I haven't-- I just haven't read the, like, GDP Eval traces close- closely. Um, it's not clear to me that, you know, like, what, like, what does the job of an accountant entail, and, like, what kind of context needs to be in the product-
Yeah
... so, so you can actually do it.
They have, like, PDFs.
I see, I see.
Like, so they, they try to go as close to source documents as possible.
I see. I see.
Yeah.
Yeah. So yeah, I think, like, roughly operating in this kind of thing is what I envision.
Yeah.
But-
Because I think it can't, it can't be, like, an artificial like, "Oh, let me clean up this data for you-
Exactly. Yeah
... to make it easy for the LLM to process." No.
Yeah.
PDF in an agent, go.
Uh, yeah, I think that's sort of, like, roughly the right shape of the thing, and I guess how I imagine this being, um, you know, operationalized is that you'd want to, like, co-design the product and the model so that, like, the product for whatever it is, like, I mean, coding is kind of maybe the easiest first step because most of the context that you care about is just your code base and, like, being able to run stuff in the terminal and that kind of stuff.
And still, like, we're not that close really to automating it necessarily. But, you know, for, for, like, all the other jobs, the context is, like, insane, right? It's, like, all the conversations you had with your coworkers, like, your Slack messages.
You know, like, for my, um, so at OpenAI, I was working on kind of like hyperparameter scaling research, and I actually wrote not that much code.
Like, like grid search or, uh-
Um-
... neural architecture search or-
Uh, no, more like, um, understanding how diff- like, um, science of deep learning in, like, 2020 where it's like, oh, you have to, like, initialize the layers in a particular way to get good scaling law, kind of the analog for that for RL.
Okay.
Um, the thing is, like, I didn't write a ton of code, so the LLM, you know, like, writing code is not the, uh, bottleneck, but it's more like, you know, over the course of a year, I, like, run sweeps, look at, like, the interaction between different hyperparameters, and kind of build up that knowledge for, like, a year of, like, just different graphs.
Mm-hmm.
And to do my job, the model would also need all those things in context, you know, to, like, successfully, like, you know, kind of automate my job, and you'd kind of want a product that allows you to, like, bring all that context in.
Mm. Did you have to build it for yourself, or-
Uh-
... was there an existing one?
Uh, no. I mean, like, uh, no, I mean, I, like I, you know, those graphs are just sitting in my head.
Yeah.
Right? So, um, I think it's-- it would be pretty hard to, like, go automate that job. But I think what you need to do is build a product that kind of, yeah, has-- brings that context in, and then you want to RL on top of that to, you know, understand, like, to teach the models to, like, use that, um, that context.
Yeah. Another conversation that I think has really come to a phase this year is kind of the, the death of one model fits all.
One Model?16:48
Mm-hmm.
I feel like the point of the G in AGI is, like, one model fits all.
Mm-hmm.
I think OpenAI has, like, clearly abandoned that this year.
Oh, where do you see that?
Fiji Simo writing a blog post with the title that we are no longer doing one model fits all.
Oh, okay. Interesting. Okay. Uh-
Uh, and I think Mark Chen or one of the other senior people that are not Sam also saying this in a, as-
Mm-hmm
... in a podcast. Uh, so, so basically, like, the idea was you, you started with Codex, someone else was doing Instruct GPT, then we launched GPT-4, 4o-
Mm-hmm
... I guess, o1.
Mm-hmm.
And o1 was kind of a supposed to be, like, a reasoning one model fits all.
Mm-hmm.
And then we merged the 4o and o1, o3 line into 5.
Mm-hmm, mm-hmm.
And now we're splitting it out into 5 and 5 Codex again. It's like it's just a weird-
Well, uh, OpenAI is very guilty-- I mean, you know, I, I don't think you should interpret those as, like, scientific facts about the universe. It's just more like, um, OpenAI has a tendency to shift the org chart, basically.
Yeah.
Uh, right?
The world has the tendency.
Yeah, exactly. So I think, um, a lot of it is related to that. But yeah, I see what you mean by, like, um, yeah, actually I do, I do wonder if, um, yeah, like the current reasoning paradigm, like the current reasoning paradigm is just kind of fitting itself to this kind of peaky in certain areas thing, right?
Yeah.
I don't think it's so much a matter of, like, model capacity, though. It's just more of a, another kind of organizational thing that, like- If you care really, like a lot about coding, you probably don't have the data to do all the other stuff.
I don't think it's so much a matter of like if, if you had all the data, probably you would benefit from just like training on all of it, and you'll get some generalization between these. But it's hard to find like one organization that cares about all these at once.
Yeah, yeah.
Yeah.
So before I double-click on just like the o-series in, in OpenAI-
The Blip18:37
Mm-hmm
... uh, I do like to p- ask OpenAI people-
Mm-hmm
... who are there, uh, do you have a favorite Blip story?
Yeah, the, the Blip was crazy for me. Like, uh- Yeah, I was, um ... It was like Thanksgiving.
Like everyone remembers where they were, what they were wearing.
Yeah, yeah, exactly. Uh, I was, um, at Thanksgiving with, um, two OpenAI friends actually, and then one of them on like Friday, uh, afternoon is like, "Oh, like Sam Altman just got fired." We were just like co-working together.
He's like-
I'm like, "What? Oh, ha ha, like good joke." And then, yeah, no, it was crazy. And then, yeah, it was, it was just like a crazy weekend of just like ups and downs. Like, you know, we thought, yeah, like, uh, like Sam came back.
You, you signed, you signed a letter, um-
Um, yeah, I did
... where you say 95% of people signed it.
Yeah, yeah, yeah. Um, yeah, I thought, you know, like I, I-
Are you going to Microsoft or
Well, I think maybe I had a slightly more compli- like I actually do think that governance feels really important to me.
Yeah.
Uh, because it does feel like no matter if we hit AGI in like two years or 10 or whatever, it's not clear that we have a good structure for the governance of it.
Okay.
And so it, it is a question that I think we like probably should spend more time on, and I was like, during that period, just pretty willing to be like, "You know what? Like, let's forget about the like equity and stuff."
Like, you know, I think it's like good and healthy to have a conversation about like how exactly the governance should work.
Okay. You care about this, yeah?
Uh-huh.
Right. So now the OpenAI nonprofit has this like secret shadow board of members that determine when, when we've reached AGI.
Y- yeah. Yeah.
Is that better?
Um, yeah, I, I don't have a like, like maybe ... I, I would say I don't have an answer.
Yeah.
Like, you know, like it's just, it's not big-
It's above my pay grade, but like-
Yeah, and w- even, even, even back then, I was kind of like, "Well"-
I don't care
... I, I do care quite a lot.
Yeah.
Um, when the Blip happened, one of my reactions was like, well, you know, this nonprofit board stuff, like actually if it takes such somewhat like surprising, maybe erratic actions, like maybe you'd rather just have like, you know, a thing like the Microsoft board, which is kind of like, you know, like probably like all the pensions of the world.
Like serious people.
But wait, like sure, like serious people, but also like, you know, the stakeholders are kind of like the whole world because everyone's kind of, you know, through their pensions or something invested in it. Like maybe that is a bit more of a democratic way to run things than having like seven people, uh, run it.
But yeah, I don't, I don't really know. It feels like we haven't solved governance like at all though, right? Like if forget AI, um, even stuff like unhealthy food or like social media, it kind of feels like, like whatever the kind of like capitalistic incentive is, like doesn't actually like, uh, capture kind of good outcomes for society maybe.
o1 Origins21:22
Yeah. So about like the, the transition into reasoning, right?
Mm-hmm.
Um, you shocked me by when, by mentioning that the reasoning team is 300 people.
Uh, I, it's, uh-
If anywhere you draw the line.
It's, it's kind of like, you know, now that, um, like, you know, when o3 was kind of showed as a product, like I think it just like kind of gets like larger and larger how many people worked on it, so.
Yeah.
Yeah. Uh, I think I've like, like lost track of the numbers, but yeah, like a lot of people contribute to the different aspects of like safety and whatever, eval, those kinds of things.
Yeah, so like original o1, like I saw the video.
Yeah.
It's like a dozen people, you know, like-
Uh, yeah, e- even then, like if you look at all the contributors, it was probably more like 50 to 100 people.
Okay.
Um, yeah.
So, so I mean, like let's, let's tell that story from your point of view.
Mm-hmm.
Uh, figuring out what does RL mean there.
Mm-hmm.
And I guess was this a branch of any other prior work that you want to credit?
Yeah. Um, so I think, um, yeah, like setting the scene, I guess, you know, in like 2023, people were kind of talking about, oh, like, is, uh, are scaling, scaling laws dead, this kind of stuff. Um-
Every year, every NeurIPS.
Yeah, yeah. But especially- I think especially that year it f- it felt pretty like serious, you know? Yeah, I think, um, in general, OpenAI is really good about like having conviction in something and just like really like from first principles, like going after it, and I think like the people who are kind of most responsible for that is probably like, like Ilya Sutskever and Jakob Pachocki.
I think e- even like, uh, like, uh, Dota was kind of more or less the same template in some ways, right? And that was, um, 2017. And so a lot of the people there have kind of this like AGI in their bones kind of point of view.
Um, and they've basically been convinced that like RL would be the way to get there. So I think for a long, long time people have been convinced that something like that should work, and it's just that it started to work once the, um-
Human feedback
... like kind of pre-training got good enough.
Okay. Yeah, yeah.
Yeah. Uh, I think human feedback is kind of like a, a bit of like a side branch because-
Yeah
... you can't really pour that much compute into it, right? It's like you, you take the model, and you like elicit it to be a little bit better, uh, in terms of personality. But like the people there are really convinced that like at some point, you know, it's not about copying the internet.
Like you can go, um, yeah, do RL, and, uh, you, like, you know, that's like the path to g- like getting much better intelligence. So I think it, it, there was kind of like a long line of k- kind of like returning to RL in for, in like different ways.
Um, and then it's just that around like, yeah, 2023 is when it started like really clicking, and it was kind of interesting 'cause even, you know, it's, it's not like those initial models performed like, um, way better than the existing models 'cause they're like s- smaller scale.
But people were very good at being like, "Oh, like this is kind of interesting." Like, you know, the, the reasoning trace that you see here is kind of not something that you've really seen be so accurate, um, in other models like this, like this one.
Kind of similar to how I think a lot of people didn't really think of GPT or GPT-2 as something that was like super Compelling probably. I know that I personally didn't like, uh-
GPT-2?
... that much of GPT-2.
Yeah.
I was like, "Okay, whatever." And then I think-- And then GPT-3 happened, I'm like, "Oh, whoa," like, I feel a lot of FOMO sitting in my PhD. It's kind of that where I think it takes a bit of, like, first principle's conviction to, like, um, yeah, decide that, like, oh, this, this thing, like, there's something here, and we should really go scale it up.
And OpenAI's really good about once you decide that something is good, then you just, like, scale it up all the way.
Yeah, is-- Was there an internal prototype pre-o1-
Mm
... that was like, "Okay, this is the thing. We'll fund it to scale it up," right? Like, there, there usually is.
Uh, yeah, yeah, exactly. Yeah, like-
What was the thing?
Uh-
What, what was the demo that, like, really sort of sold-
J-just this, like, you know, like, uh, like, running RL on even, like, a pretty small model producing, like, very interesting reasoning traces and, like, getting, like, uh, surprisingly good scores on math.
Yeah.
In a way that we couldn't have done without, like, a bunch more pre-training. Um, and then, you know, once, once that looked good, then, you know, more and more resources went to just, like, scaling up that new, like, uh, law.
Yeah.
Um, and, you know, things like adding, um, tool use and this kind of stuff.
Yeah.
Um, yeah.
Did-- I think a lot of people make a lot of headlines on the, the large models.
Mm-hmm.
But I think a lot, uh, it's very underappreciated, the minis.
Mm-hmm.
Uh, how well-
Mm
... this solution works.
Mm.
Any comments or just, like, discoveries on...?
Yeah. No-nothing much to say there.
Yeah.
I was also, like, not super involved in-
Yeah
... the mini stuff. Um, I think maybe one thing, not, not exactly related to that, but, like, it seems like, um, externally, people are kind of very, like, oh, uh, like, research seems to come in these, like, big leaps.
Okay.
Uh, but I think internally at OpenAI, it feels very smooth.
Like, you have a bunch of experiments.
Yeah.
Some of them have inconclusive results, but maybe you stack them.
Yeah, exactly. You stack them, and just, like, you keep scaling, you keep, like, having, like, different runs that, you know, get a little better each time.
Okay.
Um, so I think that's maybe one other aspect that's, like, a little underappreciated is that, like, I don't know. Like, in the media, there's just these wild swings between, like, oh-
It's so like-
... Google is wearing it. Yeah, exactly. And I think, like, internally at Big Labs, it's just kind of like, oh, we're just, like, chugging along. Like, maybe this month is a little better than last month or something, but it's, like, not as, um, crazy up and down.
I think the question is, there used to be more of this, and now I know there's less.
Mm-hmm.
Which is, well, the stuff we've released, we're like, you know, internally, we're like six months ahead of-
Mm.
Like you say, part of the reason why people, uh-- ChatGPT-- OpenAI wasn't that excited about ChatGPT's launch was 'cause they already had GPT-4.
Mm-hmm.
They're like, "Oh, we'll just, like, put this out."
Mm.
Like, it's we-we're already way ahead.
Mm-hmm.
I think now people are just releasing things as they have them. Like yeah.
I think, yeah, especially 'cause there's some, like, competitive pressure, right?
Yeah, yeah.
Uh, I think people are probably pretty worried that, like, if you, if you let a lead linger for too long, that'll, like, grab a lot of market share. Like, I don't know, like-
I, I-
... Nano Banana Pro right now is probably, like, you know, it's like it's pretty good.
They improve it every month.
Yeah.
So I would say, like, now the lead, internal to external lead time is about one to two months.
Mm. Yeah, yeah.
Which is-
Exactly
... tiny.
Pretty, pretty sure, yeah.
Tiny.
Yeah.
Anything else on reason- on reasoning side? I guess you can talk about on, on, on, so say the work on coding, um, anything surprise you or, like, is an external misconception on o1, o3 side-
Um-
... before we go to Cursor?
Well, um, not really. Like, uh, yeah, I mean, uh, it's, yeah, pretty cool. Like, uh, I, I think, you know, it felt already by, like, maybe early twenty twenty-four like, "Oh, wow, like, this recipe, like, really works, and we can see how far we take it."
And, um, so I think, um, you know, it was, like, very steady progress and, you know, by that point, it was probably pretty pred-uh, predictable that we could, like, you know, really, like, uh, smash, like, you know, things like IMO or IOI.
Yeah. One, one funny thing that kinda happened is, um, while this was happening, uh, I went to this conference called The Curve-
Yeah
... um, which is about, like, kind of AI progress.
And Joseph Gordon-Levitt, like ran it.
Yes. I went last year. Um-
Yeah, yeah
... uh, this was before the o1 stuff was released.
Yeah.
And I, like, went to this thing where people were kind of making bets on, um, where we'd be on, um, epoch AIs, like, um, the, the math, the epoch A math ex-exam and, like, humanities last exam and stuff like that.
And, um, their estimates were like, oh, we'll be at, like, ten, twenty percent in, like, twenty twenty-seven. And I think at the time, there was, like, you know, models internally that were, like, already better than their estimates, so they're like-- it's, like, off by, like, you know, two years or something.
And the interesting thing is, like, those are also people who are kind of, like, you know, predicting that there'd be, like, Dyson spheres by, like, twenty thirty-five or something.
Okay. So-
So, so, so, like, their, their current estimate is, is way under-
Yeah, they're too pessimistic in the short term, too optimistic in the long term.
Um, yeah, well, I don't know if the-- like, I, I-
Yeah, yeah
... I didn't-- There might be Dyson spheres by twenty thirty-five. Like, I, and I don't, I don't really know.
Yeah.
Uh, but, uh, I think that, that is, like, one interesting aspect, um, is that, yeah, I think people still seem pretty miscalibrated in different ways. Uh, I, I do really appreciate how that community, uh, makes predictions though. Like, um-
Yeah
... 'cause I think most of the rest of the world just kind of, like, cynically says, like, "Oh, I, I saw this the whole time." Like-
Yeah.
Um-
And so-
... I do appreciate that
... is this EA adjacent?
Uh, yeah, I think, I think it's-
That strong-
Yeah, exactly. It's like-
That, that, that group
... it's like, it's like that group. Yeah, yeah.
Yeah, yeah.
Um-
I like that they re- they like to sort of register their opinions-
Yeah
... a-ahead of time, and then they-
And, and I think, like, broadly, uh, the people who've been, you know, uh, the capabilities predictions in that group have been broadly correct if you look, you know, from, like, twenty fifteen to twenty twenty or something, like where I think a lot of people kind of thought that AI was, like, a sham or, like, you know, not really gonna be that useful for a long time.
And actually, you know, it is-- it's somewhere in the, like, twenty thirty-ish thing that, like, it will probably reach, like, human level intelligence.
Yeah. It's weird. So, like, I, I, I, I feel like a skeptic when I keep saying, like, everyone always predicts that AGI happens in their lifetime.
Mm-hmm.
That's very convenient-
Mm-hmm
... for whoever. And, like, we have a consistent view of history where you make-- see, like, people in the eighteen hundreds and nineteen hundreds making predictions.
Mm-hmm.
It somehow always lands in their lifetime, whatever the, the thing is.
Yeah.
But, like, this time it might happen.
Almost surely, right? Like-
But no.
I'm, I'm pretty sure.
Yeah.
Yeah.
Uh, yeah. So, so i-it's a, it's an interesting observation, like how different are we-
Mm-hmm
... from our predecessors-
Mm
... in, in terms of, uh, developing our technology.
Yeah.
Um, did the DeepSeek moment this year, also this year-
Mm-hmm
... crazy Uh, change anything internally?
Uh, not really. Yeah, I think that was-- I think mo- more so just, like, surprised that, um, it created such a moment. Like, it was kinda confusing, right? It was like DeepSeek shows that Nvidia chips are actually more useful than previously thought, and, like, Nvidia's stock, like, goes down a bunch.
Like, it, it was kind of like a-
I think it's m- more like, okay, well I'll, I'll, I'll do the steelman-
Yeah, yeah
... that side, which is, well, you don't need the top-of-the-line Nvidias. You can just use-
Mm
... the, the sort of previous generation or the shackled ones they sell to China to do an equivalent amount of work, uh, for, for a-
I see
... recent model.
I see. Yeah, but then it was als- uh, I guess the feeling at OpenAI is that, like, well, I think we, we had a better model already at the time, right? So, um... And it was quite valuable. Like, uh, like smarter models were clearly quite valuable, so you kind of wanted to be at the frontier.
Okay. I, so I wasn't quite framing this as like a race-
Yeah
... uh, dynamics thing between labs. It was just also more like, well, were they right? Were, were their approaches right?
Mm.
They had R10, which is kind of like a really cool branch.
Mm-hmm.
So more like commentary on what we learned about RL this year in particular.
Yeah. Yeah. Well, it does seem like basically, um, a lot of the labs have kind of like converged onto some similar-ish way of doing RL, and they're all kind of back at the same level of, like, frontier again.
Like, even the Anthropic, uh, models, like the, uh, Opus t- 4.5, it has this kind of like, uh-- There's this, like, RKGI-2 plot that looks exactly like the OpenAI ones, right?
What?
Like, so I think everyone seems to be converging on a pretty similar, um, form of RL. Um, yeah, it's kind of interesting. I think people basically figured out in one way or another to, like, achieve more or less the same thing.
Yeah.
Yeah.
Cursor Move32:28
Let's talk about the move to Cursor.
Yeah.
Why is Cursor accumulating and drawing so many cool RL people?
Yeah. Um, yeah. So I'd actually kind of like already, like, uh, talked about this so far, I guess. Like, so yeah, I think from the perspective of Cursor, it's like, you know, nice not to be so like, uh, dependent on, like, external labs for everything.
And like, I think there's also, like, um, unique opportunities to co-design the product with the model in ways that I think we couldn't do unless we actually, you know, built the model ourselves and, like, had access to, yeah, making it good.
Yeah.
Um, so yeah, that's kind of like, uh, broadly why Cursor's so excited. Um-
Okay. I'll, I'll push back a little bit, right?
Uh-huh.
Uh, OpenAI is, has infinity resources.
Mm-hmm.
Uh, infinity data, has Codex. Uh, you could have just stayed.
Yeah, yeah. Well, uh, actually right around when I was, um, leaving is when, like, I think people started actually, like, using Codex, uh, a lot. So that was kind of like a-- It like, it like happened right after I left, so that was kind of funny.
Like, yeah.
So, so mostly people are using Cursor internally.
Mm.
Maybe a bit of Windsurf because-
Yeah
... it was left over from the previous thing.
Sure. Yeah, yeah, exactly. So it, it wasn't that obvious. But, um, actually I think more to the point, um, this thing I was saying about, like RL is kind of a tool that doesn't really generalize that well. So what you wanna do is bring the entire like, um, kind of test distribution inside your training distribution.
I saw the opportunity to do that at Cursor kind of like directly, and I think the Cursor folks are also just like really excited about that kind of vision. And it's just like a small place where, you know, like the product people sit like right next to the ML people, and I think there's a lot of potential there.
Um, you can kind of see that, um, recently, uh, Jakob Jackson had this blog post about, um, like online tab where, uh, like, you know, we're doing policy-
Cursor updates every two hours.
Exactly. Like, like a policy update every two, two hours or something. And I think that's the type of thing that, you know, I think it's like a little hard to do-- Uh, it's like very hard to imagine that at OpenAI, for example, just 'cause like, you know, it-- the product is this like kind of complicated thing and also like the product people and RL people are pretty like, you know, on like different sides of the org.
I, I think if you put your mind to it, you would. It's like, you know, tab is an autocomplete. It's a smaller model. It's, you know, it's not-
Yeah
... as complex, I guess, as-
Um
... below them.
Yeah, but I don't think that's really this like, you know, I think we-- I don't think that's why Cursor was able to do it. It's actually more about like just the org itself being kind of like smaller and a bit more like focused.
Yeah. Yeah. Okay. Well, I mean, since you're indulging this-
Mm
... uh, I think the question about continual learning, which obviously is a big theme-
Continual Learning34:56
Mm
... it's always been a big theme, is bigger this year, is, well, don't you need to curate your data? You can't just like chuck whatever your users are doing in, straight in-
Mm
... because that tends to get you towards the middle of the distribution. They actually want to spike it.
Mm. I guess it depends how you're thinking about continual learning. I mean, I don't know, like, like humans are quite good about dealing with bad data too, right? Uh, like you can see something-- like you can see someone doing something dumb and decide, like you're not gonna do it.
Filter it out. Yeah.
Yeah. But, but like it's not even actually filtered out. Like you have, you know, presumably some kind of value function that like says that if you see someone touch a hot stove, like you're not gonna go, you don't need to-- Like it's not just filtering it out, you're actually not gonna do it, right?
Um-
You could rediscover hot stoves on purpose.
Yeah, but like you don't, you don't need to. So, um, I think there's something pretty deep there. Um, yeah, like it seems like we're kind of like a few orders of magnitude of like kind of data efficiency, basically, away from like that kind of like, you know, you, you, you do something once or like you, you make a mistake, uh, like you, yeah, you, you introduce like a bug in your code.
You're not gonna do it again. Uh, but the models will happily just like keep doing it, um, even within the same context, but definitely, you know, of course acro- across context.
Yeah.
Um, so I think there's something like interesting and deep there is like maybe, yeah, I suspect that it'll be kind of like paradigm shifting in the next like year or something, but I have no idea, like, you know, what it might be.
Um, yeah.
So is-- Primarily you worked on Composer-
Mm-hmm
... Tab, and maybe Search?
So I, I have-- I've, I've actually just worked on Composer.
Yeah.
Um, and that's kind of like the main focus of the company basically.
Okay.
Uh, or like the ML group, um, is-
Which is-
... shipping a better-
Yeah
... um, shipping a better Composer.
Can you describe, I guess, the impressive, uh, brag a bit about the ML group?
Yeah, yeah. I mean, I, I think the ML group is great. Um, it's like, uh- You know, it's just like twenty, twenty-five people and, um, you know, I was like honestly like pleasantly like very, very surprised at like how good Composer is, like given the size of the group and, you know, it's not like a big research lab yet.
And, um, uh, yeah, it's-- I think it's like a really good model. You can kind of see that in the reception, and I think it's kind of the start of hints of like co-design with the product in some ways 'cause I think one of the reasons that people really like it is it's smart enough, um, that peop- that you actually wanna use it.
Um, and it's also fast, so you kind of like stay in the loop with the model while you use it 'cause I think all the other smart models have this kind of-- they're, they're slow that you wanna go kind of context switch away and come back, and that sucks, you know?
Like, uh, just as like a, like programmer, it just sucks to kind of context switch. It kind of like gives you ADHD. Like, it, it's like really terrible.
Yeah.
Um...
I agree.
And I think, uh, yeah, it's like one step in the direction of like being able to be more sync, and I think that's-- like basically the whole company is just really, you know, full of people who want to, you know, code, even like the co-founders, you know, uh, like actually, uh, the co-founders are often some of the best like, like high taste testers, which also kinda gives you a lot of like reassurance that you're gonna ship good stuff.
So yeah.
Any example test that like maybe Composer doesn't solve yet but you're really motivated to solve?
Well, yeah, ironically, uh, I feel like I'm actually like a low taste tester in some ways. Because I don't know, like, you know, I just like write like slow like machine learning code and just like think about, um, algorithms and stuff all day.
Yeah.
Um, I think more broadly, I'm super excited about co-designing the product so that you can actually, you know, not just-- Right now we're getting better and better at like answering user prompts, um, and I think that's why Composer One is like quite good.
But, uh, you know, what we're really aiming for is like more like, you know, automate software engineering as a process where you like write code, you go look at Datadog, uh, look at what's like happening, then come back and like, you know, maybe have some hypotheses about what's better, like re-rerun stuff.
I think that's the type of thing that we actually want to make the model do.
Hmm.
And I do think that Cursor is kind of like uniquely positioned to do that in the sense of like, you know, if we can kind of-- if, if a lot of what a software engineer does kinda ends up in the product, um, I think we can use that to like get better and better at, you know, not just writing code, but kind of like the whole job.
Yeah. I think that's very inspiring. Just to double-click on just any sort of, uh, RL insights, uh, y- Sasha and Lee have talked a lot about like the internal tooling that you've had-
Mm-hmm
... for all the like the cluster visualizations.
Mm-hmm.
Is that helpful? Is that what every lab has?
Yeah, I think, um, the tooling at Cursor is actually really good, um, I think because, you know, it's just kind of like a-- people are just down to like vibe code stuff. They like do-
Of course
... test their own stuff. Like, um, so we just have like a lot of good tooling where you can, you know, like have like a SSH session into like, um, our own like, um, user environment or something and like, you know, see if like, uh, code runs the way that like u-users got it to run, like this kind of thing.
I think that's actually, yeah, quite nice. I think basically one of the big lessons in ML in general is that you wanna be like really close to your data and understand your data well, and, um, yeah, I think Cursor's like kind of, yeah, again, kind of like uniquely positioned to do that well, especially-
It's all internal tooling. You're not buying anything.
Yeah, it's just like internal. Um, and part of it is just that we're also working on a product where you can understand it really well because it's a code product. Well, like, you know, if-- I don't know, in, in, in, in OpenAI, if I was like to look at like a biology question, I have no idea, like, you know, what, what this is about.
Yeah. Yeah. Yeah. Interesting. Okay. So I think that's a good overview of, of everything. I guess other than the-- we covered OpenAI and Cursor, just interesting RL work that other people are doing that you're, that you're like still mulling over, it's influential to your thinking, good papers, anything like that.
RL Future40:39
Yeah, you know, um, uh, unfortunately, I've like kind of gotten the habit, especially at OpenAI, of like not reading that much external work and just like reading people's like Slack posts internally -
Nice
... as like the main like way to like, you know, um, uh, like learn new stuff. Um, no super inspiring recent things have popped up to me. Um, I do think that this like kind of vibe of like, yeah, continual learning just like does feel like, uh, I think there's something super interesting there, and like it feels like, uh, maybe even in academia people could make like a pr- big crack at it.
And continual learning specifically meaning kind of what TAB is doing?
Yeah, maybe what TAB is doing, but also just like kind of like in context learning but with like infinite memory or something so that you don't-- Once you experience something in context, it should just like be in your weights, and you shouldn't have to like-
Yeah
... make that same mistake again, that kind of thing.
Why do you think there's-- Okay, but y- so it sh- it should be in your weights, but there, there's a finite capacity for the weights to remember things.
Yeah. Yeah.
You will forget things, uh, if you do that too much, right?
Not r- I mean, you know, you, you start out by memorizing or, you know, like learning from trillions of tokens.
Yeah.
Now you're gonna experience like thousands or maybe millions of tokens, and somehow, like, you know, we can-- and those thou- the million tokens are kind of in deployment.
And it's on- you only need one epoch.
Yeah, exactly. So-
Crazy.
Yeah, so, uh, it feels like if you, if you could learn enough about those million tokens that you're actually in deployment on, um, I don't think you should need-- like I don't think there's a risk of overloading the capacity of your model, right?
'Cause y- you can train on a trillion tokens, and it's like fine.
Right. Right. So there's proportionately it's a drop in the water.
Yeah, exactly.
Water in a bucket.
Yeah.
Unless you run it for years and, you know, at some point it start-
Maybe, yeah.
Yeah. So basically, I, I, I find it very curious. I've only had one podcast on information theory of language models.
Mm-hmm.
Like, what is the theoretical capacity? How much are we using?
Mm-hmm.
And you should probably track that.
Yeah. Yeah. That's a good idea, yeah.
Like treat the, like the weights.
Yeah.
If you want to store things in weights, okay.
Yeah.
Treat it as a hard drive. What's the capacity of the hard drive? How much can be stored in there?
Yeah.
We know, we know the capacity. It is the number of bits that, you know, th-this-
Yeah
... occupied by the language-- by, by the parameters.
Yeah. Yeah.
Physically cannot store more than that.
Yeah. Yeah. Yeah. And it's-- Yeah. I've, I've heard that there's this kind of like someone recently at Cursor, Jacob, kind of brought up this view. I don't know if it's like a more public view that's like, oh, th- there's kind of like a hard drive view of, um- You know, uh, neural networks and kind of like a CPU view of neural networks where, you know, is, is what's happening the weights?
Yeah.
Like, yeah, memorizing stuff or is it like you're, like, having some, like, few circuits that, um, do a lot of work?
Yeah.
This kind of thing? And yeah, I don't know. Yeah. You know, I would love to, uh ... Yeah, there's like actually so many of these kind of more science-y questions that I would, like, love to explore sometime, but then it really kind of conflicts with, like, empirical stuff, you know?
Like, uh-
Mm
... unfortunately, at any given moment in time, it doesn't seem like the most, um, fruit for like, you know, improving something in the sh- especially in the short run, but even in the next, like, couple years is, like, understanding some of these questions.
Um, yeah, I mean, I guess this is technically supposed to be the role of academia, but it's, like, also hard to explore those ideas there without enough compute. Um, but yeah, actually, I would love to, like, go at some point, um, you know, like return to exploring these kind of, like, fundamental science ideas.
Okay. This is a ... I'm just kind of springing this on you, so you can take some time. What is a good RL interview question that if somebody can answer, they should join Cursor immediately?
Hiring Cursor44:00
Ooh, it's a hard, uh, question. Um-
I'm assuming you do interviews here.
Yeah, yeah. Um, well, actually, at Cursor, we do, like, work trials.
Yeah.
And it's, like, two-day work trials that I actually think that that's, like, more representative.
'Cause you plug in and-
Yeah
... you see how they behave.
Exactly. Um, so I actually think it's, like, more valuable. Um, this is honestly less of a thing about how you understand RL and a bit more like were you around in the, like, 2017 to '22 era. But, um, it's like why is off-policy RL unstable i- is kind of, I think, like, a, a good, uh, question to, like, yeah, dive into.
I don't actually know, so I'm going to have to dig into it.
Yeah.
Cool. Thank you. That's, that was great conversation. Uh, do you have any sort of call to action?
Yeah, I mean, uh, you know, we are definitely hiring at Cursor, so, um, yep, if you're interested in working on especially, like, kind of, uh, data and rewards for code, I think that that's, like, a huge need. Um, yeah, please, like, get in touch.
Um, uh, yeah.
That's it?
Yeah.
Thank you.
Sweet.





