Intro0:00
So METR stands for M-E-T-R. Uh, first two letters, model evaluation. That is, we think about what the capabilities of AI models might look like today and tomorrow, as well as their propensities, what they'll, what they'll actually do in the wild given that they have some level of capability.
And then threat research is the final two letters. We try to connect those, those capabilities and propensities to, uh, particular threat models that we have in order to determine whether AI models pose enormous or, or catastrophic risks to society.
So the secret, if you, if you read this article about how I became the number one most profitable trader on Manifold, mostly comes down to this one market where-
Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, founder of Kernel Labs, and I'm joined by Squeaks, editor of Latent Space.
Hello, hello. We're back in the studio with Joel Becker from METR. Welcome.
Thank you very much, guys. It's a, it's a great pleasure to be here.
So Joel, uh, your work has impacted the AI field a lot, especially over the last year. I invited you for the AIE, uh, Summit, which thank you for speaking as well and, and doing, doing the extra workshop. And there...
You have a lot of papers that have been very impactful, but I guess upfront, uh, a lot of people... Like, METR just burst onto the scene. Could you explain and introduce METR?
Yes. So METR stands for M-E-T-R. Uh, first two letters, model evaluation. That is, we think about, um, what the capabilities of AI models might look like, um, might look like today and tomorrow, as well as their propensities, what they'll, what they'll actually do in the wild, um, given that they have some level of capability.
And then threat research is the final two letters. We try to connect those, those capabilities and propensities to, uh, particular threat models that we have in order to determine whether AI models pose, um, enormous or, or catastrophic risks to society.
Yeah. Would you say that you've done a lot more ME and TR is, like, kind of the next phase? Or is there a TR side of work that I've missed-
You know, I think-
... conscious?
I think there's some TR. So, um- ... so, you know, some of the most publicized work does look more like the ME. Um, it looks like this, this time horizon stuff and, and the, um, developer productivity RCTs, stuff like that.
Um, but there's this, there's one full report on our website, GPT-5, um, report and, and analogous one for GPT-5.1 as well, trying to make this more sort of structured case that it doesn't pose these really large scale risks, you know, eventually coming to the conclusion that it, that it doesn't.
But, but it's worth thinking like w- why exactly is that the case? Like, if, you know, you and I work with GPT-5, it does seem very capable. That matches up to, to benchmark scores. You know, why, why is it not, um, able to do really so- something really, um, enormously wrong?
Um, well, you know, w- we, we go through the evidence. We find... We, we think it's not capable enough, you know, on, on the basis of some of this capabilities evidence that you've alluded to, to, to commit these catastrophic harms and so, you know, it's not, it's not going to be able to do this.
But perhaps in future, you know, we, we will think it's capable of doing, um, pretty extraordinary things. You know, the k- the kinds of things that would be necessary to, to provide really, really serious threats, and then maybe you'd lean more on the propensities parts.
Like, you know, are there protections that we have against, uh, these dangerous capabilities sufficient for it not to, not to pose an existential threat? Um, that sort of thing. So, so I think it's, um... I think threat research, you know, very much, very much is there, very much is something that we're, that we're aspiring towards.
In some ways you might, sorry, see the capabilities evidence as a kind of input for that.
Yeah.
Ha- the threat model's been updated a lot? Or- ... do you feel like you're still using the same threat models as GPD2 of like, you know, paperclip factory, blah, blah, blah, you know? But, like, how, how much are you ri- you know, increasing the bar?
Yeah. So I'm not an expert in, in the threat modeling piece.
Time Horizon3:33
Yeah, yeah.
More in the, more in the capabilities piece. Um, I, I do think they've been changing to, to some extent. So something like the autonomous replication threat model, that is being able to, uh, set yourself up and, and, and control resources, so- something like that, has been deprioritized relative to, um, R&D acceleration.
That is, you know, the possibility there could be some capabilities explosion inside of, inside of a lab, and that could be destabilizing for, for all sorts of reasons that we could talk about. Um, so, so mainly, mainly we're focusing on that latter one, although, although we do think about a, a, a number of threat models.
Yeah. Let's talk about the ME side, I guess. Uh- So I would say the, you know, model time horizon chart is probably the most quoted, I would say, both in, um, investment decks that I see and, uh, just general on, on Twitter.
What was the origin story of it, and, um, any other color you wanna give on it to introduce it to the audience?
Yeah. So there are a couple of different ways to tell this story. O- one way is there's this, um, uh, PowerPoint, uh, internal METR PowerPoint from 2023 where we're trying to lay out our ambitions for, for, for, for what METR research might look like in the future, and there's this graph.
It has a Y-axis that's like, you know, some measure of autonomous capabilities or dangerous capabilities or something like that, and then an X-axis that's labeled time or compute or, you know, what- whatever, whatever resources that we, that we want the, the Y-axis to vary over.
And then it has a bunch of sort of scattered points that kind of go, go up and to the right. You know, we think ca- capabilities are, are improving over time. Many of METR's research bets have been, uh, sort of try- trying to make this ever more concrete.
And then when we, you know, w- when we actually did the full thing, w- when we had something like this Y-axis, which turned out to be this task difficulty as measured by the length of time it takes for humans to do at which models can complete these tasks with 50% reliability, when we actually, uh, uh, got that data and plotted it over time, it turned out to be remarkably straight.
You know, as, as straight a, as straight as you're aware of from, from the, from the f- familiar graph. Part of what makes it so extraordinary is that this, this pattern does seem to be so regular. Um, in fact, it's just, you know, way more straight than this incredibly scattered graph that we, that we had at the beginning before, before my time, before I joined METR.
How did you pick the tasks? I would say that's one question that people have. You have some labels kind of like train classifier, fix bugs in small Python library. They all seem kind of arbitrary, you know? Like, what's the process of task selection?
People are right to be worried about, uh, task selection, um, or, or, or there are many, um, many finicky details in, in here. I would say the aspiration was to pick economically valuable tasks Relevant especially to sort of general autonomy and R&D, the, the, the threat models that we're primarily, primarily interested in.
So, you know, o-one misreading of the Time Horizon graph is this is referring to, you know, the full distribution of, like, a-any, any tasks that you might give AIs, and I think that's, that's clearly not right. And, you know, in particular, tasks that are requiring of vision capabilities, they're probably, um, to take one example, they're probably much less capable today, um, as measured by Time Horizon as, as for, for, for these tasks that are typically not requiring vision, vision capabilities that we give them.
So we, we try and, you know... W-we sample these tasks by having people inside of Meta create the tasks and by, um, having a bounty so that people from, uh, from outside of Meta can, can, can provide us with these tasks, stuff, stuff like this.
You know, that, that's not a sort of perfectly random selection process. In particular, it's a, it's a process that has a bunch of constraints. You know, in, in order to be able to, uh, scalably run our evals, it's helpful, not necessary, but helpful for the, for success on the tasks to be automatically gradable, and that means, you know, some, some types of tasks are included and tasks that are harder to make that happen for are, uh, are not included.
Um, but yeah, th-this is the aspiration.
Yeah, the computer vision point was interesting. Any other disqualifiers, so to speak? What are, like, other things where, like, you would expect the chart to be a lot worse at?
One thing is, um, fairness or, or we, we want s- we want tasks to be, in principle, completable by a model. You know, it has access to sufficient information. It's not, it's not impossible given the information it has.
The way we think about that is, um, you know, could a low-context human who was, uh, sufficiently skilled at sort of the general skills, but maybe, maybe not, uh, maybe not the particulars in the background, um, would, would they be able to achieve success on, on this task?
And I think that, that rules out a lot of, a lot of real work, um, because, you know, a lot of real work involves peoples... people having, um, careful mental models of the situation that are, that are not all sort of fully listed in an issue description or, or, or, or the equivalence, the equivalent of that.
In some ways, you might think of us as, as, as not measuring things like that. Another thing is that our tasks tend not to be... They vary a little bit, but they tend not to be so sort of open-ended or, like, interacting with the outside world or, you know, th-this sort of thing.
Messy as we, as we, as we call it internally, which refers to a bunch of different, um-
Right
... a bunch of different things. But, you know, you, you, you broadly get the picture from, from the descriptor messy. You know, relative to tasks that you might find in the real world, our tasks are, you know, somewhat nicely scoped.
They're, um, uh, they're, they're quite, they're quite neatly contained. You know, indeed, I, I think we're gonna talk about some of the developer productivity stuff later and, and, and some of the, some of the interesting findings, uh, you know, make more sense in light of the fact that those tasks are a lot more messy than the, than the Meta tasks.
Are there any that you would want to highlight in terms of task, uh, distribution? I think, uh, I've c- I've come across RE Bench before-
Yeah
... and you have a particular affiliat- affinity for RE Bench. Um- I don't, uh, for... I don't know if you want to introduce your, your side projects, RE, RE Bench Warmers?
I have a, I have a soccer team called the RE Bench Warmers. Um, we are the, you know, most enthusiastic and, and possibly least technically skilled- ... soccer team in San Francisco. We made the playoffs last season.
Ooh.
Shout out, shout out to the team for that. We're certainly gonna make the playoffs again this season, but, well, possibly by the time this, this podcast is out, we'll find out that we have not made the playoffs. That's the, that's the-
Is this the same one that you... same league that you're in or?
Sa-same organizer, but different field.
Okay. All right.
We play Mission Bay.
Okay.
We play Palomar.
Uh, yeah, and, but, like, uh, HCAST was, like, it was, like, the first time-
Yeah
... I, I come across it and, and the others. I, I, I mean, uh, SWA? Is that, is that the Meta proprietary ones? And then, you know, anything else that you're considering adding?
Yeah, so there are private tasks in, um, in HCAST as well. But yeah, S-SWA is this list of sort of atomic tasks or these kind of very small, um, uh, software actions. Uh, may-maybe one example is here's a list of four files.
One of them contains the passwords. One of them is called passwords.txt. Which file most likely contains the, contains the passwords? You know, I think GPT-2 can, like, sometimes do that task and, and, and sometimes not. Opus 4.5 I'm sure, I'm sure can do that task 100% of the time.
Then we go up to HCAST tasks, which span from, you know, only a little harder than, than those, than those SWA tasks all the way up to, you know, something like, uh, 20, 30 hours, which are requiring of more autonomy, more sort of, um, more sort of sequential actions.
Many of them are much more challenging. Perhaps in some sense they're built out of these atomic actions, although I'm not sure quite how clear that is. And then these RE Bench tasks are these very challenging, novel, um, machine learning research engineering, uh, challenges.
So totaling 170 tasks and, uh, I mean, I, I think, like, uh, this is very good. What, what, what's really interesting is I think the... people don't understand, like, when, when people quote, like, the number of hours, it is the human equivalent hours, but machines will probably take a lot less time for that.
You know, one thing I've always won- wondered was why didn't you publish a second chart where it was just like, well, here's the difference between what machines can do versus what humans can do?
That's a good question. I think you can kind of think of Time Horizon in some ways as, as, you know, a summary statistic, a single number for how good models are plotted, plotted over time. We could have done, um, how long the models can, can work for productively.
It's not quite clear how to operationalize that. Like, you know, you do want some notion of success, otherwise what, uh, how exactly do you, do you threshold this how long they can work for? Um, but, you know, but in principle, we, we could do, we could do something like that.
But this is the, uh, you know, this is closer to the, to the, to the first thing we tried. This is the thing with the, the clear empirical trends. Um, I do think it's right that a common misconception about Time Horizon is that it's about how long the models should work for, and the models are, you know, as, as we all see, working for longer periods of time autonomously, you know, in the wild when we, when we use them in, in a Cursor, a Claude code or a, or, or, or a Codex, but that's not the, the primary thing going on.
Opus 4.511:37
In some ways, I think it would be easier to explain Time Horizon if you assumed that the model solved all these challenges in like zero minutes or like five minutes or, or, or something. Just, just to emphasize, that's really not, really not the thing that's going on here.
You know, instead, we're just plotting what's the difficulty of tasks they can do over time, and that difficulty is measured in human time.
Yeah, I do think there's some collision when people say, "I ran Claude Code for five hours"-
Mm
... which is like, you know, the top of your chart right now.
Mm, mm.
But that would mean five hours of, like, a Claude Code run would be the equivalent of, like, a-
30
... fif- 30-hour task-
500
... in your-
Yeah
... thing, basically. And yeah. I, I, I think that's interesting to-
Or, or it might not.
Right. Yeah, yeah, yeah.
Because it might-
Totally. Exactly. Yeah, totally
... spend three hours doing absolute bullshit. Like
Yeah. And a lot of these claims about, um, uh, you know, um, Claude Code-
30 hours
... was running for 30 hours or something.
Yeah.
You know, I have a lot of questions about that. Like, how, you know, how, how good was that output really at the end?
Yeah.
Um, we, we have and we haven't to, to, to some degree. I c- I can talk about, talk about particulars. Um, th- there's also the question of, you know, if I attempted that again, like, how cherry-picked is this example?
W- if it succeeded the first time, would it fail, would it fail the second time? Um, I think in some ways those anecdotes are, are interesting, but, like, not so, not so scientific.
Yeah. That is something for people serious about AI to understand. The state of people making claims on agent performance is very unscientific and much more a- a- anecdotal, and sometimes influenced by marketing desires. Let's just put it kindly.
Yeah. You know, I think, um, METR's out there trying to, um, uh, support civil society, trying, trying to, um, provide high-quality independent information to the public. I, I, I can agree more that the information environment is, is, is less than perfect.
Let's talk about the four-- Opus 4.5.
Yep.
It's a very, very big jump. This was the first... I, I called it out when, when, when you guys put it out, where I was like, "Uh, this is the first time... As far as I understand, you're the first people to call out how much better Opus 4.5 was than the, than, than the status quo."
And I think, uh, this almost ties into your background as, like, a super forecaster- ... a little bit because, because then, basically over the entire holiday period over New Year's, people discovered that, what you ha- already discovered. Uh, what is your reflections on that?
What, what's your reactions? Any, any stories to tell about that?
That's very kind. I do wanna attack you on, on two claims. Firstly, I, I have not been a, not been a super forecaster. I think that's... There's a particular group of people- ... who work for TechLock or something that are supposed to be.
Okay, no, no. But you're broader, broader, broader.
Okay. And, and, um, um and-
What I'm referencing is your number one on-
And actually, you know, Opus 4.5 is a, is a big jump on benchmarks as well. I think in some ways, like, METR's time horizon is, you know, it's highly correlated with a bunch of, with a bunch of benchmark scores.
It's, in some ways, a kind of more understandable, um, a, a way of thinking about what benchmark performance really, really means, slightly, slightly more interpretable. Yeah, I, I, I do, um, I do feel intuitively like, um, like Opus 4.5 was, was a big bump.
Productivity RCTs14:27
I've seen some of the most talented engineers I know go from, you know, being picky about not using, um, not using AIs for coding to, to, to practically not, not write a, not writing a line of code. You know, I'm sure many other people at previous model releases have, have seen th- similar things happen to them.
I, I'm not sure that implies it's so, it's so discontinuous. You know, in some ways, I think the story of time horizon is that progress has been remarkably continuous over, over so many years, so many orders of magnitude of, of compute and effective compute.
But yeah, I, I, I think model capabilities are astonishing. It points to model capabilities being even more astonishing in, in, in future.
But I mean, it broke your trend line, you know? The, the trend line that you so hard... That you were so... working so hard to build over multiple years, and it just, you know, did that.
Yeah. I'm not, I'm not sure about the characterization. I, I think it's, um... So there was some speculation even when the paper came out that maybe the appropriate trend line to use is this faster four-month doubling time-
Yeah
... which Opus 4.5 would be, would be perfectly in line for.
And you picked seven months.
And picked seven months. You know, I, I was more of a believer in, in seven months. Um, a- and so it is kind of falsifying my, my, my trend line in si- in some way. There's also... It's, you know, it's, it's slightly confusing to think about, um, whether differences from the trend line represent sort of differences in the difficulty of our task distribution at particular points versus something sort of more fundamental, more like latent capability.
Um, and I don't feel like I have a, I have a perfect handle on that. You know, i- in general, I think the, the, the tweet sphere pays a lot of attention to, to particular model releases and, um, you know, really the informative thing is, like, over a period of a year, over a period of three years, uh, what, what the trends look like.
Well, I mean, you know, it was, it was a pretty significant update, I would say, for all of us. Um, and, and yeah, I mean, I would, I would cosign what you said there with even very sort of cynical, more senior developers being finally pilled into-
Yeah
... into agentic coding. And now very serious people are telling me that they want to commit their organizations to full "Don't write a single line of code by human hands," um, and just commit to 100% agentic coding, which is not something that you would've said a year ago.
Yeah. That sounds right to me. I, I feel it. I feel it in my own case.
How do you validate previous research? So, like, take the developer productivity study, right? You had AI slowed people down. If you were to redo it with Opus 4.5, would you expect the results to be dramatically different? And should we redo the study?
Like, should we stop coding the study? Like, how do you think about that?
We have been redoing it in the background. I think it's, um... And I, I won't comment on exact results, but, but I think it is, it is much harder to do it today than, than it was in the past for all sorts of reasons.
You know, the first is as AIs get better at coding, it's harder and harder to find, you know, developers submitting tasks who are willing to, to be randomized to, to AI disallowed. There's a, quote-unquote, "selection issue" where, you know, maybe we end up only observing the tasks that they, that they thought AI wouldn't greatly uplift them on, uh, ahead of time because they're, you know, those are the tasks that they're willing to be, to be paid for to, to, to be flipped into AI disallowed.
There are other issues like, I think today a common workflow is to work on multiple issues or multiple lines of work at the same time concurrently, and that wasn't, wasn't really true before. It's, it's difficult to know how to capture that in our, in our study design.
If you flip a single task to be AI allowed or AI disallowed, you know, you're sort of supposed to work on that single task, but actually that's not how developers are working today. I think basically these weren't threats to the previous study design or, or in, you know, approximately like March 2025, people weren't really working concurrently or, or not, not nearly to the same degree.
They basically were giving, giving us, giving us all of their issues. Yeah, I, I think that's an enormous challenge. I, I have some ideas about, about novel study designs, but, um, repeating the same one does seem tricky to me.
Yeah. We have Quentin Anthony, who was part of the study-
Yeah
... before.
The, the only productive developer.
The only, the only productive ...
I, I, I have some questions about that. I think, you know, um, Quentin is very talented as, as all of the developers-
Mm
... in the study are very talented, but, uh, we don't measure developer effects very precisely.
Yeah. No. Well, I'm curious- ... you know, um, you know, I don't know if he's part of the new study. You don't have to share that.
Yeah. Yeah, yeah.
But I think we'll... it'll be interesting to have people on again who've been in the study. I, I do feel like things are changing. Like, even three months ago, I was, like, using Cursor a lot more, like, in pair with Cloud Code.
Like, I think today it's like I do a lot of just Async Cloud Code and then review and iterate and... I don't know, man. It's much better. And I don't know how to quantify. You know? I think, like, that's part of some of your points before.
It's like people maybe overestimate. Like, if you were to ask me how much did it speed you up, you know, it's like, I don't know, 10X. But it's, like, probably not, right? But, like, I don't know how to calculate the actual percentage, so it's hard for everybody involved.
Y- yeah. So he- here's some issues you might think about. You know, if you, if you took the tasks that you were completing personally in, in March 2025 and then submitted them to our, to our Uplift study now under the, under the previous design, you know, we might reason about how much faster those would go.
You, you might expect them to go, to go somewhat faster because AI capabilities have improved. But you're doing a sort of different and larger set of tasks now. Like, I can think of a couple of side projects that I have that, um, you know, I simply won't be doing-
Right. Yeah
... we- we- were it not for AI existing. And, you know, in so- in, in some sense, the, the speed up there is, like, maybe infinite because these are, these are things that I simply could not have done otherwise.
B- but if you were to equate speed up with, like, the additional value that these projects are providing, these, these wouldn't really line up. You know? There, there's a reason I wasn't getting the expertise to do the other projects before.
It's just, like, it, it, it's less, it's less valuable to me. Another problem is the, is the concu- concurrency thing that we just raised. Yeah, I do think that very bullish estimates of speed up today are, you know, to some extent inflated by, um, by what we document in that original paper, that people's expectations of speed up tend to be too optimistic, it seems.
They also tend to be inflated, I think, by, by not quite grokking that the value of the additional tasks that they're able to complete are lower value than you might think. You know? That there's a reason that they weren't doing them previously.
That said, I don't doubt that, um, that those tasks do have value, that people are being sped up on even the tasks they would've done before. Um, it's a complicated issue.
Yeah. I do think that a lot of companies have issues absorbing additional productivity, especially when you're, like, a real product organization. Like, if you think of the AWS console, right? It's like if you gave AWS AI and everybody's 10X more productive, even if they ship 50,000 more services, you know?
Like-
Right
... customers can't really absorb 50,000 more server. So I think there's some... You know, you shouldn't really expect your engineers to do 10X more because your organization cannot push out 10X more product. And I agree. I, I spent a lot of time more doing side projects and things, which have been fun, but not that valuable in economical sense, but valuable to me, to my soul, right?
It's like...
Yeah, yeah, yeah, yeah, yeah. I, I mean, I, I don't want to overstate that. Like, like, I think probably people at AI companies today, they are being significantly sped up by, by access to AIs. I think you for your not side projects probably being-
Right
... sped, sped up by access to AIs. Um, but yeah, it's, it's, it's tricky. It's easy, it's easy to overstate.
Yeah. What's the Cognition internal tracking? How do you guys measure... What, what are, like, the... Yeah. How do you measure sped up? How do you measure how much impact-
How much do you think people-
... you have? It's like... Yeah.
Yeah.
What's your number?
I mean... Oh, me personally, quite a bit, except that I am doing a lot of non-technical stuff like organizing a conference-
Mm. Mm. Mm
... which is mostly dealing with contracts and booking guests and all that other stuff that has nothing to do with code. I would say what I've seen internally in Cognition is a lot of, like, just velocity of commits regardless of-
Yep
... whether or not you had autho- authored them. Um, and I, I do think, like, weirdly enough, like, number of PRs, let's call it, is, like, a pretty decent, like, how, how engaged are you in terms of, like, you know, like, shipping products and then also debugging and maintaining, uh, things.
Safety & Trends22:22
I don't think that there's a good measurement of, like, quality. Like, there's no story points, you know-
Right. Right. Right
... like the other, uh, guest that we had at AIE was talking about. Like, we pay people by story points. You do more-
Yeah
... you complete more story points, we'll pay you more. And, you know, there's no upper bound to that. And I think that's, like, a really interesting thing, except that you have to have a very confidence, uh, relationship between the engineer and the person assigning story points.
Um, which is effectively what you're doing. Like, your hour is a story point.
Yep.
And, like-
Right
... you know, you... we'll, we'll award the-
Right
... we reward the models based on the, the story points that they complete.
Right. Yeah. In some sense, ideally, you know, you want to get Cognition, get a bunch of other companies. You know, you randomize the companies to, to use AI or not use AI, and then, and then the, um, the, the outcome metric for your, uh, randomized control trial is how much profit they make or something-
Yeah
... or the, or their valuation after, after some period of time.
Yeah. I think, like, basically no one is stopping to do science except for you guys.
Mm.
Uh, because, uh, we know RCTs are the best, right? But sometimes human intuition is good enough that you're like, okay, I mean, when we lack data but enough humans agree, either it's mass psychosis and we're all wrong, or there's something here and we just cannot articulate it, but the benefits are...
outweigh the cost of slowing down to do the science first. Like, this is not where we're introducing, like, a new, I don't know, uh, food to the general population where we have to do a lot of safety testing.
Like, here, it's just like, it's just software, guys. Like, let's just ship it.
Totally. I, I mean, you know, thinking, thinking at Meta about why, why models today aren't catastrophically dangerous, you know, it's interesting to get the uplift numbers. It's interesting to get the time horizon numbers. But really, um, uh, why don't I believe they're dangerous?
Well, it's a mix of I watch the models do things in transcripts, and sometimes they're kind of derpy. Like they, they don't use resources well or, um, you know, they, they just sort of clearly have some of these obvious faults.
In broad deployment, only slightly worse models in the past six months have not been doing anything crazy or, or, or causing, causing great danger. The, the next model was only sort of a little bit better, and so it seems s- sort of surprising on, on, on priors if it was, if it was so dangerous.
Yeah, I, I totally think that anecdotes and intuitions are real evidence. People should totally be taking that into account.
I do wanna comment on this, this whole thing about how you are, uh, you know, the, the, the, the, this threat assessment side is in your name. And typically, I expect, let's say, EA-affiliated companies or organizations to be sort of the, on the Eliezer side of the world, where they're sort of banging the, the drum about, uh, danger.
Whereas here you're actually saying like, actually it's like a pretty balanced, like we care about AI safety, but also we're not there yet, and we, we are actually the watchdogs, uh, looking out for it. Uh, and I would say you stand out as someone not funded by the labs where-
Yeah
... let's say Arc. Is it Arc? Or some other groups that also do threat evaluations before model releases. They would typically be funded by OpenAI or some other big lab.
Meta came out of Arc, so, so I think, um-
Right. But now you're a separately funded organization. As, and, and as far as I know, it's like a big deal that you're not funded by-
Yeah
... a big lab.
Yeah. I think it's, I think it's, uh, you know, vital to have this independent source of expertise. I can, I can, I can bang that, bang that drum forever.
Yeah. Um, the other thing also is just like this, uh, concept of capability explosion, which i- which is a word that you use. That's also something I, I wrestle with, right? Like, if you believe in emergence, you believe in, uh, multiple capabilities sort of fusing together p- to produce generalized capabilities that you may not be able to, to detect, it's hard to predict based on trend lines.
It should be discontinuous in some sense, and, uh, I, I don't know that going like, "Oh, the m- N minus one model was fine, therefore the N model is probably fine." It's really hard to tell. The thing that gives me comfort is, um, yesterday I was at the OpenAI live stream, and even Sam Altman was like, "Yeah, I just let Codex just YOLO dangerous permissions, whatever, on my computer, and I, I don't approve the model anymore.
Like, it just does whatever it wants to do on my, on my laptop." And I think like, I guess the, the guard is every model lab leader dogfooding, and if it screws up their personal permissions, then, then, you know, they, they have the skin in the game is what I'm saying.
On the continuity arguments, you know, I, I'm not sure what I think. I, I, I agree that it's kind of flimsy or like this, um, you know, there's only so many models, so many data points on this, on this time horizon trend.
You know, how much should we expect it to be, to be continuous to, to, to keep going like this? Well, I'm not, I'm not sure. You, you know, maybe, maybe an intuition that it's, um, that something might be discontinuous because, you know, um, models are providing so much effective labor in, in improving the next generation of models.
You know, may- maybe that's a reasonable thing to think. On the other hand, I've been pretty surprised so far about the degree to which it's continuous and, and that gives me some faith that it's, um, that, that it might continue to be, to be continuous in future.
Um, seems kind of ambiguous to me.
I mean, we have, you know, break points in physics, right?
Mm-hmm.
It's like I'm curious if-
Yep
... it doesn't seem... Like you know, when you... It's funny. It's like when you think about water, right? It's like, water does boil-
Yeah, yeah
... at this exact temperature. It's like, I mean, maybe we do know, but I feel like we don't really know. And I feel like with models, I don't know if there seems to be the same thing because it's all just like compounding of the same thing, if that makes sense.
It's just like scaling the same thing over and over.
Yep.
But yeah, maybe we will see it. But I, I'm curious like what you would need to see to feel that is here. Because I mean, even if you look at the Opus 4.5, it's like, well, that's clearly out of trend, you know?
And so you were saying four months instead of seven months. But if then the next model is like, oh, maybe it sh-should not be four months, it should be two months, like would that make you change your mind about whether or not the months thing even makes sense?
Or like if we maybe we pass some base level after which it accelerates and will keep going. I, I don't know. I, I feel like you must be having this discussion internally.
In some sense, the thing that would really concern me is if ARND was fully automated inside of, inside of some lab, that, that would totally seem like, um, uh, the conditions are there for, for potentially a, a capabilities explosion.
If I saw, you know, time horizon of a year, I, I would still find it ambiguous, I think, at the moment whether that was the case. Because for things to be fully automated, you know, 90% automated isn't enough.
Um, uh, you need, you need some s- for some full loop to be closed, and perhaps we're missing some sort of task that points to that missing 10% or, or something. S-so I think it's, I think it's a tricky issue.
I, I, I think I, I think I can't give a number. Um, but yeah, my, my intuition for, for where water boils is at some point where this, this loop is fully closed. There are interesting debates about what exactly that loop is.
So some people talk about software-only intelligence explosions, which means, you know, even holding hardware fixed, we could get to the point where, um, just from, um, models improving themselves, they'd then be sort of smarter in this next step to create even better models with sort of even fewer resources, this, this sort of thing, and, uh, and this could lead to some extreme takeoff.
Or maybe that fizzles out, um, uh, uh, somewhat quickly, and instead you need, um, in addition to the, to the software-only capabilities, you need chip design, or maybe you even need chip production, and that's sort of this, this larger loop that can, that can close.
You know, if you, if you think that, I think you maybe should still think that closing the chip production and, um, and software-only and chip design loop is potentially very destabilizing and, and, and concerning. But yeah, tricky, tricky issues.
I, I think that is the actual paperclip factory. Like, if you incentivize a model to go build its own compute, and it, and it, it would just builds whatever it needs, and it will turn the planet into chips.
Well, I don't think it can do it. We will stop it before that. Question mark? Uh, but-
I don't know if we have the power. There's no off button, you know? Like there's...
I, I think it's, I think it's super hard to foresee. Um, but, uh, you know, a model that had that-
Yeah
... ha- had, had those kind of capabilities, you know, it's hard, hard to rule out. There would be something like a, like a capabilities explosion and, and, and who knows what happens after that point.
Yeah. So, okay. I mean, the, so there is a bunch of other benchmarks that actually directly track this, right? Like, uh, OpenAI has like Paperbench, I think, uh, which di- directly tracks its capability to reproduce papers. And I think there's a, there's a lot of...
Other than, than RE Bench, there's a lot of other sort of similar sort of ML, um, self-improvement benchmarks. They've directly, like Yakun, uh, from OpenAI has like directly prioritized, like we will have an automated AI researcher. I did a podcast with, um, Yitae from Gemini, who is also like, basically he's plugging his own training logs into Gemini to like improve his, his own code.
And I, I'm like, "Well, at some point you don't need to be here." Like um, I mean, I think this year
I, I'm, you know-
You're skeptical. Okay. Say more. Say more
... I'm, I'm not speaking for everyone at Meta.
Yeah, yeah.
I'm a relatively longer timelines, quote-unquote, person at, at, at Meta. We have Nicolo, my colleague, who, who, um, helped out with AI 2027, who's on the, who's on the shorter timelines end. Um, th- this is not, not only to you-
Which is officially AI 2028 now.
That's a, that's a conversation.
One year pass and move it back a year. Okay.
Um, yeah, I, I, I think my view would be, um, uh, not, not that Nicolo's view is necessarily different, just, just so I'm not, not, not speaking for other people at Meta. Um, that, you know, a paper bench, let's say, perfectly, perfectly measures, um, you know, not only reproducing papers, but, you know, in fact producing novel, novel research papers.
That's just a, a, a part of, um, of this R&D production process. There's also, like, your GPUs are constantly failing. Like, can you get someone to go to the data center and, and fix them in the appropriate way?
You know, can you call up the water company when, when, when, when the, when the cooling breaks down, um, uh, et cetera, et cetera, et cetera. Not, not aware of benchmarks tracking that in particular. My point is more there's this very long tail of things potentially involved in, um, in R&D that would, that would perhaps need to be fully automated i- in order to lead to capabilities explosion.
I expect we're measuring, you know, in some ways only, only a small proportion of, only a small proportion of, um, of those capabilities. And so I expect, you know, the capabilities needed for the full loop to close to, um, to, to, to come somewhat later.
Yeah.
Um, that's a contrarian view.
I, I don't think so. I think, I think that's, that's a reasonable take. Um, something I do that does surprise me in terms of when I'm talking to capabilities researchers-
Yeah
... is that, um, you guys don't have, like, a enumeration of the capabilities that matter. I, I, I mean, I think you, I think you implicitly do-
Mm
... in the, your choices that you make.
Mm-hmm.
But I think it's almost import-- Like, I always imagine, like, a wagon wheel. And this is, like, the terminology. I don't know who came up with this term, but you know what I mean. Like, that, that, like, here's, like, the 10 things we care about, and here's where everything is on, on those 10 benchmarks.
And I, I, I feel like capabilities tracking is just tracking, okay, what's that list, and then how, where are we on that list? And, uh, uh, I think I almost feel like this need to reduce everything to a single number-
Mm
... is actively working against that-
Mm-hmm
... because it, uh, reduces any form of nuance of, like, well, it's insufficient here, like the, the calling a data center thing.
Yep.
Uh, so, like, we're, we're fine. And it's like actually we should just not invest anything in that area because that's the danger, danger zone.
Yeah. I think I could not agree more that, you know, time horizon is, for instance, but many other single numbers, is, is one number and that's, like, collapsing an enormous amount of, um, uh, really important detail.
Dimensionality, yeah.
I don't know how to come up with that list of 10 and I, and I challenge you if you're-
I'm working on it
... if you're able to come up with that list of 10.
I am working on it for code.
I'll be very interested to see it for code. Um, my, my intuition is that we'll come up with a list of 10 and it'll turn out that there's a secret 11th thing that's, um-
Yeah
... that we thought was important but, but it was difficult to pre-specify ahead of time. And, and now it seems obvious that even ahead of time, if we'd had that foresight, um, that would've been helpful to add.
Well, I think that the security community does this by versioning year by year, right? So, like, this year the top 10 are blah, and then we'll just publicize it to everybody so everyone knows what top 10 is. And next year we'll have a different top 10.
But, you know, like, it obviously is stochastic and, uh, we should update our assumptions, but it, it's u- broadly useful to have that list as a public service.
You also had this research on the, um, uh, slowing AI improvements based on, like, AI compute, and you mentioned that it's like, in a way you could tie the AI time horizon to, like, the growth in compute. Can you say more about that?
It's in a way unintuitive because the compute growth is not always tied to, like, how much every single model compute needs. It's kind of like a broader market thing.
Yep.
Yeah. How did you get the two together and then some of the findings that you had?
Yeah. Maybe for a second let's take time horizon very literally. We don't have the, the qualms about it that we've, that we've just been discussing. Um, it makes sense to continue extrapolating it into the future. What are some important forces that might cause it to rise more quickly?
Some of the things we've just been talking about, automated R&D versus, um, versus go more, go more slowly. One of the most obvious forces that might cause it to go more slowly is if inputs slow. Um, one important input is compute.
You know, I, I think, I think we all have the intuition that to some extent if, if compute growth slows, um, which we expect it to at some point in, in the not so distant future, then capabilities will slow.
But by how much? It's a big, it's a big question. The suggestion in this, in this paper is that if you think that algorithmic progress, you know, that, that it's coming up with a transformer, coming up with RLHF, you know, MOEs, this, the, all, all, all of, all of this stuff, better learning rate schedules is, is, um, uh, is itself a function of compute because, you know, you, you need compute to sco- to discover it.
You know, the, the transformer, the, the gains from, from transformers, um, show up much better with scale. If you don't s- if you don't put in those resources, you know, you'll never find out that this is, um, this is the, uh, superior algorithm.
You need to run a ton of experiments. You know, e- each of the experiments can be quite compute expensive. N- not to say that no labor is involved. You know, obviously people are, people are working on this. But, um, if you think it's ultimately bottlenecked by, by compute, then algorithmic progress too slows down, right, if, if compute growth slows down.
So then if you think about, um, time horizon or, you know, whatever your favorite measure of AI capabilities is being a function of algorithms in some sense and compute in another sense, and both of them, uh, both of those components halve when compute halves.
Compute halves sort of trivially because compute is halving and algorithmic progress halves because compute is this, is, is this important input and, and compute halves, then you might expect time horizon growth to half. And then some of these major capabilities milestones that we might be interested in would be significantly delayed.
I think there, there are so many caveats to that picture. I think there clearly are some types at least of algorithmic innovations that did not require a lot of compute to, to go about creating. Some, some that took, um, uh, some, some that took a lot, a lot more compute inputs.
Prediction Markets36:44
You know, if you, if you expect that, you know, no compute inputs are required, we could just, um, sort of survey researchers for the, um, for the best ideas and then immediately put those into training the frontier models, then there'd be no slowdown of algorithmic progress from, um, from compute growth slowdown.
And of course, all of this is counteracted by the possibility of capabilities explosions or AIs providing, um, e- even short of capabilities explosions, A- AI is providing significant l- labor at making AIs better. But just analyzing the compute force o-on, on, on its own, um, you know, it might, it might lead to, to significant slowdowns depending on the degree to which it makes sense to call algorithmic progress basically determined by, determined by computes versus not needing compute to come about.
Do you think of compute on a per lab basis? Because there's kind of one, one way you can model this out is like, you know, the improvements slow down, not every company is able to stay in business, and then their compute gets recycled back into the other labs, which then kind of grow compute again.
There's almost, like, benefit to, like, the heterogeneous distribution of, like, researchers and compute. But I'm curious to, like, how much you care about, like, just a broader compute, compute is out there for people versus, like, the big labs have more and more compute.
Yeah. So s- for the paper, we use OpenAI data and, and OpenAI projections. Um, so you know, I, I think this applies more broadly, but, but we used, we used that as a, as a kind of case study.
You know, I think the, the argument I just laid out sort of goes through if you're not interested in compute at all, and you just talk about dollars, you know? What, what are the dollars going into-
Yeah
... um, going into models? You know, will algorithmic progress slow if dollars that goes into them, uh, uh, uh, slows? The, the whole, the whole argument kind of works, and that works, I think, at a, at an industry level or at a lab level, so on and so forth.
You know, I, I agree things like, um, certain labs going out of business or labs consolidating or, you know, these kind of industrial organization things, uh, would be very important. I'm, I'm laying out an extremely simple picture and the, and the, and the real picture is not s- is, is, is not extreme.
Um, but that's the, that's the basic.
We have examples of, uh, you know, xAI has been said to be distilling from Claude, right? So, like, people kind of share compute in, uh-
Hmm
... indirect ways-
Hmm
... let's call it. Uh, I think it's also very interesting. I'm, I'm just kind of curious, like, what, uh, OpenAI numbers did, did you have? Is this the, the 500 billion for Stargate or something else?
This is from their, uh, previous tax returns, the, the amount that they've spent-
Ah
... on R&D compute, and then-
Sure. Fine
... from a, um, uh, from a information reports earlier, earlier this year, some projections that OpenAI have for, um, how much they'll spend on, on compute R&D in the future, um, converting that from dollars back into FLOPS.
Yeah. Yeah. Yeah. It's i-interesting because, like-
Back to FLOPS, sorry.
Right. Um, and, and obviously all the labs, but particularly OpenAI in the last, like, three months, have, uh, basically thrown $10 billion each to every single compute provider on the planet to develop, uh, uh, alternatives to their, their current approach, which is very interesting.
But I will s- also say, like, uh, you know, don't discount Meta compute spend, don't discount xAI compute spend, and don't discount DeepMind compute spend. Uh, all of which you have basically zero visibility, right? Like
Mm-hmm. Mm-hmm.
Um, i-if you're looking at a single company, maybe that's authoritative, but then the total spend could be a lot higher.
Yep.
It's interesting. I do think-- I do also observe that, like, people like, uh, Dylan from SemiAnalysis, uh, do tend to very strongly time model progress with compute clusters coming online. Uh, which is like... Or which is like... The, the people on the model sort of API side don't see it, but this is all downstream of like, "Well, our, like, 10,000 GPU cluster just came online," and like, "Well, it takes six months to do it, and therefore Grok 5 will be here."
And like, it's pretty mathematically, like, deterministic there.
Yeah. Yeah. It seems right to me.
Yeah. It's, it's fascinating.
Yeah. I mean, they must... Yeah, because from the lab side, they must see something in the early checkpoints to, like, go ahead and keep investing 18 months from now. Because I mean, I, I, I wonder what the time gap is between finishing-
Yeah
... a good pre-training run and, like, going live. That's probably like 9 months, 12 months, something like that. You know?
Yeah. Uh, I, I think Mistral is actually pretty open about this. Um, the, uh, the plans for Mistral 3 and 4, uh, I think, like, they've been pretty open about the number of GPUs and, like, the direct timeline from coming online to when they ship the model.
It's, like, pretty, uh, pretty set. I don't, I don't have a clear timeline in my, in mind, but I would say four to six months.
Yeah.
Um, but like, yeah, it's... The, the, the compe- competition is very tight.
Right.
And one of those things is, like, um, it's also very interesting to see when labs throw away models because they failed. Like, their run was, like, came behind someone else's run that was better.
Then they were like, "Oh, well, we can't release this anymore."
Yeah. Or like release it as a... Yeah. Yeah. That, that's the biggest risk with the prediction markets on model performance, actually. Just to tie back, I'm always-
Failed runs? Yeah, yeah.
Well, it's... Yeah. It's like, okay, like, you know, when... I, I think in December there was, like, the who's gonna have the best model by end of 2025.
Yeah.
I think there was, like, a lot of activity, like, in the last few months where, like, the GPT 5.1 model came out and it's like, okay. Then I guess Gemini, because they just threw that out. I mean, said Gemini's coming out next week, and so trade that.
Um-
Do we wanna talk about Manifold, about while we're on the topic?
Yeah. Yeah. You were, like, the most profitable Manifold markets trader. Like, I, I mean, there's obviously, like, a lot of talk about insider trading on, like, these markets- ... especially in AI. Uh, I mean, I've seen it with, like, a lot of the embargo news that we get.
I'm like, man, people are trading, like, a million dollars on this market. It's like there's thousands of people that know the actual information. Uh, how... I-if you didn't have insider trading information- ... how would you think about modeling these things out?
And do you think it's, like, a worthwhile thing? Like for example, like, who's gonna have the best model in three months? Do you think that's a prediction market where you can build some sort of strategy alpha or?
I guess the, um, naive prior without, without any extra information is just, um, uh, you know, in 2025, for what percentage of time did, did which model providers, um, uh, have the, have the top model as measured by Time Horizon?
Eval Frontiers43:04
But, you know, you could do it for, for, for any old benchmark. I think that's something like 5% xAI, 50% OpenAI, 45% Anthropic. Don't, uh, don't shoot me if I'm, if I'm incorrect.
No, no. DeepMind.
I think, I think it's, it's not the case that a DeepMind model was, was at the frontier of Time Horizon-
Ah, Time Horizon. Yeah, yeah
... at any point in 2025. Um, but yeah, yeah, different, different, different things for different measurements. Yeah, may-maybe that's the same, the same prior that you want to, that you want to apply. You know, xAI was kind of coming online at the beginning of the year, so, so, so maybe, um, uh, maybe naively you want to raise, raise xAI a bit.
Uh, yeah.
I'm always curious, like when I see people betting on these things that are obviously not... There's no like real basis to-
Well, I mean, I was just gonna say-
I was broader, yeah, broader scope
... what's, what's your secret for Manifold Markets', uh, Alpha?
Yeah.
Yeah, I, I see. So, so, so the secret, if you, if you read this article about how I became the number one most profitable trader on Manifold, which sounds very nice and impressive, and you know, like I must be so good at predicting things, but actually it mostly comes down to this one market where Manifold had opened up a charity program, and the market is on how much is going to be donated through this charity program by the end of its first month.
Okay. And the, the, the market opens or, or like I, I first see it sort of five days in, and it's giving a kind of linear projection of how much has been donated so far. Let's assume that per day amount keeps getting donated every day until the end of the month.
Um, but you know, as a, as a person who gives money to charity sometimes, um, uh, I noticed that you can, you can manipulate this market in a way, right? By, by giving more to charity and so, and so, and so moving it more up.
Um, and so I think the strategy was to a ton of Mana this, um, this, this fake currency that's, that's used on Manifold into, um, the option that was above the linear projection. People keep betting against you because it doesn't look like that's happening.
I haven't actually done any donations yet. Eventually they, they caught on to what's happening, that someone's, you know, going to, going to make this, um, donation to, to move those over the edge and they're betting on that. Um, and then I did it again into the, into the next category once people had started betting on that, um, on that category above the linear projection.
And again, people bet against that and against that and against that. Um, and I mopped up those, those fake internet points. And then I think I did it once more a- as a bluff. The bluff failed, but the previous, previous two worked out and then I ended up donating, I, I can't remember exactly how, how much it was.
Not, not so much. Something like $5,000. Um-
Oh, it's all for a good cause.
Yeah, yeah. To, to, to give, well, I think, and won lots of fake internet points on the market and so became the number one most profitable trader. Um, uh, you know, s- slightly, slightly legitimately. You know, there's nothing, nothing about that that was outside of the rules.
Exactly.
Uh, but-
This, this is called, uh-
Also, I should have less respect for my forecasting abilities.
Yeah, yeah. This is, this is called prediction markets with high agency as you actually go in-
Yeah.
... and, you know, uh, the future is what you make it. So, so I think like the, the, the, the, um, to, to me the broader lesson is like, well, the, the, the classic difference between Manifold markets and Polymarket is that Polymarket is only real money, right?
Like, and so is the whole fake internet points thing a worthwhile pursuit or a waste of time? Like should, you know, do, do people actually w- want to use real dollars? And, and maybe like that's one question. The other question is obviously like prediction market ethics, which like I think always just...
it's gonna indirectly come to assassination markets. Like it's just like e- even if it's like you, there's... you ban the, the word like will someone die, like some other proxy to will someone die will happen to be an assassination market.
Like,
Yeah. I, I, I'm good friends with the Manifold Markets co-founders. Um, uh, I, I, I love them very much. I, I do, um... you know, my, my view on the, on the social value of prediction markets, which was always the dream, right?
Like it, it would be nice to have, uh, well-calibrated probabilities on, on events that matter. Um, you know, this country going to war with that country. You know, thing- things that, um, things that really matter to people. Uh, it'd be nice to have high quality information.
But you know, when I, when I look at, um, uh, real examples that have, that have come out in the past year, it doesn't seem to me like those examples are, are so, so socially valuable. Um, I, I'm not sure about assassination markets in particular.
I'm sure that, you know, I'm sure those would be, those would be overruled hopefully. Um, but um, I think gambling-like behaviors are, are socially costly and, um, uh, the value of higher quality information is, um, is, is, is real.
But you know, i- is it worth that disbenefit of, um-
Ah
... of, um, of, of people trading away their money? It's not, you know, it's not, it's not so clear to me.
Yeah, yeah. I mean, well, you know, price discovery has a cost-
Yeah
... and sometimes that is gambling and th- th- uh, that, that is the stock market. Like it funds a lot of, uh, corporate America.
Yeah. Yeah. I mean, you know, a lot of the stock markets-
Like I used to be one of those
... big firms-
Yeah, yeah
... big firms pla- you know, playing against other big firms. It doesn't have the same, you know, versus, um, versus s- sports betting markets, take an example on the other extreme, has this like very different character. You might imagine that, you know, at least one direction that prediction markets could go is this sort of big players playing against retail.
Um, and that, that maybe has a more worrying dynamic. I, I'm not, I don't-
Yeah, yeah
... such, um, so closely in touch with the space, but um, at least something like that you can imagine being concerning.
I think at a large enough scale it becomes profitable for some of the companies to do it-
Yeah. Right
... if they can get the markets-
Right
... on it. I think now it's still small enough numbers compared to the rest. Like yeah, which company has the best AI model by the end of January? $28 million of trading volume.
Oh my God.
Wow.
It's like why are people trading 20... Like peop- you know what I mean? It's like it's just crazy. You're like... But I think there's some, you know, I would... I was having dinner with somebody this week and then-
Wait, wait a second. I think it's totally possible to... Uh, I, I'm not going to do it. I think it's important that, um, you know, METR employees not be, not be, um, uh, not be making bets on prediction markets like that.
But I think it's totally possible in principle to, to have a guess, um, uh, for the answer to these kind of questions.
Oh, well, but other people know exactly. Like, you know, Gemini Three Score on Frontier MAD Benchmark by January 31st.
I see your people signed the-
It's like some people, people alre- people at Google already know-
Right, right, right, right
... what the number is.
Yeah. Yeah.
You know what I mean? It's, I, I, I think-
No, if you just-
Well, you know, if you believe in the benefits of price discovery, then this is, uh-
Yeah, exactly
... you know, this is a legitimate, uh-
Right.
Yeah, yeah. I, I mean, I, I... they actively encourage insider trading, and this is a way for insider information to sort of pseudonymously leak out, and as long as, you know, like they bear the consequences of leaking the information, whoever is like-
Right. Yeah, yeah, yeah
... is traceable to that thing. And people I think have been fired for, for doing, trading on insider infor- information. I mean, that's okay. Like the, the only step up from that is the government coming in and saying, "This is actually illegal.
Put you in jail for that." But I think for now it's self-policing.
Retroactively you can't really do that, I guess. All right. If you have any embargo news, uh, press and...
We do, we do work with people on embargo and we don't trade them, so.
Yeah. We have not made... We would've made a lot more money trading on embargo news than we've made on anything else.
Um, what else? What are other interesting model evaluation trajectories or, like, anything that you're not doing at METR that you've maybe seen other people do that you find interesting or, like, you would like more people to, to do?
Yeah. One, one project that I think is interesting is, um, AI Village that, that I think possibly both of you would've come across. These are these, um, very open-ended goals given to a village of agents, and they try to accomplish them.
It's like, um, I think set up a merchandise shop is maybe one of them. Organize an event in a park, build a human subjects experiments, this, this sort of thing. I think I, I have a number of questions about the, um, uh, about exactly what I should learn.
You know, they're using old models as well as new models in this, quote-unquote, village. Um, the models are relying a lot on vision capabilities, which we, we spoke about models, um, uh, not being so capable of today, this, this sort of thing.
But the vibe of models trying to achieve open-ended things instead of benchmark-like tasks, you know, the vibe that's a bit more like that, um, a vending machine bench, um, in some ways. Se-it seems like a very interesting direction to, to me for the, for the science to go or seems like, you know, something that, that comes with a lot of cons, but attacks some of the, some of the cons of, of benchmarks in a pretty interesting way.
I think seeing the, the ways in which these models trip up, seeing the ways in which they're, in which they're derpy is a, is an important source of information. Uh, I'd be interested in, in, in more work like that coming about.
I think that's one of them. Another is, um, transcripts as an extremely interesting, uh, source of information. This is, you know, the models taking actions and then seeing outputs and then using those outputs to commit the next action and, and, and, and so on and so forth on benchmark style tasks or, um, even more interesting on sort of in the wild deployments, um, uh, like you might find on, on your own Claude Code usage, Codex usage, HCAST usage, et cetera.
You know, that has the con of being less experimental, sort of less, less clean and scientific in some way. It's, it's, um, uh, it's more, it's more selected, quote-unquote. Like, the tasks-
METR Roadmap51:39
Mm-hmm
... that you get AIs to do are obviously the tasks you expect it has some chance of succeeding in. So it's, you're not just giving any sort of task. Um, if I see the models doing something extremely impressive or, um, or potentially unsafe in some sense, subverting user preferences, uh, it's not clear how often that kind of behavior would happen given the previous history, but it's like it's a massive data source.
There's a huge amount of, a huge amount of information there, and I'd love people to be, to be working more on that sort of thing. Um, as we mentioned, there are a lot of problems with, um, time horizon and, and our developer productivity work.
You know, I, I think it's, I think it has, um, been important evidence. I think it's, I think it's mil- moved the field forwards, but it's, but it's, but it's far from perfect. I think there, there are lots of other directions there that look, that look very interesting to me.
Maybe one that I'll call out there is this difference between whether models pass, uh, unit tests, whe-whether they, they succeed by, you know, SWE-bench like scoring, um, kind of METER like scoring, benchmark style scoring, versus whether their solution would be merged into main.
That is, you know, whether the solution adds tests where it should or, or doesn't, whether it follows existing patterns in the code base, whether it makes sure, um, that, that its changes sort of speak to other parts of the code base in, in, in, in appropriate ways.
Um, that, that seems very interesting to me. You know, I think model capabilities probably are l-lagging behind there somewhat versus, um, versus that which you might see on SWE-bench like scoring. Um, I can keep going on, but
Yeah, yeah, yeah. These are the novel research things that you were, you were referencing earlier, right?
These are, these are the novel research things, yeah.
I, I... Just to comment on the AI Village thing first. You, you, you, you, you mentioned a lot of stuff. I even want to double click on the transcript stuff.
Yeah, yeah.
Uh, the AI Village ties back to our, one of our highlights of last year, which was Noam Brown's conversation on... That, and that he's actively working on multi-agents that are co-operative instead of competitive. And, um, the, the, the basic idea that, you know, we can do more to, as a team than we can do individually or, like, the, you know, the agents are the friends we made along the way.
And, uh, I, I think that's great. Uh, and, and I think, uh, on the DeepMind side, the way they phrase it is literally having open-endedness team, which I, I think, uh, is, is a, is a topic that reemerges once a year.
I just like... Yeah, I think, uh, it's unclear what open-endedness does for us, and this is a core divide in terms of studying these things as life forms, potentially new artificial life forms versus tools for our-- that serve us.
And maybe there is... Open-endedness means that there is no goal. And, like, if you're just trying to eval this as like, "Well, what does it do for me?" That's completely wrong. And, like, you will never get anywhere with that because they are just living their lives as artificial life forms.
You know, in some sense, the, the gold standard evaluation that's, um, uh, that, that I would like to do if I was, if I was looking to learn most about the questions I'm most interested in, um, the degree to which, um, AIs might automate or accelerate, um, R&D.
Closing54:24
I'd quite like to just sort of, you know, give the AI a bunch of affordances, type into the AI, automate R&D, go- ... and, and see, and see what it does. And, um, I suspect that wouldn't work today, even, even with all the affordances, because it would fall over on its face and, you know, in, in, um, uh, when working with resources, handling, handling resource use, um, in ways it's not so capable of today.
Um, it would struggle at some types of long horizon tasks, et cetera, et cetera, et cetera. And, you know, in some ways, I, I, I think, I think benchmarks have a, uh, face difficulties in capturing this sort of thing.
And AI Village seems like a... or AI Village style things, these more sort of, um, uh, open-ended goals, seeing how, seeing how models pursue open-ended goals. You know, they, they give some color to this sort of thing, to seeing models fall on their face.
You know, I do, I do think that this will become these more open-ended goals more and more important over time. I agree to some extent, like you're going to provide them, um, you know, in, in the extreme case that I, that I just mentioned.
You're going to provide them, you know, documentation about how this part of the company works and that part of the company works and so on and so forth. It's not, it's not purely open-ended. Uh, but it's, you know, it's, it, it's, it's pretty open-ended.
It's more open-ended than the kinds of problems that we're, that we're giving them today.
Yeah. Bounded open-endedness.
Yeah. And, you know, if, if models are, um, excellent when, you know, when specs uses them, um, given some, uh, detailed issue, uh, uh, description and, you know, some, some very clear, clearly spec thing on, on, on what they're supposed to do, that, that's interesting.
But it's, but it's a very different thing, I think, from being, being able to automate R&D. Um, I'm interested in how, how far we are away from that. And, you know, in some ways, this speaks more directly to that sort of thing.
So we had the Terminal Bench guys on the podcast. How do you think about, yeah, the harness benchmarking in a way? Because if you look at their leaderboards, the same model with different harnesses, there's like 10 percentage points of difference.
Does that seem interesting? Like, I don't know if how you build the harness in Meter or whether or not you always pick the best harness or compare them.
Yeah. So, um, let's say how we pick harnesses at Meta. Um, this is not what I work on in particular. I, I, I'm not an expert. But roughly, we, um, uh, build harnesses to, to get, uh, models to be as performant as possible on a dev set of tasks, some, some held out set of tasks, and then we use those same harnesses trying to make sure they're not overfit for, um, for, for our main suite of tasks.
On the one hand, I, I do have the intuition that there is-- there's a lot of juice in, in, in scaffolding. It, it's easy to overstate how much juice there is because of this overfit problem. Or if, you know, if we were building a scaffold to do as well as possible on our test tasks, then it would do much better than, than the scaffold that was, that was, that was built only on our dev tasks.
And in some sense, that would feel sort of illegitimate or not interesting, or like you wouldn't expect that to generalize to so- to some, to some other set of tasks potentially. You know, on the other hand, a lot of work has gone in at Meta to, to, to building scaffolds that are as, um, you know, make models as, as performant as possible because we are interested in upper bounding the capabilities of models, you know, when thinking about whether these models, um, might or might not be dangerous.
So, so, you know, I do, I do have faith that these, these scaffolds are a lot better than, than, um, you know, the, the, the first thing, the first thing that people might try because so much effort has gone into them.
Yeah. It's interesting because I do want to overfit, you know? Like, as a, as a customer of the models, like you do want to overfit to your task specifically, and I think sometimes people underestimate maybe how much value you can get out of it.
But, um-
Yeah, I think, you know, if, if you have a, um, a kind of mechanical workflow or, or, or something that, that you're imagining automating it, you, you imagine automating that workflow and there's, you know, some place where more sort of stochastic intelligence would, would, would be nice inside of that, uh, like deciding where to route customers to on, on, on, on, on customer calls, something like that.
There, I feel like that makes a lot of sense, but for this sort of, you know, more general, you know, um, uh, in particular thinking about sort of helpfulness and software engineering thing, I, I'm not sure I have that same-
Yeah. Well, take an example is like I work in TypeScript, right?
Yeah, yeah.
If I build a better linter that is private to me or like a better, uh, test suite, like a better Playwright replacement, like in theory, I'm kind of overfitting the model to perform better, right?
Yep.
Like, it doesn't really matter to me.
Yep.
Like, I'm not trying to report on the model performance.
Yeah.
I'm trying to build the best thing. Yeah. I, I think like that-
Yeah, but if you have a model, build the linter.
Right. Well, no, I agree. I agree. I, I think like that's kind of like the question of like-
Yeah
... okay, should I just wait for the next model?
Yeah, yeah, yeah, yeah.
You know what I mean? It's like-
Yeah, yeah, yeah
... at what point should I be building the better scaffold? Like, Noam Brown was like, "All scaffolding is gonna get washed away."
Yep.
But on a realistic-
He, he would say that.
On a realis- realistic schedule, it's like, what am I supposed to do this week? You know, it's like-
Yes. Those simultaneously can be true.
Right.
That all scaffolding will, will be washed away, and that scaffolding today is valuable.
Totally.
Yeah.
Totally. Or within model generation it's valuable, and across model generations it's, it's, it's not so valuable. Yeah. You know, I'm, um, I'd say at best an acceptable software engineer.
Right.
Um, uh, in-intentionally not investing in engineering skills because, because the AIs are getting so good. Uh, maybe that's the wrong decision. Um, yeah, I, I, I agree. If you, if you expect, as I, as I think you should, um, capabilities to, to keep going up and up and up, forces difficult trade-offs about how you spend time today, because maybe it won't be so helpful in six months' time.
Take a sabbatical. All right? If you live in Europe, you can just take, you know, six months off or something.
Yeah. But then six months you might wanna take another sabbatical for the next six.
Perfect.
Um, cool. J-just to wrap up, like, I guess, what do we expect out of Meta in 2026? What does success look like in 2030? I don't, I don't know if you have-- there's like a sort of broader vision.
Uh, a-and then maybe on a personal side, we can talk about the karaoke stuff. But- ... let's, let's, let's talk about Meta as well.
Yeah. So from, from Meta, I think you're going to see, um, more hopefully high, high quality capabilities evidence. You know, the kind of thing you saw in the past with, um, Time Horizon and the developer productivity work. S-sort of along the lines of, of what we've been describing.
So some of these, some of these future research directions. We also have some monitoring research directions that I'm, that I'm, that I'm not so, um, not so expert in, that, that is thinking about if we can, uh, successfully apply safeguards to models attempting, attempting dangerous tasks.
There's a whole line of work there.
Is that an interpretability dimension or what kind of safeguards?
So usually this is black box not, not white box in, in, in my understanding in, in current work. So, so not using interpretability, but, but you can imagine in principle doing, doing, doing something, doing something more white box.
And then this, um, this risk assessment work that is taking into account, um, how capable we think models are, uh, what their propensities are, you know, whether we can, um, uh, track using safeguards, the, the kinds of, the kinds of things that the models are doing.
You know, do, do we think that these models pose large scale harms? Um, you can expect to see much more of that in 2026. Maybe, maybe now is a good time to say that we are hiring.
Yes.
Um, on my team we're, we're hiring for research engineers, research scientists, people from, you know, startup backgrounds, from, um, from ML backgrounds, of course. I'm originally from sort of economics or quantitative genomics background, so, so we, we are accepting a pretty, pretty wide range of, of people, um, who, who, who get stuff done, the kind of, the kind of stuff that you've seen in, in, in past Meta work, um, as well as a, a director of operations, I think, I think.
On the Meta jobs page, you can, you can find out more.
Yeah. Everyone I know is hiring a director of operations, including, including myself. And, uh, I feel like that's, you know, probably the, the, the one agent that everyone wants that we cannot have. Um, yeah, I mean, like, you know, whenever people have a hiring pitch, I always try to push for, okay, the average candidate comes in, you reject them.
Why? What's the, what's the thing you're looking for?
So there, there are lots of different shapes of people we can look for, so, so different, different stuff for, for, for, for different folks. One thing is, is, um, you know, good kind of basic research intuitions like, um, like checking your data.
You know, we, we don't work on pre-training at Meta, but if you're working on pre-training, you should look at the corpus to get some sense of, of what's going into the models. Even working on this uplift RCT, um, that was, that was, that was pretty important.
You know, really, really having a shape of, of these issues in your head. I think people who are communicating in writing with, with sort of a lot of transparency, not overstating their results, um, you know, my hope is that, that your sense of, of meat work in the past is that it's, um, it's trying to be level-headed, not, not to, um, not to understate, not to overstate what the, what the, what the science says.
That's important internally. It's, it's, it's important, it's important externally. You know, and, and then I think productivity or, or, or something. That there are a lot of, um, a lot of people with, um, uh, great talents who, who are not going to, um, work quite as well in a, um, scrapper environment working on, working on sort of frontier science and, um, and that's, and that's the thing we do.
I just want to prime people for, I guess, what are the valuable skills in this new age? Uh, because I think the, the more people articulate what the positive directions are-
Mm
... what is hired, hard to hire for-
Mm
... that's what we guide our audience towards, you know, improving themselves in. I think that's important.
Hell yeah.
Did you have a karaoke question or... Uh, I don't know. There's a... Well, I mean, like, you know, I think there's a-
Are you gonna sing on, on podcast?
There's a... I've never done it.
Come on.
I, I've been, I've been wondering-
Can't Help Falling in Love With You.
Uh, what, what, what is this karaoke thing that you organize? Is this like a, like a music... You, you're a, you're a musician.
Um, a musician might be exaggerating it, but, you know, I, I hit instruments and noises come out. Um, um, so I, I've hosted a couple of these, uh, live band karaoke events, that is like getting, getting a group of friends together and people, um, accompanied by, by a band, um, singing karaoke to, to an audience of, you know, 50, 100, 200 people.
It's great fun. I think, I think people should be, should be doing more of this. Um, I look forward to seeing you both at the next one.
Yeah. I, I will, I will do that at, at y- one of your events. Uh, yeah, it's one of those things where it, it's weird because I used to be in a, a cappella a lot.
Oh, wow.
And I just think it's like a dying form.
Yeah.
Uh, you know, I just watched this video that was really good about like 20 time- 2010's wave of a cappella from like Pitch Perfect. Uh, Glee to Pitch Perfect-
Yeah, yeah
... to... That's, uh, what's the, what's the, what's that group?
Pentatonix?
Pentatonix, exactly, and that's where it died. And then uh, it's, it's very interesting to, to see how it's like dying as a, as an art form in general, and like how new formats have taken over. And I don't know.
It's, it's, it's weird for humans also because, like, now I'm also like, let's say, let's call it more interested in synthetic song generation or DJing, anything like that.
Yeah.
That, that the human voice is actually more commoditized. Like, it doesn't really matter who sings it.
I don't know. I feel like there's a kind of transcendence to singing in person, um, uh, that, uh, you know, the, the AI generator songs are not providing me.
That's good. That's good. Yeah, yeah. I mean, I, I, I do, I do think that, like, we humans always like want that. Um-
Yeah, yeah
... but I'm not sure humans in the year 3000 will want that. Um, it's one of those weird things. Uh, well, thank you for coming on. It's, uh, it's great to have you as a human in person here.
Thank you so much for having me as a human.
Yeah. So someday we'll, we'll interview AI versions of you.






