LALatent SpaceJul 18, 2025· 39:41

⚡️ARC-AGI-3: The Interactive Reasoning Benchmark

Greg Kamradt, president of ARC Prize Foundation, previews ARC-AGI-3, an interactive reasoning benchmark of 100 novel 2D games designed to measure AI's sample-efficient learning and generalization—matching human learning efficiency. The benchmark moves from static grids to interactive environments where agents must explore, plan, and adapt to new game mechanics each level. Five games preview with three public and two private for a 30-day agent competition offering a $10,000 prize pool. Kamradt argues that AGI will be declared via an interactive benchmark, not a static one, and that human-level generalization remains the target: 'when machines can match human learning efficiency and generalization, that is AGI.' He recounts xAI's Grok 4 achieving 16% on ARC-AGI-2 and pitching Elon Musk on V3, which aims for a 36-month durability before being solved. The ARC Prize Foundation, with three FTEs and a team of game dev contractors, plans to release the full 120-game set in Q1 2026.

  1. 0:00Intro
  2. 2:16Defining Intelligence
  3. 7:02Locksmith Demo
  4. 12:23API & Agents
  5. 17:03Efficiency Metrics
  6. 23:30Team & Roadmap
  7. 32:01Grok 4 Recap
  8. 38:47Outro

Powered by PodHood

Transcript

Intro0:00

Alessio0:03

Hey, everyone. Welcome to another Latent Space Lightning Pod. This is Alessio, partner and CTO at Decibel, and I'm joined by Swyx, founder of Smol AI.

Swyx0:11

Hello, hello. We are just all recovering from the Grok 4 live stream, uh, and, uh, we're so happy to have Greg Kamradt from the ARC AGI Foundation. Is that, is that the full name of it?

Greg Kamradt0:22

That's exactly. ARC AGI Foundation. No-

Swyx0:24

Yeah

Greg Kamradt0:24

... I'm sorry. ARC Prize Foundation.

Swyx0:25

ARC Prize Foundation. That is the sponsor of the ARC AGI challenge benchmark. We're also recording this-- So we're recording this one day after the Grok 4 release when... And you were there, you spoke with Elon about, uh, their progress on ARC AGI, but then you're also doing this launch event for ARC AGI V3 next week that we'll also be at.

So this is a little bit of a preview, a little bit of a recap, a little bit of an introduction to people for ARC AGI. It is clearly becoming... I think, like this time last year, people weren't super taking ARC AGI that seriously, and now they really, really are.

First of all, congrat-congrats on picking the right horse. And then, and second of all, maybe let's, let's get a high level, like how are you introducing ARC AGI these days?

Greg Kamradt1:05

Yeah, absolutely. So ARC AGI, we are a nonprofit that wants to act as the North Star towards AGI. And the way that we do that is we build benchmarks. Because w- the way that I think about a benchmark, it's kind of like a target that's out there in the sky, and it says, "Hey, this is where we need to go, and this is where we need to incentivize research to go for it."

So the first benchmark Francois Chollet came up with in 2019, and he came out with this paper on the measure of intelligence, which is kind of-- it's an interesting way to come out with a benchmark because he first started to try to define intelligence, and then you define a benchmark that can actually go and measure that definition of intelligence.

So he did that in 2019, and then in 2024, Ike Knoop, one of the co-founders of ARC Prize, he went to Francois and he said, "Francois, either I'm wrong or this benchmark is drastically underrated." So Mike actually put, um, he put up a million dollars of his own money and said, "Hey, I'm gonna put a bounty.

So for anybody who can beat this benchmark, they're gonna get a million dollars." And that's where ARC Prize, the competition, came out in 2024. And then, uh, this year we expanded the mission and went into a full-blown nonprofit.

So now not only is it the competition, we're also doing a lot of the leaderboard stuff. We're building another benchmark and we're about to come out with our first platform, or at least a preview of the first platform on, uh, interacting with that benchmark.

Defining Intelligence2:16

Alessio2:16

What has been the most surprising thing working on it for you in the last year or so? Like, y-you went from underrated to obviously, as Shaun said, very popular. Do you think there's something fundamentally different about either the models that you're testing against or people's interests or just right time, the space is growing?

Greg Kamradt2:32

Well, I, I think nothing happens in a vacuum. Right time, the space is growing. However, the way that we think about our company is if we were a company or a nonprofit, if we're gonna equate it to a startup, the product that we have is the benchmark.

And so you're not gonna convince the community. You, you can't BS your way through a poorly, inappropriately designed benchmark. And so the great work that Francois did, like the foundational research that he did and the flag he put in the ground, that is what the differentiator was for it.

And so in 2024, um, we were solving for awareness and we're very happy where that ended at the end of last year. Um, we left, uh, 2024 with OpenAI, uh, inviting us to come join them on their o3 model, uh, preview for it and like desc-- and, uh, an-- co-announcing the results for that.

So, um, solve through awareness, and now we have to solve for a whole lot more.

Swyx3:19

I, I, I think also, I think there's a little bit... There's a, there's a fit between the reasoning paradigm in ARC AGI and in some sense, like ARC AGI one was, uh, almost like too early. Like or honestly, just like well-timed because it was like already a well-established benchmark and people found like, "Oh yeah," like actually this somehow tests abstract reasoning and learning, uh, acquiring intelligence on the fly.

I, I forget what, um, Francois's specific term for this, like his definition of AGI.

Greg Kamradt3:47

Yeah. Well, so I mean, this is where we get into the really interesting stuff, and you're gonna get me going here. But Francois's definition of intelligence is not necessarily how well you do at any one test. So you can buy, buy arbitrary levels of skill.

Just if you go practice more, you're gonna get better at a skill, right? Sorry about that. You're gonna get better at one skill. Francois's definition of intelligence is what is your ability to learn new things? So can you go and learn new things?

So we already know that AI can beat chess, and we know that it can beat Go, we know that it can do self-driving, but those systems, they cannot generalize to other domains outside of their training data. And that's the important part is can you actually generalize for that?

So the way that Francois calls it is skill acquisition, but then the, the kicker that comes with this, which is really important, is skill acquisition efficiency. And so what is your efficiency to actually learn new things? And so when you talk about efficiency, you're talking about a ratio and there's a denominator in that, right?

So what is the denominator of intelligence? Well, there's two things. Number one is the amount of energy required to actually go and learn new things. So the reason why this is so cool and important is because ARC, uh, ARC Prize Foundation, we use humans as the benchmark for what general intelligence is because that's our only proof point of general intelligence out there.

It's the human brain, and we know how much energy the human brain takes. You can literally measure like the number of calories and like wattage that the human brain actually does. And what's really cool is with that, you have an intelligence output and you have an energy input, so you know how much that is.

And you can directly compare that to the energy required for current AI, like how much enerd-energy do they, uh, do they take for it. Now, the second denominator is how much training data do you need in order to output that intelligence that you see from there.

So humans, they do not have an Internet's worth of training data, um, in their training data, but yet they can still do generally intelligent things. Whereas we're not seeing that with AI right now. So energy and training data are the two denominators we u- we use for intelligence.

Alessio5:38

So if somebody were to maybe have a still man argument of this would be that everybody's saying energy cost is going to zero, and while humans are limited by clock time of like their life on what they can learn, on machines, we can just kind of parallelize all of it.

So maybe can you talk about why it's actually important to generalize and versus somebody might say, "Hey, you know, we can just do RL in all the different domains." Is it because the data is not there, it's hard to build the environments?

And how does that tie into the games that you're building for the new challenge?

Greg Kamradt6:07

Yeah. Yeah. Yeah. Yeah. So I hear that argument a lot. Oh, you can just RL, like, on anything, right? And I think one could argue that if you had a proper environment, then yes, you could RL on a whole bunch.

And humans are RL-ing on the ultimate eval engine that there possibly is, which is physics, reality, right? That, that is the ultimate eval engine that we're all going for. Now, if you're not using physics or an eval engine, then you have to have a manufactured simulation RL environment.

And often, what happens is the human or developer intelligence is often injected into that environment itself. And so the model isn't actually intelligent. You're just almost, like, taking the intelligence from the developer, injecting it into the environment, and then, then you're transferring it over into the, into the AI from there.

So the important part of the intelligence definition is your ability to learn new things, but that's on unseen tasks. So it's really difficult to create environments for things you can't anticipate, but yet humans are good at that, right?

Like, we come out the womb and all of a sudden we can... Not all of a sudden, but throughout our lifetime, we can learn to drive, we can learn to play chess, we can learn English, we can learn all these different things.

And so generalization is the key piece of that. Now, what you're bringing up though is what we're announcing, um, either today or however it may be for this podcast we're recording, is we're coming out with ARC-AGI 3. And what this is gonna be is it's gonna be a series of 100 different novel environments, or you could simply call them a n- 100 different novel games that we're making ourselves.

Locksmith Demo7:02

Greg Kamradt7:23

So these are simple games, 2D games, that humans can do, that are really easy for humans to be able to play and intuit about them, but AI's still gonna have a really, really hard time with this. Now, the reason why we're moving to actual games is because of this movement that goes from static benchmarks in- to interactive benchmarks.

My hypothesis is that when AGI is declared, it will happen via an interactive benchmark. We're not gonna know that AG- AGI is here just via a static benchmark. And the reason why is because humans are very good at intuiting their environment, understanding what the goals are, making a plan, making a long horizon plan, and you're not gonna get any of that from a static benchmark, and you're required to have an interactive benchmark.

So that's what ARC-AGI 3 is gonna be.

Alessio8:04

Yeah. Let's run through whatever you can share. I, I know some people, a lot of people are watching on YouTube, so it's always nice to have a follow along.

Greg Kamradt8:11

Yeah, absolutely. Well, uh, so I just got word from my a- uh, developer that Sean should have good access for this. Sean, are you there and you can share your screen and try this out for us?

Alessio8:19

I don't think Sean is here. He just texted me that it froze. But you can share. You can share.

Greg Kamradt8:24

All right. Yeah.

Alessio8:25

We can go through it.

Greg Kamradt8:25

Let me go for... Let's go for it here. We're gonna do... Awesome. So here we have an example of one of our games, and each one of the games has a, um, has a fun name that goes along with it.

So as a part of the preview, we're launching five games. Now, three of them are gonna be public on day one, and two of them are gonna be private. Now, the reason why we're gonna be, uh, two are private is because, uh, we're having an agent competition.

So we wanna see how good can the community be against, uh, trying to beat these things. And so we'll release the, um, two, uh, two extra ones at the end of the competition. It's only 30 days long, and so you'll get these immediately.

Now, here we are looking at, uh, Locksmith. Now, you just look at this one screenshot and it doesn't make a ton of sense, but that's on purpose because you need to explore in order for it to, like, intu- i- in order to intuit what's going on.

So we'll show the same thing to AI, except the AI is gonna get a JSON grid of, uh, a list of lists. So they'll just get a bunch of numbers, 64 by 64, and they can choose to turn that into an image if they want to.

We're agnostic. Do whatever you want with it, if they wanna do multimodal. So what I'm gonna start doing is I'm gonna start clicking around and trying to figure out what are the rules of the environment that I wanna do.

So I see there's some sort of walls, and I'll go over this dark thing. Nothing there. Let me go over this one. And you can see here that in the bottom left, this object changed, and now it matches what's in this, uh, black area over here.

Okay. Well, let me go over there. Nice. That looks like it was... My guess is that you're gonna have to match what's in this bottom left to what's in the, um, to what's in the black square over here.

So let me go try that out. Let me make that match. Okay, cool. Now I got a match. Let me go down. Oh, dang. I ran out of energy. I ran out of life. Oh, there's this purple thing up-

Alessio10:04

Oh my God.

Greg Kamradt10:06

I need to go get that. And so now you're seeing how there's rules about the environment that force you to explore to understand what it is. I just picked up more life. Now I can successfully go down to the dark one over here, right?

Now, what we have is each new level should introduce a g- new game mechanic. Now, the reason why we introduce a new game mechanic is because we're testing your ability to learn on the fly. So now there's new rules about the environment, and this is what humans are very good at, is sample efficient learning, right?

Now, let's go through here. Let me try this one. You'll notice that this dark square is a different color than what's in the bottom right here. So now you need to switch the color from what this thing is to try to make it match.

I wanna be mindful of my life right here. So let me go, let me get some more there. Then let me ma- change the object, and let's go down. Let's make sure that... Ah, see? I just ran out of life.

I, I play this all the time and I, I need to redo it. Let me do that one more time. Okay, we have that. Let me go up. All right, let me try to match this object. And oh my goodness, I...

This is probably an annoying demo if you don't know what you're doing here. All right, we'll do it one more time. Get more life.

Alessio11:14

I think it's positive that you know the game and you still struggle to complete it.

Greg Kamradt11:19

Well, so-

Alessio11:20

That is part of this.

Greg Kamradt11:21

What's interesting here, though, is even though I know the game, that doesn't

omit the fact that I need to pre-plan what I need to go do because, like, we need to actually do some planning on this one, and that's what humans are very good at. So as you start thinking about CoTs and reasoning here, reasoning in AI is gonna have to plan their actions and what they're gonna have to go do for it.

So now that we've matched up the symbols, okay, great. Looks good. Now we go to this last one, and, uh, this is where I'll stop right here. Let me make sure this matches up. Let me get the right color.

All right, let me get some more life. Let me match the shape. Okay, so I matched the shape, but notice that it's out of rotation right now. So now I need to go to this last one and actually need to rotate this object And let me get some more life.

Let me go through here and, and go all the way down there. So we're forcing exploration, you're forcing long-term planning, you're forcing understanding what the different actions do and how the actions map to the action space. And then you're getting just, um, minimal sparse rewards for if, if you actually did something.

Anyway, so th-this is Locksmith, our first demo for it.

API & Agents12:23

Alessio12:23

Nice. Are you limited to only-- to the easy-to-translate-to-text environments basically, or, um, how, how are you thinking about designing the rest of the games?

Greg Kamradt12:34

Yeah. So the API actually, so what-- when agents play this, what they will get... Let me, let me jump out of here. Um, what agents will get is agents will get a, a series of frames, and those frames will be sixty-four by sixty-four.

Now, generally, it's just gonna be one frame, but you might be able to get, like, maybe two in a row or three in a row, and that would show an animation. And so beginning state, middle state, end state.

And so you can kinda get a perception of over time there. What we ask for back is just one through six, a, a series of one through six actions, and it's really just integers. And so you pass back a one, two, three, four, five, or six.

One, two, three, four, five, those are just, like, we call them basic actions. And then six could represent a click. So at six, you submit a, a six action and a coordinate somewhere on the screen, and that represents that you actually clicked on something on there.

Now, we li- we like a really scoped environment because that removes a ton of... We'll call it, it, it's really easy to analyze the learning efficiency for humans and AI if we scope down the environment really well. So to your point, yes, it will just be two by two in the beginning.

Swyx13:33

So another, another thing that I think is a consistent finding from people that I just wanted to double-check is that you pretty much always emphasize a visual representation, but most models do not benefit from visual learning, right? Is, is that-- does that still hold, actually?

Greg Kamradt13:49

You know, we get that, we get that pushback all the time. We are agnostic as to how the model represents the data to themselves or anybody represents it. So we've had people turn the j-- Like, we give a JSON grid.

We've had people turn that into emojis. We've had different text representations. We've had people turn it into pictures. Anything is on board if they wanna go for it. Now, we communicate these puzzles to humans via pictures 'cause it's so much easier to do it via visual.

Um, it, it's a valid, it's a valid pushback we hear all the time. We're not concerned about it because we're agnostic as to how you represent the data to the model.

Swyx14:19

I wasn't asking about the pushback. I, this is me using some prior knowledge of the solutions to ARC-AGI that-

Greg Kamradt14:25

Yeah

Swyx14:25

... uh, apparently most people say that adding multimodal vision doesn't actually help.

Greg Kamradt14:30

Yeah.

Swyx14:30

Uh, and I'm just-

Greg Kamradt14:31

Sure

Swyx14:31

... I'm just wondering if that's still true.

Greg Kamradt14:32

Yeah, yeah. It's, it's still true. We have n-- We have yet to have an AI successfully beat any level on any of these games, so it hasn't happened yet.

Swyx14:41

Okay.

Greg Kamradt14:42

Yeah.

Swyx14:43

Yeah. No, I mean, uh, you know, I, I think that there's just the, the previous ARC-AGIs... I'm just right looking for, like... Because, uh, for example, in the game that you showed us, you're representing life, and you're rep-you're representing sort of planning.

Presumably, there's this sort of, like, sort of path planning algorithm, like an A* or something, that, uh, people are gonna have to end up intuiting or, you know, finding. If, if they don't explicitly program it, they're gonna have to model it, basically.

And there's just, just a question of like, you know, how much does this vision help and, uh, you know, uh, I, I definitely am using it a lot in, in solving it, but language models might act differently.

Greg Kamradt15:19

Yeah. And, um, that's one of the reasons why we're having our agent competition too, 'cause we wanna know-- We wanna use the collective intelligence of the community and say, "Hey, what is-- what are the types of s-- like, what are the types of scaffolds that are needed in order to be competitive on this?

And what are the types of scores you can get for it?" So the way that we're gonna judge the agent competition is we will have a gener-- like a generalization score, which is like, how well do you do on the private test set?

We know people are gonna overfit to the public stuff, but we'll, we'll, um, place a higher weighting on the private side. But that's a good point. And one of the things I'm actually really curious about is what role does scaffolding play in future AGI?

Now, I know that sounds kind of like, like a weird question, but my current hypothesis is that AGI will be heavily scaffolded. And why do you have that? Well, my hypothesis is that you're gonna need different components that are working together in order to get the effect that you want.

So, for example, in this game, you need to have a concept of memory because certain actions that you took back then are gonna inform your future actions. Is that all just gonna be held in the context window? I don't know.

That's a lot of tokens that come from that. Is that the most efficient way to do it? Probably not. You're probably gonna wanna do some compression. Okay. Well, then all of a sudden, you have the beginning of a harness that's sitting right there.

So, um, one of the questions I have going into the agent competition is: what is the minimum viable harness that you need in order just to swap out the model and see how they do on this? And then if you're gonna have the Porsche of harnesses, like, what-- if you're gonna go crazy, what are you gonna go do for it?

And so that's why we're gonna throw up a little bit of prize money, incentivize people, and see what they do for it.

Swyx16:43

Yeah. I think it's a really interesting open question. I, I would say definitely the researchers, you know, like we, we specifically asked this question to Noam Brown, and his answer was like, "Yeah, scaffolds are all, all going to die."

Greg Kamradt16:52

Yep.

Swyx16:52

So he's definitely on, like, the sort of thin agent camp.

Greg Kamradt16:55

Yeah. I think that the a-- um, the jury is still out for that, but I think that... Well, yeah, the, the jury's still out. I could argue both ways for it.

Efficiency Metrics17:03

Alessio17:03

I know in the B2, you also had the dollar per task thing. How do you think about keeping that in B3? And then are you recalculating B2 based on the price drop of the, of the models? Because I think like, for example, o3 was like, I think, the most expensive one, but I know they dropped the price by a bunch.

So I'm curious how you balance the how much to weight it versus how much it's kinda like a temporal thing.

Greg Kamradt17:27

Totally. Well, so two things here. If a model-- if a provider drops the price, yes, we will update our pricing. And it just-- we just wanna reflect reality. So we, we would say in our reporting, "Hey, the price is up here.

Now it's down here." And that, and that, that's what it is. Now, to the, to the first question, though, is around efficiency, right? 'Cause like I said, if intelligence is a, is a fraction or is a, uh, efficiency metric, the denominator's energy and training data.

We don't get that for closed models. Obviously, they're not gonna tell us how much energy it takes and how much training data they use. We can only guess it. So we use cost as a proxy for that because, in theory, cost is market efficient.

Now, it's not perfect 'cause it's not a complete commodity yet, but it's a-- in, in theory, we-we're okay with the, with the error buffer there. Now for V3, that's a different, that's a different question. What's cool about interactivity is we get a new efficiency metric.

So yes, we still get cost. Yes, we still get training data. But now we actually get action efficiency. So the way I like to demonstrate this is imagine if you're playing that Locksmith game I just did, and you had a brute force random agent.

In fact, that's one of the templates we'll launch with is just like a random agent. It's gonna be extremely inefficient. And one of the quality checks that we do is we run a random agent at a million steps to see if it beats it or not.

And no, it doesn't beat Lockstep at all or Locksmith. And so what we'll do is when we come out with ARC-AGI 3 in Q1 of next year, 2026, we're gonna go and test hundreds of people on how they do on Locksmith right here.

And what's nice is that you can get-- on the X-axis, you're gonna get number of actions, and on the Y-axis, you're gonna get levels completed until game over. And you're gonna get a little line that goes up and up and up and up.

And that slope, you can kind of think of that as your learning efficiency, right? If you take a lot of actions to go a little bit of levels, well, you're not really learning the rules of the game, and we're gonna know who's the best, we're gonna know where it lands.

Now, the brute force agent, the random agent that I was just talking about, that's gonna be a flat line all the way out. It's not even gonna be even close. So when we report, um, learning efficiency for this, especially with AI versus humans, it's all gonna be around how many actions do you take in order to complete the goal of the environment.

Which not only does that encompass learning what the environment entails, but also executing what you perceive to be the goal of the environment as well. So it's almost learning and execution there.

Swyx19:31

Yeah, that's, that's, uh... I like that you're consciously designing with that chart in mind and, and specifically tracking, you know, what you have defined, uh, what Francois has, has been pushing as an important dimension for AGI, uh, which I think is very relevant.

If I had one criticism that came to my mind while you were describing this is that it is very sort of embodied. It's very single agent, and that may be not the case. We might want to do multi-agent, we might want to do branching.

Uh, I, I don't know if you have responses to that kind of stuff.

Greg Kamradt20:01

Yeah. So what's actually really cool is-- So the one game I showed you is agent-based, like you have a little thing that's moving around. We have a requirement that games must be novel from each other. We have a large percentage of games that are non-agent-based.

So think of it as like, like Solitaire or Connect 4 or like Simon or Memory or something like that. Those are non-agent-based games, and so game mechanics will be completely, uh, different across those. Now, to your other point, we are also designing some games that require cooperation with other things in the environment to be able to complete the goal.

And so you can think of it as, uh, if we have 100 different games, we can tag different levels and different games with different skills required in order to complete that thing. One of those skills may be cooperation or alignment with other things in order to complete the common goal.

So just making this up, for example, kind of like prisoner's dilemma. If you were to just go right for the exit or something, you're gonna get a lower score than if you were to cooperate. Or maybe you can't even complete the game unless you go and cooperate, and we can start to measure that.

Swyx20:59

That's, that's actually super interesting, um, because then you start to have to model other entities, and it's not just a, a map. Yeah, I mean, I think that makes sense.

Greg Kamradt21:07

And so we don't, we don't have plans yet to have multi-agent in the fact of like two LLMs battling against each other because we're really kind of focused on the one, but yet we may have deterministic agents out there in the environment that are kind of more code and program-based.

The other thing I wanted to show you really quick is going back to that learning efficiency piece. We're actually taking inspiration from a, uh, Josh Tenenbaum paper that came out, um, in 2000-- I think it was 2021. And they actually-- This is a, this is a little bit of the same thing.

They, they... They're not doing the exact same test that we have, but you can see that there's number of steps for these games, number of steps taken by agents and the number of levels that they were able to do.

And then they started to differentiate between the, um, like human players and AI players. Now, our games are gonna be different. It's gonna look completely different, but this is the chart that I have in mind for when we go and start to report the, uh, efficiencies here.

Alessio21:58

Was the 120 games number picked with any math behind or just a good, a good number?

Greg Kamradt22:04

It's kind of a balance of what do we need for statistical significance versus what is realistic from an operations perspective in order to actually go and execute these. So in order to build each game, we actually have a team of seven game developer contractors right now.

So we call it our assembly line, and it literally starts with a good game idea. Like, let me even back up from there. Francois made an architecture document that's basically, "This is what makes a good game. This is why it's gonna measure intelligence."

I hope he open sources that at some point. Maybe we'll do that with the actual, um, benchmark. Then we take that, and then we turn it into a game idea and game spec, and we go and hand that to one of our, uh, game developer contractors.

They make it for a week. We go and test it. We do some light testing, some QA, is it a good game, you know, et cetera, et cetera, and we mark it as done, and then we go from there.

Our goal is to come out with 120 by Q1 of next year. And so that means that we need to make 20 per month. We need to make five per week because there's four weeks in a month. So we need five or six game devs that are helping us making these.

And so we got the assembly line. And so a lot of folks come to us and say, "Hey, Greg, we'll make programmatic games for you. Like, we'll just go and make an LLM, go make these, and we'll just make them or whatever."

And the problem with that is that we don't wanna incentivize AI to derive the program that made the game, right? And so if we continue to have humans make the game, then the AI is incentivized to try to reverse engineer the G inside of humans, and that's kind of the whole point of what we're trying to do here.

So anyway, to answer your question long-wind away, it's an operational limit mixed with what do we need for data significance.

Alessio23:30

And so today you have the game devs, you have an engineering team that kind of builds the gaming platform. Uh, what's the breakdown of the team overall?

Team & Roadmap23:30

Greg Kamradt23:37

Yeah. So like we have three FTEs on Arc Prize right now. I'm one of them, and it's kind of like GNA and like ev-- like everything else. So I, I run Arc Prize right now. And then we have a very awesome generalist engineer.

He's an absolute rock star. His name is David. He's doing everything from serving the API to user auth to management to game serving to QA to all that other good stuff. And then we have a lead game dev.

His name is Hunter, and he is the one who looks over it and kind of does the quality check and idea inspiration for all the different ga- uh, game devs from there. Now, Mike and Francois are still very involved, but they're not full-time.

They're on the board that comes with it. So Arc Prize eng-- our engineering team is, is one person.

Alessio24:11

Are you planning on hiring? Should people reach out? Should people contract? What's the any call for action for people?

Greg Kamradt24:18

Yeah, you know, w-we're opportunistic with that. As a nonprofit, we are very, uh, budget conscious and mindful for what we can. We're not-- We don't, um, we haven't like raised a whole bunch of cash for it. So I would say it's also a bit of, uh, financial driven as well.

And that's one of the things, you know, that's on my plate as well is going and fundraising for Arc Prize with, uh, philanthropic donors. So I would say have folks reach out. We may not have a spot. We may have a spot.

We'll see how it goes.

Alessio24:39

Yeah.

Swyx24:39

Uh, uh, so I think it's a very noble cause and, uh, very impressive what you've done so far. I think there's always the, the, the question about, like, the future roadmap for ARC-AGI. Presumably, you know, you, you went to two and three really quickly, um, and presumably there is, there's thinking on four and five.

I think, like, constantly whenever ARC-AGI comes out, the, the question is always, you know, how does this apply to sort of real-world situations? It is becoming more and more real world. Does this end with, you know, you cloning Dota?

'Cause it-it's starting to be a game. It's like, you know... Yeah, how, how-- where does this end? Where does this go?

Greg Kamradt25:11

Well-

Swyx25:11

What's V4, V5?

Greg Kamradt25:13

Yeah, I tell you what, my, my view on it is at the Grok 4 livestream last night, Elon actually said something along these lines which I agree with, which is reality is the ultimate eval engine. Like, that's as simple as that.

You're n-you're not breaking laws of physics. That's what we all operate in. If I could have any benchmark possible, I would want a perfect simulator of reality and of course who wouldn't, but like a perfect simulator of reality and then that's when you go and b-- that's where you go and simulate on.

That's your environment that you can just manufacture different tests for that. In lieu of that, we have approximations that happen, um, every now and then. So the way I think about it is if you had like a linear, um, spectrum, ARC-AGI 1 and 2, it's a static list of like three or four JSON grids.

Static. Doesn't move. ARC-AGI 3, I guess on the other side, would be reality, like pure representation of reality. ARC-AGI 3 is a step more towards reality, but it's still gonna be a scoped environment. So without knowing what the exact answer is, I do know that ARC-AGI 4 or 5 or whatever it may be will need to allow us to have more axes of freedom that, um, are above a 2D 64-by-64 type of grid that comes from there.

But one thing I do wanna h-- I do wanna emphasize here is one of the things that sets us apart from other benchmark creators is we don't aim for PhD++ problems, like the hardest possible thing that nobody understands.

Our anchor and our constraint, which is actually really freeing, is can humans do this, right? Because our hypothesis and our definition of AGI is as long as we can come up with problems that humans can do and AI cannot, then we do not have AGI.

And then the flip side of that is also true, which is when us as ARC Prize, we're like we consider ourselves, like our job is to come up with problems that humans can do and AI cannot. When we can no longer do that, for all intents and purposes, that's practically AGI.

And the fact that ARC 2 is still out there and humans can do them, and the fact that ARC 3's gonna be out there and humans can do them proves that we do not yet have AGI. So the other thing I do know about ARC 4, and this is where I'll wrap it up, is that ARC 4 will still be doable for humans, but it's still gonna be hard for AI, but it's gonna need to have more axes of freedom for us to test.

Alessio27:08

Do you have a timeline on when you think V2 will get close to like 50%, then 100%? Because I know Grok 4 yesterday released 16%, which everybody was going crazy over, but it's still 16%. Uh, so, uh, how do you think about saturation of current version versus when to build the next one?

And then, you know, you're releasing V3 in 2026. Do you think that's a benchmark that will last another kinda year and a half, two years? That's kind of the rough timeline for you?

Greg Kamradt27:36

Yeah. So I, I will say that the V3 timeline also depends on how V2 goes. So like right now, like you said, if we end the end of the year and it's 16% still, w- it's just a different dynamic on if we wanna introduce V3 or if we don't.

Uh, we, we don't necessarily have a timeline on V2. My hypothesis is it's not gonna get beat in 2025. That's only six months away. But like if, if we really back this out here, like what are the things that would like surpass it?

Would an o4 Pro, whenever the heck that comes out, would, would that beat it? I'm not sure. I'm fair- I'm more confident to say that if you have a base LLM, like a GPT-5, if a, if GPT-5 is a base LLM without all the routing and COT and all that crazy stuff, uh, it, it's not gonna...

It's-- I, I'm pretty confident it's not gonna beat, um, ARC 2. Would an o5 be needed for it? I, I don't know, but o5 is not coming out this year is my guess. You know, it's gonna be coming out next year.

So I, I guess long winded saying is I, I don't know, but my guess is it's not gonna be beat for the next 12 months. And then V3, our durability estimate for that is three years. That's what we're aiming for is 36 months for V3.

That's kinda wild in this, in this timeframe. Who knows? That's our hypothesis. That's what we're aiming for. We told a big lab that our timeline for V3 was 36 months for it not to get beaten, and I won't say which lab it was, but we were basically laughed out of the room when, when, when we told it that.

Swyx28:55

Well, yeah. I mean, I think it's a cat and mouse game. You know, I think it's all towards a good cause, so it's up to them to prove you wrong and, you know, we'll all be better for it, right?

So...

Greg Kamradt29:03

That's, that's exactly how I think about it. And like, I'm even hesitant to say the words like beat ARC-AGI 'cause it's like we're not putting up this test as, as like a, like, "Look at us. Go and try to beat this thing."

It's like it is a tool to incentivize research. Like, that is why we put up a million dollars to see who can actually beat it, because our hypothesis is the thing that actually does beat ARC-AGI 2 has something worthy in it that can be applied to other types of domains.

And so in, um, ARC Prize 2024, we had over 40 papers submitted for our best paper award to try to beat ARC-AGI, and that's one of the most rewarding things for us because that's open research that's out there.

And so one of the things that came out of that was the whole, you know, test time fine-tuning adaptation for that. And so as test time compute was coming around, a form of that is test time fine-tuning, and one of...

and that was a huge proponent within ARC Prize of last year. As we think about what-- As I think about it, it's almost like I don't want ARC Prize beat per se. Yeah, sure, that's what's gonna happen, but I want cool research to happen because of ARC, and you end up performing well on ARC because of that research.

Alessio30:03

Yeah. Ideally, Mike didn't put up a million dollars just for somebody to win it, but in order to get more research out in, in the ground. I think that that's a great goal to have.

Greg Kamradt30:11

Yeah. Like tactically, people ask me like, "What, well, what am I trying to do?" Or, "What is..." like, "What's the goal?" It's like, well, once AGI is found, I think ARC Prize is setting itself up for a very good place to declare, to be the one to declare that AGI is actually here because someone's gonna have to do it.

And like maybe it might be like the Turing test where it kinda just happens and goes. I doubt that that's actually gonna be it. I think there's gonna be a before and after moment, and I would love ARC Prize to be in the place to validate that.

And then number two is once AGI is actually here, I would wanna look back at the learning path and research path towards that and say ARC Prize had a role in accelerating that progress.

Swyx30:45

That's a very good mission. I would say like-- But, you know, it's hard to declare a moment when AGI is here. I don't see why that's a worthwhile goal. Like, you guys are moving the, the goalposts, like V1, V2, V3, right?

And, and even OpenAI, I think in their current communications, they no longer say, like, at least not in the public ones, you know. They-- with whatever they have with Microsoft, uh, you know, it's like if we, like, make a hundred billion dollars in profit, that's AGI.

That's controversial, but at least publicly, they're saying that we're no longer talking about a binary, it's here, it's not here. There's, there's just levels, right? And we're like level two slash three right now.

Greg Kamradt31:19

Yeah.

Swyx31:19

Isn't that okay, you know?

Greg Kamradt31:21

Uh, I would, would have a healthy debate on that, and I would push back on that. So, like, for the first one, I agree with you. Any AGI definition that involves money has ulterior motives. I mean, simple as that.

Um, money has nothing to do with intelligence, right? Anyway, that, but that, that's a whole 'nother side concept from there. I would argue that AGI is here once an artificial machine can match the learning efficiency and generalization of an actual human.

And to me, that, that's a pretty binary place to go for it. I agree that there's gonna be spectrums and levels of that. We're already seeing it with the progress on Arc-AGI-2. They're showing non-zero levels of fluid intelligence.

Sure, okay, I accept that. But we will be able to look back at a point in which machines were basically able to learn faster or quicker or at the same level as humans.

Grok 4 Recap32:01

Swyx32:01

Yeah. Okay. Well, you're set up to declare that. Should we spend a little bit of time on, on Grok 4 since it's like-

Greg Kamradt32:07

Sure

Swyx32:08

... hot on everybody's minds?

Greg Kamradt32:09

Yeah, let's do it.

Swyx32:10

You were at the event. What, what was the vibe like? Just, just give us the on, on-the-ground review.

Greg Kamradt32:14

Yeah, totally. So I-- now I've had the luxury of attending two labs of livestreams here, and so for the OpenAI livestream in December and, and then Grok's here. They called us up... Today's Thursday. They called us up on Tuesday and they said, "Hey, we got a new model.

We wanna test it." We say, "Yep, sure. That sounds good." We have a kind of a standard testing procedure with them, so, like, no data retention, and the model that you-- that we test should be the one that's gonna go out to the public so that people can actually reproduce these results.

And they said, "Yep, cool. No problem." They gave us a little bit of credits to go for it, which is always very happily accepted, and then tested it. And one of the things that we do is we say, "What is the score that you're claiming?"

Because they do their own self-testing, and we wanna see what the-- what's their score so that we can go and validate it. And they said their public eval scores. Yep, we, we ran it on semi-private. It, it worked out, looked great, validated.

They were excited about it, and then they said, "You should come be in the audience of the livestream." We said, "Sure, sounds great." And so went over there, and the one thing I wasn't expecting is this was p.m.

on a Wednesday, right? And so, uh, I walked into the xAI place. There might've been two hundred employees there, like, buzzing, and the energy was high. Uh, there was like, you know, DoorDash deliveries happening all over the place.

Everyone was absolutely buzzing going for it. Um, and then we get into the livestream room, and they have a proper livestream setup. It's great. And then they were an hour, hour late. Like, not an hour late. It's like, I don't know what they were doing.

They were just, like, practicing, like, their slides, right? Which is totally cool. I think it was just, like, um, they were figuring out what they wanted to say, and it was just, um, not an ad hoc thing, but it was just, you know, off the cuff going from there.

And then they announced the results. It's great. I was sitting there in the audience with my laptop, like, getting the Arc Prize tweet ready and then click Enter right, right when I saw the slide come up and click Merge on the website push when we go from there.

Yeah, I mean, it, it was great. So we love when labs pull us in early to help them validate scores, um, 'cause that really gives kind of a public boost and a public, like, vote of confidence that Arc-AGI, like, has the merit to...

Like, they're saying that, "Yes, we trust Arc-AGI to help describe the intelligence of these models." We absolutely love that. So we're very thankful to the xAI team to giving us the-- giving us that access, for letting us come join the stream.

And then after that, right after the stream was over and the, you know, the cameras went down, there's a good chance to go talk to Elon. And so, uh, went up there and said thank you to Jimmy, uh, one of the founding engineers of, uh, xAI, and went over and thanked Elon, and then I gave him the V3 pitch.

And 'cause he w- 'cause he w- he wasn't as familiar with Arc-1 and Arc-2 as some other folks were, and so I quickly realized that and then pivoted over to the V3 conversation. And so I said, "Hey, um, we're coming out with V3 next year.

It's gonna be a series of a hundred video games." And he immediately did his, like, Elon, like... I c- I can't do it, but he did his Elon, like, eyebrow raise, like, right at me.

Swyx34:43

He's like, "Video games? Ooh."

Greg Kamradt34:44

Yes. Yeah. No, really. And because he had just got done in the presentation, they just got done how Grok made a video game. And he's like, "Okay, Grok making a video game one shot. Okay, cool. That's kind of impressive.

What Grok really needs to do is Grok needs to play the video game and then iterate on the game and then play it and then iterate and then play it." And so I knew that that would probably hit with him, and I said, "Hey, man, we're, I mean, we're making a hundred video games here that are gonna be easy for humans, hard for AI."

And he goes, "When's it coming out?" And I said, "Well, we're doing the developer preview next week." And then he raised his second eyebrow right at the, uh, right at that point. And so he, he was interested in it.

I th-- you know, he's a gamer himself, and so we're hoping to talk, talk to him more about it. He's excited, and so we'll see how it goes.

Swyx35:22

I'm always curious, like, you know, Elon's a, a unique figure in human history. I always figure, like, how much you actually have to deal with him versus, like, his people, you know? Like, actually, you know, you should be talking to Jimmy, not, not Elon.

But, you know, like, it's nice to have Elon, you know? Like...

Greg Kamradt35:37

Yeah. Well, I, I tell you, I mean, it, it was cool to have that experience with Elon just 'cause of his prominence in the field and in industry in general. Shake his hand, congratulate him on the Arc stuff here too.

But, like, if we wanna get xAI to, to come jump on board with us and, like, come along with this journey for us, um, Elon, like, wouldn't necessarily be the person necessarily who's gonna be doing the operational execution of getting that to happen, right?

Like, it's gonna be Jimmy and team and all those other folks. And so we're excited if we can have Elon's support. Either way, we're excited to have Jimmy's support and the rest of the team.

Swyx36:06

You know? Well, you know, he's, he's got a big, uh, checkbook that he can, that he can help to fund the nonprofit, and-

Greg Kamradt36:11

Yeah

Swyx36:11

... he has a history of that.

Greg Kamradt36:12

Yeah, yeah.

Swyx36:13

Any other takes from, from Grok 4? Uh, obviously doing, doing very well on all the evals that everybody's... Like, is, is it, is, is, is it just the new frontier LLM? Like, have they, have they come from nowhere and beat everyone?

It, it can we-

Greg Kamradt36:25

I mean, that-

Swyx36:25

Can we-

Greg Kamradt36:26

It, it-

Swyx36:26

Yeah

Greg Kamradt36:26

... I mean, that's what it looks like right now. I mean, I tell you what, I'm with everybody else. When you look at the benchmarks, it's like, okay, great. Let me get this thing inside a cursor, and what's the deal?

Let me get this thing inside of... Let me go in the chat. Let me see, like, actually how it, how it does, right? So the vibes seem great right now. I'm excited to see more adoption that comes from it.

You know, I think it, it still needs to be investigated and researched as to the reason, uh, the RL paradigm. Like, you saw that they, they 10X'd the RL that they put on top of Grok 4 for that.

And so how does that bleed down on actual performance and where are the gotchas and where are the nooks and crannies in the corners and where's the spiky intelligence actually spike? I think there's still a lot of open questions for it, but either way, right now we're seeing fricking awesome performance.

Swyx37:03

Yeah. Yeah, yeah. Um, I, I, I would say that too. I would-- I think the cursor thing, they, they did mention maybe they have like a coder-specific model, like a dev-specific model that might be coming out separately. So maybe I would hold off on, on the coding side of things, especially, you know, it's, it's, it's a huge model.

Greg Kamradt37:18

I heard rumors, I heard rumors that, that Grok doesn't wanna release the coding model until it's better than one specific other lab out there. So they're gonna wait and see when it's actually better for the-- to, so they can have that marketing point.

Swyx37:31

Yeah. I mean, sure, that makes sense, you know, but also, like, uh, the fact that they can credibly do it, you know, kind of come from a standing start i-is really, is really, I think, commendable. I think it's-- there's, there's a lot of stories to be told from inside of X, like how they did this because, you know, there's just a lot of data they have to get, and you can't all get it from Twitter 'cause like, God knows Twitter is not the best source of information .

Greg Kamradt37:52

Yeah.

Swyx37:53

So-

Greg Kamradt37:53

Well, I think it's really-- if you think about the ablation study about what does it take to get a good model out there, I mean, obviously there's people, there's money, there's compute, there's, you know, there's data, there's all these other things, and it's super interesting to have AGI labs out there that are frustrated with their external performance, and there's other labs out there who are very happy with their external performance, and what's the difference between the two?

And they all took different routes to get there. So like you're talking about coming out of nowhere, I mean, more or less. I mean, they, they've been at this for not, for not too long. They're literally building the ship as they're doing it.

And so it's very impressive for what they have, and I think other labs may look at this and be like, "Man, what is going on there that we want that type of success too?"

Swyx38:29

I know. Zuck's, Zuck's over here looking and going like, "Mm-hmm," you know. These, these X AI researchers looking very in-in-interesting.

Greg Kamradt38:35

Yeah. Well, you won't be able to blame Zuck for not taking a move either. So at least-

Swyx38:40

Yeah, yeah

Greg Kamradt38:40

...

Swyx38:41

Yeah. It's, it's a, it's a, it's a wonderful time. Cool. Well, thanks so much. Uh, any, any other last questions or thoughts, Alessio o-o-or Greg?

Alessio38:47

No, this was great. I'm excited to play the games when they come out.

Outro38:47

Greg Kamradt38:50

Cool. Yeah, absolutely. My only call to action for folks is we want agent builders, and so we're gonna come out with three different games. I would love if you could build agents however the heck you want, using whatever tools you want, RL based, LLM based or whatever it is.

We're gonna have about a $10,000 prize pool and put up some money for it. It actually may be more valuable. We're gonna like push all, like on our every single social channel that we have, push every single top-performing agent.

So I would love if you guys would participate with that, uh, call out to the audience for that.

Swyx39:17

Yeah. I mean, I'd love to hack on it. I don't know if I can-- I'll do-- I'll get anywhere, but you know, I'd love to hack on it and, and promote it. Um, but I think it's a, it's a really fun cause.

Greg Kamradt39:25

Beautiful. Love it. Thanks for having us on.

Alessio39:27

Thanks, Greg.