LALatent SpaceMar 4, 2025· 37:38

How Claude Plays Pokémon was made

David Hershey of Anthropic explains how he built Claude Plays Pokémon, an agent that uses Claude to play Pokémon Red, revealing both surprising capabilities and stark limitations of AI. He describes the simple tool-using agent loop: Claude sends button presses, receives screenshots and RAM state, and maintains a knowledge base and 30-message conversation history with summarization. Key challenges include Claude's poor spatial awareness—it famously walked up and down outside Oak's lab 12 times—and its tendency to hallucinate game knowledge. Hershey notes that newer models (Sonnet 3.7) required stripping away band-aid prompts, as simpler instructions yielded better performance. The project also showed emergent behaviors: nicknamed Pokémon made Claude more protective, and the model even wrote meta-commentary about its own mistakes in its knowledge base. While the current run is stuck in Mount Moon after 52 hours, the biggest hurdles remain vision and navigation, not prompting tricks.

  1. 0:00Intro
  2. 1:21Origins
  3. 6:28Architecture
  4. 15:29Token & Memory
  5. 20:30Knowledge Gaps
  6. 25:48Transfer & Limits
  7. 32:50Highlights

Powered by PodHood

Transcript

Intro0:00

Alessio0:02

Hey, everyone. Welcome back to another Latent Space Lightning Pod. This is Alessio, Partner and CTO at Decibel. There's no Swix today. We got a special co-host, Vibhu, which if you're a part of the Latent Space community on Discord, you've definitely seen.

Um, welcome, Vibhu, as a co-host. First time.

Vibhu0:19

What's up, guys?

Alessio0:21

And then we had David Hershey from Anthropic today, who's the person behind Claude Plays Pokémon. It's funny, I saw we at first DM'd about playing Magic: The Gathering together in, uh, in SF.

David Hershey0:32

I really-

Alessio0:32

And then people are like-

David Hershey0:33

On all of the different nerd angles you can get me. Uh, glad-

Alessio0:36

And then people were like, "David is the person doing this," and I was like, "Okay, I'll, I'll, I'll DM him," and then, um, yeah, it was cool. We already had a, a touch point. So welcome to the, to the show.

This is our second Anthropic episode. We had Eric Schwanz from the Suite agent-

David Hershey0:52

Yeah

Alessio0:52

... before, so welcome.

David Hershey0:53

Thank you. Glad to be here. Excited to talk Pokémon.

Alessio0:56

Yeah. So let's give a little background on this. So Sonnet 3.7 came out a couple weeks ago. I don't know. Time goes by so quickly.

David Hershey1:04

Monday this week.

Alessio1:04

Monday. This week? I don't know, man. It feels like two weeks ago. And then you had this Claude Plays Pokémon thing that kind of went viral where if people remember, there used to be this thing called Twitch Plays Pokémon, where people could go on Twitch and kind of type in the chat and then basically, like, figure out what the next action that the emulator would take is.

What you've done instead is given it into Claude and basically have Claude figure out how to walk through it. I'm looking at it right now. So far, it's been stuck in Mount Moon for 52 hours . Poor guy.

Origins1:21

Alessio1:32

It probably met a 15,000 Zubats. So yeah, let, let's talk about what gave you the idea for it, uh, kind of the origin story and then we can go through the implementation.

David Hershey1:42

Totally. Yeah, so I actually started working on it in, like, June of last year for the first time. And for me, so I, I work with customers at Anthropic, and I just, like, really wanted to have some way for myself to be able to, like, experiment with agents, like, in a real way.

Some framework, some harness, where I could actually just, like, go to town and try some different things and see, see what actually worked to get Claude to do, like, pretty long-running tasks in general. And so I, like, had that in one hand, and then I was like, "Okay, what is the thing that will make me the most addicted to making this work?

Like, how will I grind the hardest actually trying this?" And, uh, Pokémon was, like, a pretty clear answer. Someone else at Anthropic had actually, like, tried once to hook it up, so I had a little bit of, like, the, the shell of what I needed to, to actually put it together and to, like, kick off what, what became an obsession a little bit in the, uh, coming months.

So yeah, like, I played with it in June in this, like, just trying things out. This was like Sonnet 3.5 came out in June of last year, which is when I started kicked it around. It was very good, but it, like...

You know, you could see, like, kind of signs of life, but, like, not much really happened. Um, and then ever since then, as we've released new models, it's sort of been, like, the way that I get to know one of our new models a little bit, right?

So we released the new version of Sonnet 3.5 in October and, like, used this to, like, really kind of see, like, what's it better at. And it got better. Like, you could see it start to, like... It could get out of the house somewhat reliably, which was not always true, and it got a starter, and it, like, even named it sometimes.

Like, it was, like, doing stuff. Not great, but, like, it, it could move. Um, along the way too, like, I'm just like, we have a Claude Plays Pokémon Slack channel. Like, I'm sort of just, like, giving people updates.

So over time, as I'm, like, posting gifs and progress updates, like, I'm... It's, it's, like, slightly growing in popularity of a cult following internally of people who are somewhat interested. But then like, uh, you know, a couple weeks ago, I was bashing an early version of Sonnet 3.7, and it just, like...

You could just tell it had, like... It was a little different. It's clearly not still good, as you said at the top. Like, it's, it's in Mount Moon for its 50-somethingth hour. This is a little bit worse than average from what I've seen so far by now, but, like, this is, like, you know, about on brand.

It doesn't really have a great sense of direction. Uh, it's pretty bad at seeing the screen, stuff like that, but, like, it plays the game, you know? Like, it gets Pokémon, it catches Pokémon. Like, it caught its first Pokémon, it got out of Viridian the first time.

Like, a whole bunch of stuff happened for the first time where, like, could squint and see a thing play in the game. And yeah, like, posting updates obviously internally. It was very fun. Like, people were just, like, kind of going wild at the fact that this was actually happening finally.

Um, and it was, like, entertaining enough that I could kind of see it. And the other side is, like, we kind of just, like, got finally a sense that this was, like, an actually useful way to measure what was going on with this model.

You know what I mean? Like, there's one thing that it's, like, fun and fun follow along, but, like, internally, like, I think we got more of a sense that, like, you could actually use this as a bit of a measuring stick for what's going on in the model.

I've spent you don't want to know how many hours I've spent staring at Claude Play Pokémon. I've so much-

Alessio4:34

Oh, sure

David Hershey4:34

... I have to have seen and read, like, millions of words that Claude has generated in the course of playing Pokémon over the last eight months. So, like, you can kind of get a feel for, like, what's actually going better, what's it getting better at, and that kind of thing.

And with this particular release, like, I think the fact that it got this much better at this kind of reflects a lot of things that we wanted to be true about the model to begin with. And, uh, and those sort of lined up.

We're like, "Okay, maybe this is, like, an interesting way to actually tell people about what's going on here for a crowd that maybe doesn't, like, quite know as much about software engineering and all the other ways we've told people about agents in the past."

Alessio5:06

Yeah. Uh, were there any other games that you consider? To me, it seems like Pokémon is good because it's like, you know, isometric. You know, it's kind of like flat, so you can score, and it's... It doesn't have too many hidden facts about objects, you know?

Kind of like everything is described. Did you consider anything else, or was Pokémon just kind of like, by far and away, the first choice?

David Hershey5:26

I didn't, but it's mainly because, like, Pokémon was the first game I ever got as a kid, right? It's... This is, like, purely coming out of my own nostalgia. Um, but also, like, the Twitch Plays Pokémon, like, uh, it was also something that I cared a lot about w- a decade ago or whatever that was.

Um-

Alessio5:40

Please tell me it's not a decade ago.

David Hershey5:42

I think it's actually-

Alessio5:43

I would

David Hershey5:43

... a decade ago. I'm sorry. Yeah, painfully. Um, and-

Alessio5:48

11 years ago.

David Hershey5:49

Yeah.

Alessio5:50

That's insane.

David Hershey5:51

That's, that's-

Alessio5:51

February 2014?

David Hershey5:52

Yeah.

Alessio5:53

That is nuts. Okay.

Vibhu5:54

And Pokémon Red is 20 years ago.

Alessio5:57

Oh my God.

David Hershey5:59

2025, at least.

So yeah, I, it, for me it was that. Like since then there have been a lot of people in front of me like, "Ooh, we can do this, we can do this, we can do this." Um, I think there's like a lot of fun things you can do.

Pokémon's actually really nice because like if you don't do anything for five seconds, like there's typically not a consequence. By the nature of like doing inference on a model every, like snapshot in time, it's actually a pretty good game to be able to do this with.

But, um, but yeah, it was mostly just like, uh, my, my love for Pokémon coming through here.

Architecture6:28

Alessio6:28

Um, you put together a very nice architecture diagram. Do you wanna screen share that so people on YouTube can follow along, and then we'll put it in the show notes-

David Hershey6:36

Heck yeah

Alessio6:36

... um, if you are just listening? Um-

David Hershey6:38

You got it

Alessio6:39

... I know that Vibo had a bunch of questions on the, uh-

David Hershey6:42

Yeah

Alessio6:42

... on that too.

David Hershey6:43

Yeah, let's do it.

Vibhu6:43

Very, very straightforward question. Basically, can we just double-click into all of it?

David Hershey6:48

Yeah, yeah, yeah. It's easy.

Vibhu6:49

I found it off Twitch, and like no one was talking about it, so I started sharing it around and I, I lost the original source, but basically everything in here is like pure gold. The memory is a little interesting.

Um, but yeah, if you wanna just go through high level.

David Hershey7:02

Yeah, you got it. Yeah, I, I wanna like preface that I do not claim this is like the world's most incredible agent harness. In fact, like, uh, I explicitly have like tried not to like hyper-engineer this to be like the best chance that exists to beat Pokémon.

I think it'd be like trivial to build a better computer program to beat Pokémon with Claude in the loop. Uh, this is like meant to be some combination of like understand what Claude's good at and benchmark like, and understand Claude alongside a simple agent harness.

So what that boils down to is this is like a pretty straightforward tool using agent from my perspective is how I would frame it. So at the end of the day, like the core loop is just like having a conversation that rolls out.

Um, and it's essentially like you build the prompt, including like everything we've had up till now. You call the model, it sends back some tool use typically. You resolve those tools. And then talk about summarization, but like basically some, a few different mechanisms to maintain the information you need to do something long-running inside the context window.

Um, so like what this boils down to is like when you think about what an actual prompt looks like, it rolls out kinda like this. You've got tool definitions, which describe three tools that I'll get to in a second.

Uh, a short system prompt that's like pretty boring. It basically tells the model how to use the tools. And like there are about six facts about Pokémon that I give it, and, uh, like a few corrective things that I've seen it do like really horribly wrong that I'm like, "Hey, you might wanna consider doing this a little bit better."

Uh, but it's like n- really not a lot of system prompting going on. Uh, we have that knowledge base, which referred to, I'll talk about. This is the main way it stores like long-term concepts and memories as it's operating over time.

Um, and then the, the bulk of things is this conversation history, which is, it's like a chain of tool use. There's no like user interjections at all for the most part. So it's like go, and then the model uses the tool, and then it gets a result back, and then it uses another tool, and it gets a result back.

So, uh, pretty straightforward. Uh, feel free to like cut me off too if you've got questions along the way, but otherwise I'm gonna keep rocking.

Alessio9:13

Yeah, yeah. Go ahead.

David Hershey9:14

Cool. Uh, okay. So most of the money of this is just like in the tools themselves. When you think about what's going on,

it's really like it can press buttons, and it can like mess with its knowledge base, and that's about it. I'll talk about Navigator separately 'cause that's like a patch for how it actually can deal with some of its vision deficiencies.

Um, using the emulator is just basically like execute a sequence of button presses. It'll say like press A, B, left, right, whatever.

It gets back a screenshot and a screenshot overlaid with coordinates of the game. Uh, these coordinates are used for this Navigator tool that I'll describe in a second, but it's just basically like help Claude get a slightly better spatial sense of what's going on on a Game Boy screen.

I've been through it a lot-

Alessio10:01

And does that come with the... Sorry, does it come with the emulator, or are you adding this in?

David Hershey10:06

I add that in.

Alessio10:07

Okay.

David Hershey10:08

I have, uh, somewhat extensively reverse engineered Pokémon Red by this point to like extract roughly every bit of possible information from it. I don't use most of it, but like I have essentially everything you could know about the current state of the game, I have exposed programmatically to be able to tinker with it at this point.

Vibhu10:24

I was just reading this diagram like, yep, you just get what spaces are walkable based on what's stored in RAM. And I'm like, "Oh, you definitely reverse engineered this little emulator."

David Hershey10:31

Uh, yeah. The good news is we also released Claude Code this week, uh, if you saw that. And that has been... This would all not be possible without the help of having Claude also go figure out how to do all of this for me, 'cause I could have done it, but there's a lot of like tedious, "Here are addresses in memory, map that to a Python program" that I had no interest in doing .

So thank goodness for Claude Code.

So yeah, it gets the two screenshots. It gets like a small blurb of state, which I read straight from the game.

There's a lot of this here. Actually, like funny enough, the thing that matters is location. Claude will like pretty aggressively hallucinate that it succeeded in transitioning between zones if you don't like tell it it did not. Uh, this just comes down to like literal vision issues and, and so like most of the patching of extra help I've given it have been like attempts to make it so that it could still play despite not being very good at seeing Game Boy screens in particular.

Um, and then it gets like a handful of like reminders. This is, this reminder says a decent amount of work, but it's like

things like, yeah, remember to use your knowledge base occasionally. And we tell, we tell if it gets like stuck, for example. So if you detect that it like hasn't moved in 30 spots or 30 time steps. I once saw it see like a red box on the screen that was like the doormat and think it was a text box and spend 12 hours pressing A overnight to try to clear the text box, which you see that happen once and you add in some, some helpful reminders to not do that.

Alessio12:04

How much knowledge does the model have about the game itself?

David Hershey12:08

Yeah.

Alessio12:08

You know? So for example, types, right?

David Hershey12:12

Yeah.

Alessio12:12

Does it know about types, weaknesses, and things like that, or how much are you trying to put into it?

David Hershey12:17

Yeah, if you go to quad.ai, like it, it, it will tell you about like some stuff. I have not yet decided if the knowledge that it has about Pokémon is helpful or harmful towards it playing the game. Um, like half of the time when it's like, "Oh, I know this about Pokémon," it then like uses that to hallucinate something.

So for example-

Alessio12:37

Mm.

David Hershey12:37

... at the beginning of the run on Twitch, you saw it like go out of the lab and see like this NPC in the bottom of Pallet Town and be like, "It's Professor Oak. I found him." And it's like very much not Professor Oak, but like the fact that it has like indexed on this concept is like a little...

It's stuff like that, that it's like unclear to me where it is. But it clearly has some information about it. Uh, there's like a million game guides about Pokémon sitting on the internet. It's unsurprising that like there's a decent amount of information there.

I don't really give it a lot of extra information. It, it picks things up. I watched on this stream the other day, like it tried to use Thundershock on a Geodude and it failed, and it's like, "Hmm, I forgot about that.

That does not work." And so like clearly there's like it knows some stuff. It's not perfect. It picks some stuff up as it goes through the run. Ideally for me, like I think it's just interesting to see like what it actually learns as it's playing, so the more it does that is the more I'm like actually interested in it.

Alessio13:31

Yeah. The, one of our, the score members, Nanchang, he had a good question about the sense of self.

David Hershey13:36

Yeah.

Alessio13:37

Like sometimes it gets confused who is the actual playable character in the, in the scene. Like how, how do you steer that?

David Hershey13:44

Uh, yeah. I think like sometimes it gets confused, uh, can be applied to many things in Claude playing Pokémon, uh, in particular when it's trying to like look at the screen and understand what's going on. So

I have like attempted to prompt it all sorts of ways. Like, "You are at this exact coordinate, and you're in the middle of the screen, and you're wearing a red hat," and things like that. And like that's all neat, but Claude doesn't particularly understand like the middle of a Game Boy screen and a whole bunch of concepts like that, which means like you can prompt all around everywhere, but like this kind of like spatial awareness and where something is with respect to something else is something that Claude's still just like not great at in its current incarnation.

So one of the side effects is it sometimes loses track of who it is on the screen and thinks there's something else there. I will keep trucking through this. So I hinted at this like other tool that I give it called Navigator, and this is just like the only other patch that I have for the, the vision issue.

So Navigator, basically what it does is like Claude can say it wants to go to one of these coordinates, uh, that we provide in the screenshot, and then we like automatically press the buttons to get there. Uh, it has to be something on the screen.

Like I'm not trying to let Claude just like navigate a whole map by asking to politely. But one thing you'll notice if you run it without this tool is if like Claude wants to get from one side of a wall to another side of the wall, it like happily just tries to walk through the wall repeatedly 'cause it doesn't quite have the concept of like what's between it, and I've spent a lot of time like prompting around this and it just like isn't, it's just not, it's one of those things it's not very good at.

So in order to make it somewhat fun to learn from Claude playing Pokémon at all, we use this Navigator tool, which like helps it actually get around a little bit better.

Vibhu15:29

So since we covered a bit about the different tools, the prompting, and the strategies-

Token & Memory15:29

David Hershey15:33

Yeah

Vibhu15:33

... I'm curious how many tokens all this is using. Like there's a part to conversation history and truncating-

David Hershey15:38

Yeah

Vibhu15:38

... parts of the messages in state.

David Hershey15:40

Yeah.

Vibhu15:40

But like, yeah, at a high level, how many tokens is this using and then can we kinda go into where those are coming from, what's being truncated?

David Hershey15:47

Yeah, you got it. Uh, when you like think about the, the prompts here, um, essentially like every step, something that looks like this gets sent. So like if we just go through what each of these looks like, um, everything in the system prompt is probably like 1,000 tokens, pretty small, like a handful of paragraphs.

Knowledge base, I let get up to like 8,000 tokens, right? So, uh, I put some like arbitrary cap on it so it doesn't go to like Claude will write, put a whole bunch of BS in there if you just let it keep writing stuff.

So like the cap helps constrain it to like try to think about what's actually important a little bit. And then the conversation history, I have it like kind of finicky, but it basically rolls out, um, 30 messages. That's actually like something you can tune.

I've tuned it to be 30 messages about like the best performance I've gotten. And so what that means is it basically like uses a tool, get a response back, use a tool, get a response back. It's allowed to do that 30 times, and then at that point it triggers the summary, which takes that conversation history, summarizes it, makes it the first user message, and then we kind of roll back out again.

So the bulk of the tokens end up being in the conversation history once it's its longest. In fact, like this, the bulk past that ends up being these screenshots, which are scaled up a decent amount to, to fit in.

I do actually like, I allow it to see a number of s- the previous screenshots, but not all of them because you start, like it ends up being a ton of context if you let it see like even 30 turns worth of screenshots, so I, I trim out a few.

But that's where the bulk of the actual tokens are. So in practice, this rollout ends up like at max ending up around 100,000 tokens, I think is where it, is like the, the longest message you ever send to the API on one of these turns, and it will, it will fluctuate in like summarization depending on state of knowledge base, probably between like 5,000 and 100,000 tokens.

Vibhu17:45

And is that like per action state of the game? And roughly, do you have like a high level ballpark estimate of how long this would, how much and how long it costs to run this? Like let's say people wanna compete-

David Hershey17:56

Yeah

Vibhu17:56

... and, um, yeah, yeah, like how much would this be?

David Hershey17:59

Um, I think you'd really want to think about running this as a side project in terms of the impact on your personal wallet and how much you care about Pokémon. It's not clear to me that without the blessing of Anthropic I would have decided to take on Take on this project for my own wallet's sake.

Uh, especially if you wanna like experiment and like try 10 different things. I mean, it's, it, it's costly. I don't know, like I, I haven't spent a lot of time on the exact number. It's not that hard to estimate if you-- like I just told you a bunch of numbers, you can kind of back it out.

Uh, but like I think to like do a lot of experimentation, there's like at least thousands of dollars of tokens being consumed, so it's not a, it is not a, uh, a, a cheap rollout. Yeah. But yeah. In the scheme also of how some people use tokens, it's not terrible.

Alessio18:49

How many turns are you keeping in memory before you summarize?

David Hershey18:53

It's 30 right now. Yeah. I've tried more and less. I think like one thing you see a lot when you talk to people building agents is there's like some effective context length that actually like has the model be the smartest.

Um, and that seems to vary slightly model by model but, but for this model, for whatever purpose, like this 30 message worked better than 20 and better than 40. So, uh, kind of plot in between those that it worked pretty reasonably.

Alessio19:21

Yeah. Does that change based on location? Like how many would you wanna give it to get it out of Mount Moon?

David Hershey19:26

Uh-

Alessio19:26

So let's say, "Hey, we gotta, we gotta bring Claude home. We can't let him stay-

David Hershey19:30

Yeah, yeah, yeah

Alessio19:31

... on Moon for another 57 hours."

David Hershey19:33

I, I actually am not sure it does. Like I, I've pa- I've tried posting, like when you have a ton of screenshots, like 20 or 30 screenshots at, at a time be able to see, and it's like not obvious that like that temporal concept is actually super relevant, relevant to it.

And again, this is just like...

Trust me, as someone who has spent like a lot of hours obsessing over this, uh, you can try to prompt Claude a lot of different ways to understand how to navigate better, and anything short of telling it exactly what to do does not improve its like actual navigation.

It's just like not a skill it's great at. It's like good enough to, to like random walk its way through some of the complex mazes, and in like e- good, easy areas, it's pretty good at bopping around. But yeah.

I think I, I could tell you if there was like a way to prompt this slightly different that, uh, would navigate better. I, I would believe there is something, but it is not like, uh, it is not an easy lift.

Alessio20:30

Yeah. Yeah, ask the-- I just asked Claude AI right now how do you get through Mount Moon in Pokémon Red. It does have, it does have a plan, but I don't, I don't, I don't know, I don't know if it's the right, I don't know if it's the right plan.

Knowledge Gaps20:30

David Hershey20:42

I have seen it come up with a lot of answers to that question, and most of them aren't-

Alessio20:46

Yeah

David Hershey20:46

... right. Uh, this is part of the pain. When I talk about I'm not sure if its knowledge is better or worse, like you see it-

Alessio20:51

Right

David Hershey20:51

... usually fix it. Like, "Oh, I know the exit is on the eastern wall," and it just like spent 12 hours trying that. Um, and I-

Alessio21:00

Yeah.

David Hershey21:00

Yeah, it's like unclear to me that, that we're actually not just like harming it by having it think it knows the answer.

Alessio21:06

Yeah.

Vibhu21:06

I, I think that's the interesting part, right? Like you don't want it to just know the answer.

David Hershey21:10

Yeah.

Vibhu21:10

Like the model clearly knows a lot about the game. There's like EV, IV maxing.

David Hershey21:14

Yeah.

Vibhu21:14

Pokémon was very, very extreme. But like if that's what you wanted, we could just hook it up to a knowledge base. Like hook it up to a guide of you know, how to beat Pokémon Red. But-

David Hershey21:23

Yeah

Vibhu21:23

... the, the interesting piece here is actually like, can it figure out what to do without just memorizing the path through?

David Hershey21:30

That's exactly right. Like that's part of why, um... You know, and it-- I, I don't know. Part of what I've realized putting this out in the world is people will draw their line of where purity is anywhere on this spectrum.

Like is it, is this cheating? Like yeah, maybe. Um, who knows? Um, like frankly, like I don't particularly care. The, the main insight that I have is like when we put this out, like you learn a lot about what the model's good and bad at by staring at it, and, and that's kinda what I like about it, so.

Vibhu21:58

Evaluating the model is kinda separate than your emulator and how it can use an emulator, right? Like we can always improve those things. I'm curious, um, as you switched from 3.5 to 3.7 in sort of reasoning models, were there any degradations there?

Like did it, did it kinda get worse at anything? And was the prompting somewhat consistent? Like a lot of what we've seen with different reasoning models is like you kind of prompt them differently, right? You tell them what to do-

David Hershey22:22

Yeah

Vibhu22:22

... let them figure it out. But, um, yeah. Any, any insights there?

David Hershey22:28

Yeah. Yeah, that's a good question.

One thing that's nice about 3.7 tonnet is with like this hybrid reasoning model. So like it kinda can do the old thing and the new thing, and it's actually pretty good at just like being an out-of-the-box model and having this like thinking mode where it can spend time reasoning.

So I, I didn't like really run into any like serious degradations. The one thing I'll say is like literally every model that has come out with Pokémon, like the, the main change that I have made to this agent is deleting prompt stuff.

Like there's a whole bunch of like band-aid-y prompt stuff I've added in the past. It's like trying to like steer it away from doing a lot of the things that it got horribly stuck doing in the past. And as the models get better, I've found that just like making sure it's as simple as possible and giving them as much sort of like free reign to try to solve a problem as possible is useful.

A- and like the way I think about this is, I'm like less confident over time that I understand exactly how a model is intelligent, right? Like it's capable of all of these like ridiculous things. It does PhD-level stuff in some ways and like is unable to screen, see a screen as well as a four-year-old in other ways.

But like my confidence in like exactly what I need to tell it to do to be smart at playing Pokémon is actually like really small right now. If I tell it, "This is the way you need to solve this problem," that might not actually be the best way for 3.7 sonnet to solve this problem.

It's like just different than I am in terms of how it thinks about these things. Uh, I found that just like kinda like pulling some of the unnecessary instructions where I tried to like use my intuitions about what would make the model better out of the prompt over time is the thing that just like sort of consistently, as models got smarter, gotten more juice out of this.

Alessio24:14

I was watching the stream yesterday or the day before, and It was a very tense battle. I think they were like down to like two HP each, and like the opposing Pokémon like missed a scratch or something, and it didn't die.

And like you could tell, like Claude was like, "Wow." It was like very dramatic, and I was talking about the game. How, how, yeah, is there any thought being put into like trying to have it more... Like do you prompt it to be more rational to let it know that it's not real life, that it's a game?

It's like it, it feels like it gets very distressed when they're actually, the Pokémons are actually gonna die.

David Hershey24:51

It's funny. They, they, um... It knows it's Poke- ... Like it's like you're playing Pokémon Red, like it does know that and it has a sense of that, but it clearly has some attachment. I'll tell you a fun story.

We tell it to nickname its Pokémon now. It will occasionally do without it, but it's like more fun if it nicknames its Pokémon. So that's like in the prompt is like, "It's fun if you nickname Pokémon, you should consider it."

And one thing we found when we started doing that is it got more protective of the Pokémon it nicknamed. Like, it's pretty obvious, like when it catches a Pokémon, now that it has a nickname, it will like go heal it right away if it's hurt, and that did not ever happen before, which is pretty...

Like so there, there's some cute little things, cute quirks about Claude who really wants to protect its, uh, precious nicknamed Pokémon, which is great.

Vibhu25:36

So I will say it's kind of normal. Like, like when I was five playing Pokémon Red and, you know, I had two HP and it missed a scratch, that meant everything.

David Hershey25:45

That was existential. I agree. I agree completely.

Transfer & Limits25:48

Alessio25:48

How about skill transitioning? So one question that I had, so you're playing Pokémon Red, right?

David Hershey25:54

Yep.

Alessio25:54

Say you want to play Silver or Gold next. Um, ha- have, have you thought about how models can kind of learn from these games and like store these learnings and then use them again in the future? I'm sure it's not part of the project today, but curious your thoughts.

David Hershey26:09

I, I thought about it only a little bit, which is like, I think there's some like interest- When you actually read one of the knowledge bases that it has gained, like on some of the longer rollouts when they're good, like there's actually some like pretty decent tidbits about how it should act and try and do things and like some of the ways it succeeded in.

And actually, like one of the things that's most unique about 3.7 Sonnet that I've seen is like it will have like meta commentary on what it's good at and bad at in its knowledge base. Like I, "I misperceived this thing, and so like I need to be careful doing that again," you occasionally see show up there, which is, um, which is pretty cool.

So like I could imagine there being some way to like translate that knowledge base from one game to another. Uh, I think my knowledge base is frankly like kind of kludgy of an implementation right now. Like it's like more or less a Python dictionary that's appended to the prompt.

And I think like you could, you could find better ways if like your goal is to transfer across games and things like that to manage a knowledge base that Claude can actually like, uh, use more well in different scenarios.

Um, but there's definitely pieces there that like I think it would get, be off on a better foot on the next Pokémon game if it had that, or even if like I were to restart the stream, it would like have some, some tidbits that it would probably like, uh, speed up if it like had access to things that it learned in the past that it's interesting.

Alessio27:30

Yeah. Yeah. I, I always think of that in card games, you know. Like you have the idea of like tempo in a card game, and it's like, you know, it's the same magic as it is in, you know, Star Wars, Flesh and Blood, all these different things.

David Hershey27:41

Yeah.

Alessio27:42

I feel like w- games is similar, where like learnings you get from Pokémon you can bring over to similar kind of like open-world games in a way.

David Hershey27:50

I think it's also like particularly interesting for some of the things that are like how Claude learns how to play a game in general, where it's like-

Alessio27:57

Mm-hmm

David Hershey27:57

... pressing too many buttons at once is a bad idea. Like I lost-

Alessio28:00

Right

David Hershey28:00

... track of what's going on, that kind of thing.

Like definitely is stuff that it has learned that is like interesting in a meta way, uh, that it's like hard to give it that sense of self necessarily in training, I think sometimes. Like it's hard for it to know like what it's good and bad at in some scenarios, but it's interesting to think about how it can learn across things.

Vibhu28:19

Well, like, uh, some of this also is due to a simulator, right? So a lot of what it's learning is how do I use a simulator? What am I good and bad at? But the model internally should know quite a bit about Pokémon, right?

Like if you've played Pokémon, going from Pokémon Red to Emerald to Diamond, having played the first one doesn't help you that much in the second, right? You kind of get the general concept. You get what types are good against other types.

And the model, model knows a good bit of this, right? But it's still interesting to show this is more so like it's, it shows that knowledge bases kind of help with understanding how to use the emulator, right? Like it struggled, and then it figured it out.

So even though it knows Pokémon, it's like this thing cannot learn how to use an em-

David Hershey29:00

Yeah, which is pretty cool. That has been like part of what's been fun, seeing Miles make progress on this thing.

Vibhu29:07

I had a bit of a follow-up question to the last one with Alessio. So if people want to blow thousands of dollars and wanna, you know, improve this a little bit, is there anything else that you'd want to see done, whether that's like improve emulator, try different stuff?

Is this just anything that like anyone watching this, you'd, you'd kind of hint them towards what you'd want to work on, what they'd want to work on?

David Hershey29:28

Yeah, no doubt.

If I had to guess, like the, the biggest lift that exists around this is probably

something around the memory, which I don't think is like hyper-optimized right now. The nice thing about the memory is like it's always in the prompt. Like it's, it doesn't go away. Like some- sometimes if you leave it up to Claude to try to like read and load and save to memory bases, like it, it will underutilize it or forget things.

But I, I think there's probably something there. I will say all of the many, many hours I've spent tweaking around the edges of this thing, nothing quite does it like a new model though. Like fundamentally, I think the limitations right now are like some smarts things.

Like I've seen, uh, and I, I mean this in the kindest way, but I've seen a lot of people in Twitch tell me about ways that they could fix the navigation capabilities with a better prompt. The, uh, people would be welcome to try, but I would guess that would be like a somewhat fruitless avenue .

I don't think-- I think it's just not very good at understanding. Uh, the first time... I'll give you a very quick anecdote, which I think is like my favorite for like why this is particularly hard. I have this clip of Claude leaving Oak's lab and being like, "Great, I left Oak's lab.

Now I need to go up to the north end to go to Route One," and it just like hits up on the D-pad and goes straight back into the lab. And it's like, "Shoot, I'm back in the lab.

I need to leave," and it hits down, and it's like, "Great, I'm out of the lab. Now I can go up to Route One," and hits straight up, and it just like goes up and down 12 times. And it's like you're not, you're not fixing that with a prompt.

It just literally doesn't get it. It doesn't understand . And so it's pretty hard to make like little around the edges changes that like make a huge, huge difference.

Alessio31:09

Yeah. I mean, I've always been fascinated by the fact that Twitch Plays Pokémon actually beat the game.

David Hershey31:15

Yeah.

Alessio31:16

From a... You just look at it and you're like, "This cannot possibly work," because you have people trying to sabotage it too in the chat.

David Hershey31:22

Yeah.

Alessio31:22

Not everybody's trying to solve it. What, what... So I, I just looked that up. It took sixteen days and seven hours for Twitch Plays Pokémon to be red. How, how close do you think we are to a model that can beat it in less than sixteen days?

Um, and do you think it needs like some core, like model really big jumps, or like do you think it's like we're close?

David Hershey31:45

I think it... I think there is model stuff, at least from Claude. Like I am confident there's model stuff that needs to happen for it to be like really capable. I, I have like four spots in the game stuck in my head as like I think there's literally no hope it's gonna get through that.

So I think there's like a gap that's mostly around like its ability to like see and navigate and remember visually like what's going on that I just don't think is, like we've figured out yet. So to me, that's like a pretty big gap.

I do expect like I, I think it's gonna keep getting better. Like I have no reason to believe that this is not just like a fundamental like ability to scale, learn, and understand problems thing that I think is getting better as we train models to be more capable at sort of these like long horizon tasks.

Like I actually do think this is like a pretty reasonable proxy of that, and I think it will continue to get better for a little while. I don't know if there are like affordances around images and videos and stuff like that that we need to figure out to make it work.

It's like unclear to me if that's true or not. Um, but yeah, I think we have a little ways before we can beat the game in sixteen days. I do not have a lot of faith that the, uh, current stream is gonna, gonna beat standing in Victory Road in, uh, thirteen days.

Highlights32:50

Vibhu32:50

What's been your favorite moment from like building this to thinking of the idea to just seeing it play? Any, any like major highlight?

David Hershey32:57

Uh, I think like the, the hypest I have been is, uh, when it beat Brock the first time where I was just like, you know, I've been doing this for eight months, and then like a few weeks ago, like I kick off a run, wake up the next morning, and it's like, "Oh my God, oh my God."

And, and it was the, the other good thing about it is like I woke up at eight AM and I checked my... I, I have it send me updates to Slack. Um, this is like ridiculous things, but, um, it's like literally like about to start the Brock battle.

Like I open up my phone, it's like, "Oh, this is like happening right now," and it was like a pretty hype way to start a day. I think that was my, uh, my highlight. I have a lot of like other cute things, like some of the cute nicknames it's done over time and things like that are, are endearing, but, but that was like the peak hype for me was like, "We beat a gym leader."

Like, "We've got a badge." Like, "Claude's doing it," you know?

Vibhu33:45

A bit of a follow-up. So I noticed that you mentioned it, it eventually started beating multiple gym leaders. Were these all the same run? Was it different runs? Was it-

David Hershey33:55

Yeah. I have like... I've... The, the run that you saw that's like on the graph we put out alongside, like in our research blog, is like a single run that I have watched like get through at least Surge's gym, and then it got a little past that.

And the reason that that's where we stopped reporting is because that's like the physical amount of time that occurred between when I started it and when we watched the model. So that's like, uh, that was a very hyper, uh, hyper up-to-date graph on, on the best run we had.

So-

Alessio34:24

Awesome. Um, I know we're running out of time. My last question is, are we gonna work on Magic, on Claude Plays Magic next? Or maybe we can do like the Magic Arena intro challenge.

David Hershey34:34

Yeah. Uh, funny story. There was a project I did right before I joined Anthropic that was like training, uh, an open source model to like slightly be better at picking draft or cards in a draft. Like I, I was training it on like the seventeen lands data that exists to like learn how to, how to pick cards out of a, out of packs a little bit better.

Uh, and I, I did talk about that in my interview to get hired at Anthropic, so, so I-I've put time into this. I'm ready. I am ready for that project too. That I have that code sitting around as well somewhere.

Alessio35:07

Nice.

David Hershey35:07

You're really getting in all my nerd, her nerd ML slash, uh, slash gaming hobbies here.

Alessio35:13

Yeah. No, I'm ready. I don't know if you're planning on open sourcing any of the Pokémon stuff, but if you wanna work in open source on the Magic stuff, I'll be happy to, to collaborate.

David Hershey35:23

Awesome. We've talked about it. I don't, I don't know yet what the plan is. I think there's like a certain amount of like this is not my day job that I have to figure out how I want to, uh-

Alessio35:31

Yeah

David Hershey35:31

... deal with that, uh, we'll see.

Alessio35:33

Yeah. Um, awesome, David. Any parting thoughts? Anything people have missed?

David Hershey35:39

No. I, I think like, uh, the one thing I do like to drive home when I-I've been talking about this is like I really do think like this is just demonstrating like a thing that is going to make agents better with this model, you know?

Like this is a very fun way to see it, but like I think the thing is that it like has some ability to like course correct, update, and figure things out a little bit better than models have in the past.

And even if there's like stuff it's dumb at, like it tends to have an ability to like power through it in a new way. And so I think what's exciting to me is just like I think there will be some real world stuff that comes out of this model once people play with it, and I'm pretty excited to see like, uh, how people take the skills we put on display a little bit here or, or lack thereof in some cases and, and figure out how to turn them into actual agents that do stuff.

Vibhu36:22

I have a quick last question on that actually. Is there any, uh, guidance or any way that you like quantitatively measure the evals of this system? Like a lot of it is vibes, a lot of it is how far it gets, where it gets stuck.

But like are there, are there any lessons or any specifics about how you measure how it actually does?

David Hershey36:40

So I've done a lot of like little small tests of like put it in this scenario and see what it does. But I... like frankly, the best test I have is just like run it ten times on diff- on this configuration and like see how quickly it progresses through milestones of the game.

I mean, it's the best thing about games, right? Like it's why games are such a useful thing. There's literal like benchmarks of gym badges that are moments of progress in a game, which are like ways to evaluate what happens.

And so I think like how quickly it's able to make progress is actually a pretty reason- or a reasonable like eval if a slightly expensive one to calculate. It's an integration test, not a unit test.

Alessio37:15

Um, awesome, David. Thank you for joining. Thank you, Vibhu, for filling in on the host side too.

David Hershey37:20

Yeah. My pleasure. Thanks for having me, guys. I appreciate it.

Alessio37:22

Awesome.

David Hershey37:23

Good to see you.