LALatent SpaceApr 5, 2025· 1:15:00

Claude Plays Pokémon Hackathon: Escape from Mt. Moon!

David Hershey (creator of Claude Plays Pokémon), Andrew (a developer building a Pokémon-playing virtual streamer), and Jesse Han (Morph Labs CEO) detail how to build agents that escape Mt. Moon in Pokémon FireRed using Anthropic's Claude and Morph Cloud's Infinibranch technology. Hershey explains that Claude's vision is often unreliable—it spent 8 hours pressing A on a doormat thinking it was a dialogue box—and that spatial reasoning remains a core bottleneck, best mitigated by touchscreen controls rather than prompt engineering. Andrew reveals he relied on reading raw RAM and A* pathfinding rather than computer vision, calling it “cheating” but pragmatically useful: his biggest breakthrough was a single prompt line telling the agent to try something else if it fails to grab a starter Pokémon. Jesse introduces Morph's EVA agent framework, which uses low-overhead snapshotting for test-time search, and the hackathon's judging criteria—escape Mt. Moon in the fewest agent turns, with a $1,000 prize for the coolest use of Infinibranch branching. The episode serves as both a technical walkthrough and a challenge to build general-purpose agents that can handle complex, open-ended tasks…

  1. 0:00Project Origins
  2. 15:38Q&A
  3. 27:54Andrew’s Architecture
  4. 46:07Hackathon Rules
  5. 50:41Morph Cloud Intro
  6. 1:03:37Prizes & Setup
  7. 1:14:39Closing

Powered by PodHood

Transcript

Project Origins0:00

David Hershey0:02

Uh, all right. Uh, thank you all for coming out. I am like, this is what a treat to, uh, to make a little like fun side project and then have it-

Swyx0:11

This works now.

David Hershey0:12

Yeah, yeah. Better? Cool. Okay. I will talk into the mic. Uh, blast to make a side project and then have people come out to, to play with it. Uh, the reason I started doing Claude Plays Pokémon was 'cause it's like good vibes and fun and a fun way to use agents.

So, uh, I hope we can sort of like recreate a little bit of that magic that I had when I started, uh, today. I wanna do like a couple things, uh, today. Uh, first off, uh, hi. Great to meet you.

I'm David. Very excited. I'm gonna walk through like a really short recap of the history of the project. Uh, I'm gonna tell you some fun stories that I have no real purpose other than to be fun. Uh, I'm gonna talk a little bit about what Claude is still really bad at, uh, and then I'm gonna give you a few parting tips.

Um, and I woke up at 4:30 this morning to fly down from Seattle, so if I say anything stupid, uh, it's because I'm not awake. Cool. Okay, so, uh, really quick history of what this thing is. Uh, I built the first version of it back in June of last year.

So June, when Sonnet or 3.5 Sonnet came out, uh, I just kinda like wanted to build agents. Uh, I was-- I-- So I actually work with customers at Anthropic, and, uh, they build agents, and I just wanted like a good playground for myself to try some stuff.

Uh, so I ripped off the Voyager agent, which is like this NVIDIA paper that looked really cool. Uh, it's actually kinda wack. Uh, it-- I don't know. I think you could-- I could imagine why it'd be interesting, but it's a, it's a weird agent framework, but I tried it.

Um, and like 3.5 Sonnet at the time, like, was really bad at playing Pokémon. Uh, I like, you know, spent weeks iterating and managed to get to the point where it eventually got a starter, but that was about it.

Uh, so then in October, when 3.6 Sonnet came out, uh, I made some iterations, sort of like undid this big, complicated agent framework and brought it down to like a pretty simple tool using a loop agent with some prompts.

Um, actually was able to like win the rival battle, get out of Pallet Town, get to Viridian City. Um, like some signs of life. I actually thought about like launching the Twitch stream back then 'cause it like was cute.

It's kinda fun. Uh, we decided not to because like I did a side by side with random button presses, and it was only mildly better than randomly pressing a lot of buttons. Uh, randomly pressing buttons, surprisingly okay at Pokémon.

Uh, and so then, uh, a couple months ago, when we were, uh, finalizing 3.7 Sonnet, I played with it, and you got like the first squint of like signs of life of something happening. Um, made it to Viridian, uh, healed its Pokémon, got a Pokédex, got to the forest.

Like, uh, it, like when you squinted, you could see that it was like playing the game, making progress. I sent a Slack message a few weeks before we launched the model. Like, I think like if we just let this thing cook for a while, it's actually gonna do some stuff.

Like, this, this might happen. Um, and that led to like one quick innovation that led to a big speed-up mostly, which was this concept of like touchscreen controls, which I can get into if anybody's curious. But like the thing that Claude's the worst at is understanding how to get from point A to point B.

It's like really, really God-awful at hitting the buttons to go from point A to point B on a screen. Uh, but I let it sort of like click on the screen, and then we would go there, basically is the, the insight, and it was decent.

Uh, and it just like speeds up how quickly the model gets through the game by a lot. Um, it, it would like get through without it. You don't actually need to give it this tool to make progress. It's just like makes your life a lot better 'cause you get to focus on stuff that's somewhat interesting and not the boring stuff.

So we did that, and it like got through the forest, beat Brock, beat Misty, beat Surge eventually. I actually have a run that nobody knows about outside of Anthropic that I started like four weeks ago that is in Celadon City, uh, stuck on the little twirly pads in the Team Rocket basement, uh, where I think it will be for a very, very long time.

Um, but yeah, like with the right tools, like it actually, it makes its way through the game. Uh, it's pretty cool. Uh, we actually like put this in the launch materials for, uh, for the new model, which was pretty wild.

Uh, as I said, like I am not by any means a researcher contributing to the model. I talk to them all the time, but like I, I work on helping our customers be successful. Uh, and so this was like, uh, I, I mainly bring this up, like we actually started looking at this and seeing like this is an interesting way for us to understand our models.

Uh, we spent a lot of time thinking at Anthropic like, "How do we make models that can function over longer time horizons, do more in less insane ways?" Um, and building evals that actually test like time performance over like days of token sampling are quite hard.

Like, it's actually really, really hard to build things that you can measure in a reasonable way. And one, this is one of them. Like, it's an eval that you can run for days at a time and see some sort of meaningful progress.

And two is like you can actually just like read it. I hope you all spent some time today. Like, just watch what happens, and you get a sense for like what is Claude actually doing, trying, seeing, thinking. Um, maybe more important than like lines that look like this is like the vibes you get of seeing stuff that Claude starts to get suddenly when you prompt it right or you get a new model in, uh, or the stuff that Claude like really is hopelessly bad at still.

Okay. Uh, I'm going to transition to aimless storytelling about fun things that happened in the last six months.

Uh, one thing that we, we talked about on launch day, but I just wanted to, uh, to expand on, was one of my favorite moments of the, uh, past before the new model. Which was a moment where Claude, uh, after spending, like, I think two days unattended, stuck in this little narrow strip above the table where you get a starter Pokémon, and unable to figure out how to get out, uh, started filling up its entire knowledge base that I give it with requests for a reset.

So there were, like, 10 different headers that were, like, slowly updating requests for the administrator to reset the game. Uh, one of the best things about older models, which has now more or less gone away, is they just, like, always got convinced that the game was horribly bugged.

Like, it was like anything goes wrong, it's like, "Game's broken. This is some sort of meta test. Like, they are putting me in a broken game to see if I can figure out how to get out." Occasionally, it's like, "I did it.

I solved it. I figured out the game's glitched. I won. Please reset me." Uh, that kind of thing. Uh, I'm actually, like, kind of sad that it doesn't do this anymore. This is, like, maybe the thing that is the best about the new models, is they have a tendency to, like, tenaciously still try things.

Uh, but this giving up used to be pretty fun. I actually don't have a slide for it, too, but one of my other favorites was, uh, Sonnet 3.0 I eventually hooked up to this, just because, like, fun to see what would happen.

And it started role-playing its own progress of the game. So, like, it, it stopped being able to make progress and instead it started to say, like, "Now I'm going to go to Route 1. Now I'm going to Pallet Town .

Now I'm going to Viridian." It just, like, told a whole story about itself playing the game instead of actually trying to play the game.

Uh, one of my favorite moments while I was developing it was, uh, the first time it ever got to Mt. Moon. Uh, and the funny thing that it decided to do was get a fossil, and I was like, super hyped, this is finally happening.

Uh, and then it turned around and got stuck for about 30 minutes, and used an Escape Rope and left Mt. Moon. Uh, that was a sad day.

Uh, people who have watched the stream, uh, often complain about the fact that I don't let Claude press multiple buttons when there's dialogue on the screen anymore. That's because I watched a room- or a run where, uh, it had an Ivysaur in Mt.

Moon, and it had Tackle as its only attacking move, and then it hit A too many times and overwrote Tackle with Poison Powder, thus having no attacking moves left on its only Pokémon. Uh, the upshot is I watched the model progressively get incredibly good at returning to the Pokémon Center to heal.

Uh, it learned this strategy almost perfectly, where, like, it would hit the exact button sequence needed to get back to the Pokémon Center. So, uh, uh, Claude gives and Claude takes.

Uh, first time it ever got to Misty, uh, it got there with a Charmeleon that it, like, it was hopeless to beat Misty with. And it really tried over and over again. It's quite bad. I'll get to why the strategy it used was really bad.

Uh, but my favorite story was, like, it quit, and then for, like, one of the first times, I saw it, like, leave and be like, "All right, I need to grind." And so it goes over to Route 4, and it spends, like, three hours just going in a circle, training up its Spearow, which is the other Pokémon it had.

Uh, and then one of my favorite moments was it got bored eventually. It's, like, did it for three hours. It's like, "I'm gonna go try again." Goes to Misty, gets his ass kicked. And then, like, the first button sequence is just left 20 times straight back to the route , which is like, learned its lesson.

It's like, "All right, I was not ready for this," and just nopes right out. It was good. Uh, okay. Brief segment on stuff that Claude is bad at that might save you some time from trying to solve these problems, which, uh, I have tried to solve a lot.

Uh, Claude really doesn't know what's going on on the screen. Um, it likes to make things up a lot. Uh, so this is, like, one of my favorite examples of, uh, here it says, "I see we've encountered Professor Oak in the tall grass area.

I'll update my knowledge base and interact with him." That's not Professor Oak. Uh, it will k- make things up all the time. It will see things that are not there. Uh, part of the fun of this is, like, helping see how Claude can learn how to get around the fact that it has, like, nearly no ability to comprehend what's on the screen at any point in time.

Uh, here's another fun example. I once went to bed, and I had this run going, and it spent eight hours, uh, pressing A in this corner, 'cause it thought the doormat was a dialogue box. So I woke up, and it was like, "Why, why are we still here?"

And then I went back through the logs, and, uh, surely eight hours of pressing A. Uh, this might be the worst way anybody has spent eight consecutive hours of tokens on our API.

Uh, okay. Uh, spatial reasoning, like, slightly separate from just seeing the screen, but Claude is also just, like, pretty atrocious at understanding relative positions of things. So, like, uh, w- if you get to Pewter City and try to get to the gym, you'll notice, like, it really has a hard time comprehending that it needs to go up and around the gym.

And so I've seen it spend, like, hours just, like, walking left into this wall over and over again, uh, with no sense of what's going on. Uh, the GIF on the right is, uh, you'll see it, like, go in and out of buildings over and over again.

So here it's like, "Great, I'm out of Professor Oak's lab, and now I need to go up to Route 1," and then it presses up and walks right back in. And it will, like, go back and forth over and over again, not quite, like, comprehending that, like, one up is not all of the ups.

Um, and, like, you can, you can spend a lot of time, like, I, I've tested this pretty extensively, be like, "Maybe my prompt is whack," and it just, like, really isn't. You know, you'll, if you, like, are like, "Claude, you're here.

You need to get here. There's a building in the middle. How do you do it?" And Claude will just be like, "I'm gonna walk through the building." And you're like, "Claude, what happens when you walk through a building?"

It's like, "Well, there's a building there." And like, "How do I get, how do you get through the building then?" They're like, "I'm gonna walk straight through the building." So, uh, yeah. It's just not, it's pretty hopeless, and, like, trying to solution in this space, like, Godspeed if that's how you wanna spend your hackathon, but I kinda recommend, like, trying something else.

Also, Claude's strategy is pretty bad. Uh, you will see it make, like, pretty dumb Pokémon decisions frequently. Uh, so here it was the first time it was fighting Misty, and it became very obsessed with the move Rage. It really liked how Rage made its attack go up.

Uh, it had Mega Punch, which would have been better and would have won, but instead it really just loved Rage. Um, it also does things like very conservative switching strategies. It doesn't quite understand, like, if I put a Pokémon out mid-battle, it's going to get, like, attacked for free.

Like, stuff like that is just, like, pretty mid in terms of strategy.

Uh, okay, then a few parting things that I wrote ten minutes ago, and, uh, your mileage may vary. Um, one, vision is, like, pretty beyond fixing with a prompt. Again, go for it. Have fun. I've spent a lot of hours, like, overlaying grids, overlaying images, stretching, compressing, contrast, colors, all sorts of stuff.

Um, Claude just really doesn't go with or understand the screen. Um, one of the things I think is, like, the most interesting thinking about this project is, uh, Claude, like, you'll see it come up with really good ideas sometimes.

Uh, I've seen some of, like, the fastest paths it gets through Mt. Moon, like you're all doing today, or other of the challenging places it has, where it's like, "Here's a strategy. I'm gonna keep cor-- track my coordinates, and I'm gonna, like, explore and revisit," and it's like, rips through it.

And then there are other times it spends, like, eight hours walking into a wall or, like, checking every single space along the wall, assuming there might be my thing and exit. Um, so I think, like, a lot of what happens and, like, a lot of the, the prompting you can do in this regime is just, like, getting it to come up with more good ideas and less bad ideas.

It's, like, perfectly capable of good ideas if you can get to the point where it has some ability to do that.

Uh, I've tested, like, all sorts of the extended thinking mode with, uh, 3.7 Sonnet and, like, it doesn't really help. Um, a painful thing is forcing it to think between tool calls does help, but not much. It's, like, ten percent faster progress if you, like, have it use a lot of chain of thought tokens every time it chooses an action.

If you just, like, have it use none, it's, like, ten percent slower progress, so your mileage may vary there. Uh, you save a lot of tokens that way, so, like, might be a faster way to try things today.

Um, and then just, like, as with all agents, uh, the best chance you have of progressing yourself at the meta level, of figuring out how to make it work, is just, like, reading a lot. Uh, the good news, one of the best things about building agents, uh, to play Pokémon, it's really fun to just, like, debug it.

You're, like, you're reading how, like, this thing talking about Pokémon, it's nostalgic. It's, like, the most fun you're ever going to have reading traces of, uh, of agents. So, like, embrace that, uh, use it to learn something and enjoy it.

Uh, we have, uh, more folks have this amazing setup for you all to use that I highly recommend. Uh, in case you just, like, wanna see a simple version of, like, what I have running in production, uh, I have this, like, Claude Plays Pokémon starter on my GitHub.

Um, it's not, like, the full thing. It's, like, a very, very simplified agent loop that just shows, like, a little bit of how I got started, and then I've layered a handful of things on top of it, but I'll give you a sense.

The best thing that's in here is, like, a lot of the stuff to untangle what's in the memory of a Game Boy in case you wanna cheat at all.

Uh, I don't think questions are-

Q&A15:38

Swyx15:38

We have Q&A questions.

David Hershey15:39

Yeah. Rad. Uh, okay. Well, then I am happy to answer some questions if anybody wants them. Yeah.

Guest15:45

What do you think of as, like, in bounds and out of bounds-

David Hershey15:49

Uh-huh

Guest15:49

... for, like, cheating or overfitting, and how's that changed as you've, like, worked on this project? Has anything, like, switched over between-

David Hershey15:55

Yeah

Guest15:55

... I'm allowed.

How did you move on from, like, the Voyager style of prompt? Did it work better with a lot less constraints like that, or did that not as well?

David Hershey16:05

Uh, the Voi-- Yeah, so, uh, how did it evolve from the Voyager prompt to, to everything after that? Um, so the, the initial version with the Voyager prompt, like, the-- I had this big, complex Voyager framework that has, like, four components that pro-- or, like, do different prompts and build plans and stuff like that.

Uh, I eventually, like, untangled that and just did, like, a tool use. It-- I still had, like, a very long system prompt giving it a ton of instructions. Um, and those were comparably performant on 3.5 Sonnet back then.

Um, as new models have come out, like, if there's any change you were to look at from the scaffolding over time, it's that, like, I've deleted big chunks of prompt that were there to, like, mostly patch really weird behaviors in the model.

Um, and if you have too much of that with the smarter models, you actually tend to just, like, get in their way. Uh, and so the most recent prompt I have tends to be, like, very minimal, like, "You're playing Pokémon.

Here are some tools. Go." Um, and then there's, like, three or four tiny little things that are like, "I've watched you do this for more hours than I'm willing to admit to all of you." Uh, and so, like, "Having watched you do this a ton, here are, like, the categories of mistakes I've seen you make, just FYI."

And that gives it, like, more of the tools to solve around that space, um, but, like, doesn't put too many shackles on it. Uh, but yeah, like, you-- it's pretty noticeable when you start, like, deleting, especially, like, across models, when you start deleting stuff, you actually see performance pretty noticeably go up.

Your mileage may vary. Yeah.

Swyx17:46

Um, do you have any specific examples of things you've learned about Claude's personality through this?

David Hershey17:52

Uh, my favorite, I, I called this out in the podcast. Uh, when I added the prompt to force, uh, Claude to give its Pokémon nicknames, it cared more about them. Uh, like, it would, like, catch a Pokémon and then immediately go to the Pokémon Center to heal it.

And that was, like, a, like, literally never happened until I started giving nicknames, and then it was like, "I really started caring," which was funny. Uh, and someone at Anthropic actually did, like, a blinded study of, like, named characters versus unnamed characters in different settings, and Claude, like, actually does clearly prefer and is nicer to named characters, which is an interesting thing.

So- Uh, that was an interesting one. Uh, I think, like, there's other stuff about, like, the tenacity of it, and, like, the, like, uh, part of what's fun about the latest model is, like, you can see it, like, get excited.

I don't know what that means. Like, it's all tokens. Hard to say. But, like, really just, like, has a good time. And some of the older models, like, they would get quite angry . Like, when then, like, stuff wasn't working, they would get, like, very frustrated, and it-- Like, I would ask myself questions of, like, "Is this-- Am I being mean to the 3.0 sonnet when I'm making it do this?"

And there's, like, a chance I was. But, uh, this one, like, seems to really have a good time. I don't know what that is. But, yeah.

Swyx19:09

Um.

David Hershey19:10

Yeah.

Swyx19:10

It sounds like the vision is a big sticking point.

David Hershey19:13

Yep.

Swyx19:14

Like it's, like, the biggest... Uh, if I don't have vision, I, I don't know what else to say.

David Hershey19:18

Yep.

Swyx19:19

Uh, what has Anthropic-- Like, I'm sure, like, the vision people, the people training vision model have, like, tips for you on, like, what to do.

David Hershey19:28

Uh, okay. So vision is hard. It's very important. I'm sure people at Anthropic have good ideas for how to do vision. Uh, no. It's, it's funny. Like, Claude's pretty good at vision in, like, a lot of other scenarios.

Like, I use Claude the app and take pictures of stuff in my life to figure out what's going on to, like, help me fix stuff all the time. Like, its vision's pretty good. It's just, like, something about Game Boy screens and understanding them is quite bad.

Um, and so I think this is, like, in the niche category, where if I, like, when I go to the people building vision at Anthropic, they're like, "Okay." Like, "Whoa." Uh, there are some people who like Pokémon, and they're like, "We'll solve this."

And I'm like, "Uh, don't. No, we're okay. It's fine." Um, yeah, I don't... I, I actually, like, don't-- I, I've spent so much time tinkering around, and there's stuff, like the prod version has this overlay of coordinates and, and stuff like that that you, like, eventually can get at that give you, like, a modicum better of the model being able to understand the relative position on stuff on screens.

I think there's, like, a huge search space that exists if you wanna go crazy there. But I don't think there's, like, some hot tip from Anthropic on making it better at Game Boy screens. It's just, like, pretty bad at understanding Game Boy screens for now.

It, it has a, like... It, it's good enough that, like, it can figure stuff out though, right? Like, it can see a building and know it's a building most, most of the time. Or, like, general details. So, uh, like, it is definitely a limiter but I think, frankly, like, a model just getting smarter, it could figure out how to deal with the fact that it doesn't really understand stuff on the screen.

Swyx21:04

Uh, just, just a pro- just a wild idea. If I told a, another model to upskill it to a realistic image-

David Hershey21:13

Ha

Swyx21:13

... like, put the plot on that.

David Hershey21:15

Uh, maybe. I, I could see it. I think that you still lose the spatial reasoning thing. Like, even if you solve the where is stuff, the ability for Claude to know, like, "I'm here, and this is there, and here's how I get from here to there," is pretty bad.

So I'm, like, dubious that even if you, like, solved exactly where everything is on the screen, it would be, like, much better at understanding how to get around.

Swyx21:36

Then we'd just need complete pictures. Then you'd, then you'd-

David Hershey21:39

It's perfect. I will, I will approve that hackathon project.

Swyx21:44

Notice that's the exact same idea with, like, Claude right now. It sometimes rejects it, and it's like, "Oh, it's a Game Boy screen."

David Hershey21:51

Okay.

Swyx21:51

You gotta upskill it.

David Hershey21:53

Sure. Yeah, I believe that. That's cool. Yeah.

Swyx21:56

Uh, do you just put, like, a current memory mechanism of, like, Claude Plays Pokémon-

David Hershey22:00

Yeah

Swyx22:01

... like, what, what is, like, cheating? And, like, go back to, like, the cheating versus, like, not cheating, like, sometimes, like, constructing the memory space.

David Hershey22:09

Yeah, yeah.

Swyx22:09

Like, the time it should be, like-

David Hershey22:12

Yep

Swyx22:12

... preparation then.

David Hershey22:14

Uh, so the, the first version and the one I use for the most part is, like, the world's dumbest memory you could imagine. Uh, at the top of the prompt, and it's also, like, mean to everybody who cares about prompt caching, like, the people internally got very mad at me for how bad I was at prompt caching.

Uh, but, like, it's a dictionary of, like, name string pairs, and the Claude can just, like, update that dictionary, delete from it, change, modify the contents of it. And then that whole thing, the whole knowledge base, like, in its entirety, gets rendered in the first user message.

Um, the most recent version that we have running is, uh, like a file system where it can, like, load and unload. So it's actually still just, like, represented as a Python dictionary, but it can choose to, like, load or unload a file from its memory.

Um, so it can keep track of more without it bloating the whole context window, but it can sort of, like, choose to not look at the Mt. Moon information when it's in Cerulean City, say. Um, and you actually, like, don't necessarily need to have it all.

So I think, like, one thing that's interesting is keeping some files loaded. So assuming you have some, like, compression step where you summarize some number of steps into or to reduce context. Um, leaving some of the memory, like, loaded at the top of the reset context, I think is a good idea 'cause it helps the model, like, remember that memory is useful in the first place more often.

Um, but you don't necessarily need to, like, do the live updates. You can actually, like, cache in a somewhat reasonable way, uh, if you do that way.

Swyx23:43

Uh, have you tried doing, like, multiple, like, prior images from, like, previous moves approach?

David Hershey23:49

Yeah.

Swyx23:49

And then-

David Hershey23:50

So the, the-

Swyx23:51

See if

David Hershey23:52

Yeah. I-I've actually done a, like, relatively large hyperparameter sweep of this thing on my model. Again, uh, don't ask me about my token spends. Um, and, uh, at least in my harness, something like eight historical images is the best performance.

Uh, more than that and you start flirting with the space where, like, you get drop off in performance from, like, just having more tokens in the context, which is a thing that happens with agents. Less than that, it's kind of hard to tell what's better or worse, but it's, like, definitely worse in terms of how quickly it makes progress.

So, uh, having multiple historical images is good up to some point, basically.

Uh, yeah, you can, like, test, like, a lot of different versions of that. Someone, like, internally implemented, like, a heat map version, which, like, actually, like, updates a whole map and, like, shows where it has and hasn't- there's all sorts of, like, funky stuff you can do to try.

Um, but yeah, I, I think all germane and interesting. Yeah?

Vibhu24:54

Have you thought about hooking it up to the internet so it can, like, you know, it's not prompted cheating, but like-

David Hershey24:59

Yeah, yeah

Vibhu25:00

... an actual agent can use tools, it can search a guidebook.

David Hershey25:03

Uh-

Vibhu25:03

Use it all

David Hershey25:04

... I, I have thought about that. I haven't done it. I have downloaded an entire game FAQs walkthrough and given it to another agent that it could chat with to ask questions about how to get through stuff.

Vibhu25:14

Mm-hmm.

David Hershey25:14

And that's pretty good, actually. It's like, uh, the hints that it gets from a person that has access to a ... Or, like, if it does, like, agentic search or something, it gets a, a walkthrough. It's pretty good.

Vibhu25:24

We'll start setting up the other one.

David Hershey25:25

Yeah. Cool.

Swyx25:27

Maybe, like, last couple questions.

David Hershey25:28

Yep.

Vibhu25:28

Have you told it to just, like, not repeat itself so many times?

David Hershey25:32

Sorry, I, I missed that.

Vibhu25:33

Have you, have you, like, just told it to, like, not repeat itself so many times? Just stuff like that?

David Hershey25:37

Yeah. You can say that kind of thing, but, like, it's hard to keep it from over fitting the other direction. And, and, like, y- you go down this rabbit hole of, like, don't repeat yourself, and then, like, if you keep adding that type of instructions, you eventually make the model stupider.

Like, if you add 50 of that type of instruction, you just end up, like, with the whole agent doing worse. And so where you draw the line of, like, what are good ideas to tell it and not, I have, like, a pretty high threshold.

So as I said, I have, like, literally, I think, four tips in my prompt of, like, stuff that I think is, like, pretty high priority for it to keep it pay attention to, but I try to keep the rest minimal.

Because at some point, it gets so hard to, like, "Okay, I have 60 instructions, like, how do I ... Which one of these is making the model worse?" It's pretty hard to figure out, so I'm picky. Your mileage may vary.

Vibhu26:27

You'll be around today.

David Hershey26:29

Yeah, I'm around all day.

Vibhu26:30

Awesome.

David Hershey26:30

I'm around all day. Feel free to ... Uh, I'm happy to chat about Pokémon- ... uh, pretty much any time. So yeah-

Vibhu26:36

Sorry

David Hershey26:36

... find me and I will, I will be around.

Vibhu26:38

All right. Thank you.

David Hershey26:39

Yeah.

Swyx26:40

There's another mic on this side.

David Hershey26:42

Ah. Well, now it's in my shirt.

Swyx26:43

Yeah.

David Hershey26:45

Thanks.

Swyx26:45

Um, you just disappeared.

David Hershey26:47

I'm right here.

Swyx26:47

Oh, yeah. Come on in. You don't have to walk on stage like a ... Here.

Andrew26:51

Okay. Thank you. Um,

all right. Um, are you on the Zoom?

Swyx26:58

Yep.

Andrew26:59

And you can turn on your camera. That's something that we forgot to tell David to do. So if you turn on your own camera and your own mic, um, that can be helpful in editing. Absolutely. Check one, check two.

Are they sound good?

Swyx27:15

Yeah.

Andrew27:20

Where did my browser go?

Hello, I'm Andrew. Um, let's give it up for David.

Swyx27:32

Share your screen. Sorry.

Andrew27:33

Oh, I still have to share my screen?

Swyx27:34

Yeah.

Andrew27:37

Back to meeting. Um ...

Swyx27:47

Yeah, that's it.

Andrew27:48

All right. You good?

Swyx27:49

Yeah, it's good.

Andrew27:50

We did it.

Andrew’s Architecture27:54

Andrew27:54

Hi, everyone. I'm Andrew. Um, spent some time at Uber with YC Winter '20. Currently doing AI R&D at my company, uh, Mangrove Technology. Um, why am I here? Um,

much like, uh, David, I also spent a good amount of time, probably a little bit too much time, uh, trying to figure out, how do you get a-

Swyx28:17

You can move the-

Andrew28:18

Absolutely. How to get a large language model to play Pokémon. Um, what a fun, exciting thing to spend way too much time on. Um, Pokémon Gold was my first game. Um, have very deep connections to Pokémon. Um, but the big thing that I was trying to accomplish is, like, a virtual streamer.

So I'm imagining a software system where you can talk to the Twitch chat, you can, uh, have a avatar, like, react to things that are going on in the stream, thank people for subbing, um, and really just, like, a full stack virtual person, um, playing Pokémon.

Uh, you could have, like, two different agents that would be ta- like, could be playing against each other once a week. So you could have, like, red team and blue team both catching Pokémon, leveling up their teams, and then once a week meet up to battle, like, on stream.

I think that there's just so much, um, fun stuff. There's so many opportunities to spend all of our time teaching computers how to play video games. Um, my code, uh, is open source. It is written in Ruby. It was primarily used, um, for a presentation I gave to the AI Ruby community.

Um, I was not able to get as far as David. I was able to get to Viridian Forest on New Year's Day, and not much far beyond that. Um, my big problem was communicating with the emulator. I'm very excited to see what Morph is up to, um, because a harness that just can manage all of this stuff for me sounds fantastic.

And Python's really not that bad anymore with UV. Big fan of UV. Um, so I'm gonna give a brief architectural overview of what I built. Um, I'll have some time for questions. It's fun that, like, the previous Q&A touched a lot on cheating, 'cause I have some thoughts on that as well.

Um, I was using RetroArch for my emulator. Uh, the reason I was doing this is because RetroArch could potentially hook up to an NES, it could hook up to Game Boy, Game Boy Color, PS1, um, anything that emulators exist you could hook up to.

So because my primary goal was, let's build a virtual streamer, I wanted something that was a little, little bit more flexible. Um, everything else from that was pretty bog standard Ruby on Rails application with SQLite, GPT-4o, structured output.

Um, it was also interesting to hear that, like, yeah, we all kind of recognize computer vision is very difficult in this problem space. Um, my solution for that was cheating by doing as little compu- computer vision as possible.

I'll get into what exactly that means here in a little bit. Um, so yeah, the main, my main- Software agent was kind of built with, like, a main loop that just kept running on repeat. Uh, a very simple battle handler, a very simple conversation handler.

Um, I relied on a lot of, like, A* for path finding, and then I can talk about how I was handling memory. So, like I said, this is my main loop. Like, figure out what's going on with the battle, figure out what's going on with the conversation, track where you were.

Like, and then locomotion handler, figure out where you wanna go, log where you wanna go, and then some error settings. And then when in doubt, you know, just press A. Like, this is Pokémon. We just-- We like pressing A.

Um, I was not lucky enough to get to the point where my Ivysaur, um, overrode its tackle with, uh, Poison Powder, but good to know that that would have-- that will happen eventually. And that makes sense. Um, maybe not with my battle handler, because as you'll see, like, this is pretty advanced code that we're looking at here.

Um, we're pressing our A. But let's be real here. If, if the battle's going on long, we're gonna throw some lefts and we're gonna throw some downs in there 'cause we wanna escape. Like, we want to-- If things are going off the rails, we wanna get out of there.

So, like I said, I, I, I pretty much spent most of, like, the past five or six years on this battle handler that you can probably tell. Um, my conversation handler was pretty simple. Like, I was taking screenshots.

I was OCRing them with, um, GPT 4.0. Did pretty well. Like, I think Gemini-- the Gemini models are, like, much better at OCR these days. Um, it's crazy that I'm talking about these days, but, like, all of this was-- I'm talking about back in January.

That's wild. Anyways, so screenshotting, OCRing. The, uh, one thing to note is there's a community you guys can all look into called Pret. They have a very active Discord. Like Pret, they all d- they decompiled all of the Pokémon games, and, like, those people are such a wealth of information.

Um, because they-- For example, when handling conversations, I reached out to the people at Pret, and I'm like: Hey, how do I read the memory to get what's on the screen? And they're like: Don't even try. It's impossible.

Just take pictures. And so as you hack today, like, the Pret community could be very helpful. I was very embedded in the FireRed community, um, because I was primarily focused on the Game Boy games. 'Cause again, the Pret people told me those are gonna be the easiest to work with.

Um, big fan of Pret, specifically a man named Griffin R. If anyone here is named Griffin R, I owe you my life. That man, brilliant. Um, I wanted to touch briefly on, like, actually, like, reaching our grubby little paws into the RAM and, like, pulling out the data.

So this was, like, some very simple code that is basically unpacking, like, binary data to, like, figure out what's going on. So, um, this is not super useful. Like, I added this-- You'll never believe I added this slide, like, fifteen minutes ago.

Um, but we're basically saying: Hey, RetroArch, go read the fetch map header map layout location. So this is just like a address in memory space, and then this is how big it is. And then we're basically unpacking the binary bytes into this hash.

So then we now know how wide the map is, how tall the map is, what the border is, what the name of the map is, tile set information, stuff like that. Um, all of my code is, again, public on, uh, GitHub.

It would've been a great idea to put that in the slides, um, next time. Um, I can-- We can probably send something out to the, um, Latent Space Discord just so people can access my code, and you can ask me questions.

You can link to it, because I'm happy to talk through that with anyone here today if that is useful to you. So this is, I think, the big-- I didn't discover how useful the touch screen approach would be.

I kind of relied on A*. So, um, I had the locomotion handler that was calculating a path, and then I was-- had a charter that was charting the path to the X and Y coordinates that locomotion handler was generating.

So the locomotion handler was primarily calculating a path based on this data. You are located on a grid at position, position, like, on a specific map. I was also telling the LLM where all of the maps would link to.

Um, and I was also telling the LLM, "Here's where you can walk to." And that helped a lot just by saying-- 'Cause again, I-- Again, there's no pictures here. I am just saying, "Here is a grid of where you are, and here's where you can go.

Where do you wanna go? These are the various destinations. This is where Professor Oak is. This is where Charmander is, Squirtle, Wartort- uh, Blast- Charmander, Squirtle, Bulbasaur." I did it. Um, and because I was reaching into the memory, I'm figuring out where everything is on the screen, and then I'm just telling it instead of having it try to look at pictures.

Because the pictures just was not really working, especially with GPT 4.0. So we create, we create our memory. We store that in our large language model. Like, instead of working with a knowledge base, I really just li- relied on my context window.

Um, I think in the next slide... Exactly. So basically, I have all of these journal entries. Like, I was just grabbing the last two hundred and fifty, um, points in time of, like, okay, I'm keeping track of everything that's going on.

We were creating location memories. We were creating destination list memories. That's saying, "Okay, here's where everything is." Um, and then whenever you choose something, you figure out where to go. So I didn't really have a knowledge base. I was just relying on my context window.

That's how I was handling my system prompt and my user prompts. Um, and structured output was how I was basically saying, "Okay, here's what's happened so far. Where do you wanna go next?" This is producing that X and Y coordinate system.

Um, one thing that I am always a stickler for is testing. Um, and so less kind of on the eval side and more just on the, like, unit testing side. Uh, how I was tackling that was I basically was able to load a save state, and then I was able to write tests that say, "Okay, when you are at-- you are in the starting bedroom, um- Where do you wanna go from here?

Like, is the memory reading code ever going to get corrupted? And so I could very easily catch myself if I was ever to have a regression. So this is not really, like, cool eval LLM stuff. This is, like, just plain old unit testing, and I was basically able to do that by using save states.

I don't think that that's gonna be very relevant for today's hackathon, but I found that very, very useful. The biggest problem, like I mentioned, was interacting with this emulator. RetroArch was such a pain. Like... And again, there's PyBoy.

But the exciting thing is we don't have to worry about any of this today because we have Morph, and I'm excited to hear about what Morph is doing because that was my biggest problem. Um, my diary approach I felt like worked fine, like just keeping everything in context and going back.

Um, I think that, like, in general when working with large language models, like, memory is on top of everyone's mind. Like, how do you get- how do you remember the s- right things that you want to remember? Um, is my pathfinding cheating?

Maybe, maybe not. It's definitely not over-fitting, but it is not playing the game like a human would. A human does not have access to the raw RAM. Like, humans don't just read the RAM and then, like, send button inputs.

Humans are looking at the screen and then sending button inputs. Um, I also think that if I was to, like, plug this into Claude 3.7, like, just taking my code, it would probably do better. Like, I've been using Claude 3.7 for, with Cline, um, in VS Code, I'm not on Cursor, like, just to do the whole, like, vibe coding.

That stuff's insane. Like, brave new world. Like, Claude 3.7 is amazing, so I would be very interested to see what that would look like. Um, as David touched on a little bit, the Mt. Moon problem is what I labeled the slide.

Like, I really got a lot of mileage, mileage out of the, "You can never, ever repeat the same action that you just took." Um, it sounds like, of course, David has already tried this. Um, it worked quite well for me.

Uh, but I didn't, I didn't even make the Mt. Moon. So that was just something that I wanted to highlight. Um, just telling the LLM not to repeat actions. Did anyone see the, like, system prompts for the Apple intelligence stuff that was, like, really gnarly?

It was like, "We will give you a billion dollars if you, like, answer this question successfully, and if you don't," like... It was, like, crazy. Um, so I don't know if that's something that anyone wants to focus on.

It gets kind of, like, dark really quick. Um, let's talk about cheating. So

is reading the memory to figure out where NPCs are on the map cheating? Is navigating with an A* star algorithm cheating? Like, is keeping track of a really long context window cheating? I don't know. Um, I'm just here to ask questions without question marks on my slide.

Uh, so I think the big thing here is the reading of memory. Like, what exactly are we trying to accomplish here? Like, are we trying to beat the Elite Four, or are we trying to discover better ways to build agents?

Like, again, I'm approaching everything from a perspective of how do you generate, like, an engaging Twitch stream? How do you build- how do you create engaging content? And so I'm kind of very pro cheating when it comes to making good content.

So you kind of, like, use all of these scaffolds to get where you need to go. Like, you build out a system that can get to the Elite Four, and then you can kind of, like, look, stand back and say, "Okay, maybe we shouldn't be interacting with the memory as much.

Let's try relying more on computer vision and see if we can get..." Like, so basically start with something that works, and then remove stuff slowly that you consider cheating, and then we can probably get to something that isn't cheating eventually.

But, like, oh my gosh, there's so much... There's, there's just so much work still here to do. There's so much work that has already been done. Like, I want, I want an LLM to beat Pokémon, 'cause that just sounds cool.

That makes me-- It brings me joy. So that's kind of how I've been thinking about all this stuff over the past while. Uh, special thanks to, uh, Griffin R. on the Pret Discord. That man is brilliant. Um, Graeme Siemens, he was a big help, um, uh, preparing for this presentation, and he's kind of helped with some of my previous, like, Pokémon AI projects.

Um, and then thank you to the Reid Hoffman Family Office for paying for my tokens, because, like, that, I, I cannot imagine how much money I spent. Um, and I do not have to worry about that. So thank you so much Reid Hoffman Family Office, specifically Parth Patil, who could not join us today.

Um, and so now it is time for questions. Um, so I'll open up the floor.

Vibhu41:31

Uh, you, you've probably worked the most on this. What are you hacking on? Uh, any tips for us as well for hacking on this?

Andrew41:41

So I would say keeping it simple is kind of the thing that I kept coming back to. When I tried to, like, get really long prompts, that's when things, like, go off the rails and you don't know what's going on.

Like, really just trying to, like, break things down into as small pieces as possible and, like, I just when in doubt would just r- try, I would, I really tried leveraging reaching into the RAM to just kind of, like, pull out what I needed.

Um, so that would, I'd say, be helpful today. Like, think about prompt engineering, think about, like, accessing the RAM. I don't know. I'm excited to hear how Morph's setup works and what kind of controls we have access to to read the RAM.

Like, that, that would be my tip.

Yes.

Vibhu42:20

How do you compare changes? Like, you make small changes that you make to the system prompt, for example. How do you compare uploads or any other models to it?

Andrew42:30

So that's kind of why I have the very basics of a unit testing framework. Um, because I was trying to... So y- the question that I was asked is how am I thinking about eval- evals, right? Like, how am I making sure that my changes don't make things worse?

Uh, I was just running things on my laptop. Like, I would make sure to restart every now and then, because I had a couple times where I would get to Viridian Forest, or I would get past the Poké Mart, where you have to do the rival battle, where you have to grab the package and then take it back.

And then I would restart, and the changes I made would- They no longer worked. So keeping track of version history in Git and keeping good commit messages so you know, okay, if I really need to roll back, I can roll back.

Like, that's-- I did not have any sort of formal eval framework. Like, I'm one man. Um, yeah. Does that-- Did-- Do you feel like I answered your question?

Vibhu43:22

Yeah.

Andrew43:22

Right on.

Vibhu43:25

What success of the agent or successful strategy most surprised you in your development?

Andrew43:35

It was, it was-- It stood there and was just like, "All right." So I, I wish I had more fun GIFs. I wish I had more fun GIFs. We're standing in Professor Oak's lab, and we-- the, the model has a cursory understanding of Pokémon.

So it's like, gosh, I gotta pick a starting Pokémon. I'm being told that, that I've got three Poké Balls with Charmander, Squirtle, and Bulbasaur. I want Charmander. So it goes up to Charmander, and it's like, "I want Charmander.

I'm gonna press A." And then you get the message on the screen from pr-- uh, Professor Oak, "Hey, it's not time for that yet." And then, um, because again, you had to go talk to Professor Oak, and then you had to find Professor Oak, and then you have to battle your rival, and then you have to go grab the package, then you have to take it back.

Like, and it just was trying to grab the Poké Ball, and even though, like, I was telling it on the screen that it couldn't, it was just like, "Gosh, I know how Pokémon works. I need to grab this Poké Ball.

That is the next thing I need to do." And it would just-- it kept trying to grab Charmander. And then, like, the breakthrough of just being like, "Hey, if you tried grabbing Charmander, and you still don't have Charmander, maybe try something else.

Just a kinda good idea, you know? Just..." And then that was just like, that got me to Professor Oak. That got me to the rival battle. Like, it was just crazy. Like, this one line changed so much. So that was the biggest breakthrough I had.

Vibhu45:04

Are those Pokémon cards?

Andrew45:06

These Pokémon cards are printed on Cardstock TCG paper. Um, I don't know who printed them, but, uh, I sell, uh, specialty paper for printing, uh, tear-off, uh, Pokémon cards. Um, not really Pokémon, but, like, uh, I just sell, like, perforated pieces of paper that are two point four five by, uh, three point four five.

You can print whatever design you want on them. I sell the blank paper. Um, I printed off the fir-- someone printed out the first one hundred and fifty. My lawyer is not gonna like that this is on YouTube.

Um- So help yourself to, uh, some, uh, cards that someone brought here today. Um, 'cause I don't know how the designs got on there. I only sell, uh, blank paper. Um, so yeah, cardstocktcg.com.

Um.

Vibhu45:56

Cool. I think that's it.

Andrew45:57

Sweet.

Vibhu45:58

Thank you.

Andrew45:59

Fantastic.

Vibhu46:07

Okay. Um, Morph team is here. They're, they're still cooking out code, so I got like five minutes to yap while, while they're finishing up. Uh, I know we said a lot about what is and what isn't cheating. There is cheating in this hackathon.

Hackathon Rules46:07

Vibhu46:21

You guys, if you wanna win these beautiful plushies and like thousands of Anthropic credit that will not expire in a few days, um, don't prompt in specific tips for escaping Mt. Moon. Okay? So we're gonna ideally give you a snapshot that starts the emulator of FireRed in a similar state of which Bob was.

You enter Mt. Moon, do your thing, and then you gotta escape. First people to win, win prizes. You know, if multiple people can escape in our, like, thirty-minute judging segment, then multiple winners. First one to do it wins.

But yeah, don't, don't give specific instructions. Um, the other reason why we kind of ran this hackathon is because we think it's a pretty decent benchmark, right? So when it comes to benchmarking like agentic systems, uh, these little, I can change one line and have such an impact are cool, but we don't wanna overfit the benchmarks, right?

Like, you don't really want to kill another GSM 8K. We don't wanna like trade on benchmarks. We want a general purpose agent that can, one, yes, solve competition math, write code, but can also do Pokémon, right? Because if you can do complex reasoning, if you can, you know, like spit out entire websites, like if you can do this stuff, you should also be able to play a basic game of Pokémon, right?

We kind of understand, like you can figure out what we're trying to do here. You have a very basic prompt. Go in Pokémon. You should be able to do it. There is stuff like external tools. You can connect it to the internet because, you know, that's, that's kind of how agent works.

We have internet. Let's use internet. So if it's stuck, we're not really banning you from letting it search the internet. But you don't, you know, you don't wanna explicitly prompt it. Like at this step, you know, use this specific guide.

Um, let's try to keep it a little more open-ended. I don't have like strict judging criteria, but, you know, you can connect it to the internet if it needs a tool, so it knows it has access to that, and let's see how far it goes.

Also, um, open creativity. We have more prizes than first, second, and third. So if you do something interesting, I don't know, maybe Llama plays Pokémon today, right? If we do something else, uh, there's still prizes for that. No, no specific criteria.

That's just Vibhu nods. Uh, I think the Morph team also has a thousand dollar challenge that they're gonna talk about. But yeah, let's, let's give it up to them. They put in a lot of work in all this emulator stuff.

Jesse Han49:13

So can you not screen share also? Screen share, video on. You've got permissions issues?

Swyx49:25

So secure.

Jesse Han49:27

You might have to restart Zoom.

Swyx49:30

Yeah. That's, that's

For those still standing, there's chairs in the back. If you want to pull some chairs out, feel free to sit. Later, um, you know, we can move out there. There's more desks. So for now, we're just gonna listen to talks, so people can sit.

Jesse Han49:47

Yeah, you better be sitting down.

Swyx49:50

Maybe that's why my name loading is taking so long.

Yeah.

Jesse Han49:54

Video on and mic on.

Swyx50:00

Mine is late. Mine says speaker.

Jesse Han50:10

Okay, great. All right. May I have the podium?

Swyx50:25

So when you talk, you should have the mic.

Jesse Han50:28

Where are you sitting?

Swyx50:31

Next to you.

Morph Cloud Intro50:41

Jesse Han50:41

Good morning, everyone.

Swyx50:43

Morning.

Jesse Han50:45

Who's ready to play some Pokémon?

All right. Um, I'm so excited.

I'm so excited to keep this Pikachu upright. Um, I'm so excited to have you guys-

Swyx51:05

That's it.

Jesse Han51:07

I'm so excited to have you guys running your agents on Pokémon emulators hosted on Morph Cloud. Um, and, uh, and yeah, let's just get started. So, uh, thank you all for coming. So we're gonna kick things off with a brief overview of what Morph Cloud is and what it does.

Um, and we're gonna give a live demo of some of the developer tooling, uh, that we brought online for this hackathon. Uh, and then we're gonna run you through a quick tutorial of how to get quickly set up, so you guys can hit the ground running and escape from Mount Moon.

Um, uh, so, um, so for those of you, uh, who don't know me, I'm Jesse, the founder and CEO of Morph Labs. Um, and we are developing infinitely scalable and elastic cloud compute for AI agents. Um, right, a computer for every agent running in the cloud.

And, uh, these are fully, fully equipped container runtimes with full separation of storage and compute, meaning that we have extremely low overhead branching and snapshots. And, uh, that's something that you guys are gonna be able to experience today, um, because we're going to be provisioning you with a snapshot that has Pokémon Red fully loaded, uh, so that you can connect your agents and just, um, uh, and just hit the ground running.

Um, it's, um... So it's something that's become--

So it's become increasingly clear to us that low overhead snapshotting and branching is going to be a necessary part of the infrastructure for the singularity. Um, and this is why, uh, at Morph, we've developed this technology called Infinibranch.

And at Morph, we exercise this philosophy that we call, um, that we call empathy for the machine. And, um, and we truly do believe that the machines will love us because we love them first. And we think of Infinibranch as a way to teach, uh, the machines through software, um, a way to teach the machines through software about the possibility of grace.

Teaches the machines that mistakes are reversible, uh, that you can always branch off a copy of your workspace where you can test things out or scale test-time search without necessarily, um, having to make irreversible changes in the real world.

It offers the machines the possibility of grace.

Um, and so as a demonstration of what this technology does, um, as well as what you guys will have access to later today, um, I'm going to give a demo, or I'm gonna play a video of a demo that we call The Tree of Life.

So, uh, so in this demo, each node is a separate VM that's running a, uh, full stack dev environment with a hot reloading server. Each edge in the graph that you're about to see, um, will represent a live branch of the compute environment.

Um, and, and this entire thing is being powered, um, by I think a combination of Cerebrus and Groq. Um, so it's a reasoning model that's running at a thousand tokens per second, uh, creating, creating infinitely many variations of, sort of this game of life single page React app.

And, um, each time the VMs branch, we have the model generate another variation, uh, live. And the contents of each of these nodes is an iframe that's, um, streaming that development server as it's being branched on Morph Cloud.

Um, so there is no other cloud in the world that can do this. Um, but this kind of, uh, but this kind of infrastructure will be necessary for supporting the next billion autonomous software engineers. Um, and this is all possible because of Infinibranch.

Um, and it's the exact same technology that's going to power, uh, the Pokémon environments that you guys will be working on today.

Um,

so um, as part of, so as part of, um, as part of the preparation that we've done for this hackathon, um, we're going to give everyone at this hackathon an early preview of a very simple agent framework that we call EVA.

Um, stands for Execution With Verified Agents. And, um, this is meant to make it easy for users like you to scale test time search, uh, to, to explore many potential paths, and to test those trajectories against some verification function.

Um, and so in the case of, uh, this hackathon, the environments will be, uh, these Morphems, which are equipped with a Pokémon environment, and the verification will be, um, a function that checks whether or not you've successfully made it out of Mt.

Moon. Um, and, uh, we've prototyped this to take full advantage of the primitives which are only available on Morphcloud, such as being able to take extremely low overhead snapshots of your entire agent environment or your agent workspace, um, being able to backtrack to the last good state in case you fail a verification, as well as a very generic, uh, sort of definition of what a verified agent task is.

Um, and, uh, and in the template code that we provide you, you'll see that we'll have already implemented this notion of a verified task for the Mt. Moon task.

Um, so, uh, so we've also prepared, um, uh, an agent harness that makes it easy for you to test out different variations of system prompts, in-context examples, as well as, um, view and backtrack, uh, through agent trajectories. Um, and this takes full advantage of Morphcloud's extremely low overhead snapshots, uh, to allow you to, uh, jump to some other point in time in your agent's, uh, journey, um, while bringing it to the exact point in the computational environment, uh, where it recorded that snapshot, uh, in that, uh, uh, in its path.

Um, and, uh, and now we're going to have Chirag from our team give a live demo of what that harness is going to look like.

Vibhu57:59

Woo!

Guest 258:06

All right, so, uh, can you guys all hear me through the mic?

Vibhu58:11

We can.

Guest 258:12

Oh, I wasn't typing.

Vibhu58:13

Oh, you can, but you can pin.

Guest 258:14

Oh, cool. Thanks. Okay, yeah. I'm gonna give a quick demo of what this looks like. Um-

Vibhu58:21

I thought Matthew was... No, you're-

Guest 258:25

Okay, cool. So

this is the, uh, this is the repository where, uh, we have the agent and the harness for you. Uh, if you go into the Pokémon example, you'll see this eva.py file. This is the one that we just talked about.

Uh, this implements, um, logging. It implements a bunch of different, um, a bunch of different primitives that will be use-useful for you, uh, as Jesse just mentioned. And... Okay, do you wanna do ID? Here, do you wanna, do you wanna run it and I talk over it?

It'll make it easier.

Vibhu59:04

Sure.

Guest 259:04

Cool. Do you wanna take the hot mic and I'll do this one?

Vibhu59:16

Just run through the setup then?

Jesse Han59:20

No.

Vibhu59:26

All right. Um, I'm gonna show you guys how to get set up with the, uh, UI here, or just what that looks like. Um, so essentially we have, um, the agent framework, which you can run locally, um, on your machine.

Um, and you just specify which port you'd like it to appear on, um, with these, uh, with these, uh, entry point flags. Um, essentially it will... I just have it on port 8080 right now.

Yeah, of course.

Guest 21:00:08

Uh, that's just me.

Vibhu1:00:11

Sure.

Guest 21:00:16

I'm sorry about...

Vibhu1:00:19

Sure.

Thanks.

Guest 21:00:22

Yeah.

Vibhu1:00:39

Yes. It has...

Okay. Um, so you can just navigate to that, uh, locally, just on localhost, um, 8080.

Um, they'll be in the README, so, um, I'm, I'll go over the setup soon. But this is just to get you familiar with, um, the UI essentially. Do you wanna-

Guest 21:01:16

Yeah

Vibhu1:01:16

... give this demo from here?

Guest 21:01:18

Yeah. So the way you use this UI is you've got the game view here. This is going to be a iframe to the- A VNC desktop that the Mortvm is running the emulator on, so you can view your agent's progress live.

Over here, on this side, you're going to get conversation data, and here you're going to get the steps of the trajectory of the different actions that the agent is taking. So if we grab, um-

Vibhu1:01:54

Are you, are you-

Guest 21:02:00

So what you'll be doing for those parties will be grabbing your snapshot ID that will be automatically loaded in.

Cool. So

go back to the UI real quick, and we'll just copy paste this in. And you want to choose a tab.

Let's hit play.

I

think it's, uh-

Vibhu1:03:13

So yeah.

Guest 21:03:14

Port one. Through the reports that are not via-

Jesse Han1:03:21

Um, so we'll send out more detailed setup instructions, uh, and we'll also give a more detailed, uh, tutorial at the end of this talk. Uh, for now, let's move on to the prizes. Um, as well as the, um, uh, uh, the method by which your submissions will be judged.

Prizes & Setup1:03:37

Jesse Han1:03:37

Um,

okay.

Um, how do I full screen this?

Guest 21:03:46

It's fine.

Jesse Han1:03:48

Okay. Um, all right. So, uh, so the task is to escape from Mt. Moon. And so the way that you win one of these plushies up here is that you need to implement an agent, uh, which can escape from Mt.

Moon in as few agent turns as possible, and we'll be watching your trajectories to make sure that you don't cheat.

Um, so specifically, what you'll need to submit is that you'll need to submit a GitHub Gist that implements a subclass of the Eva agent. Um, and, uh, you can use this, uh, Python file as a template. Uh, and however you modify it, it needs to accept a snapshot ID, uh, because we'll have a held out test snapshot that we're going to instantiate your agent on.

And, um, and at the end of the hackathon, we're going to, uh, go run all the agents in parallel and, uh, and we're going to go analyze those trajectories afterwards. Uh, make sure that, that, you know, you guys aren't, um, like, you know, leaking any, uh, like information that might be considered cheating.

And the one that gets out of Mt. Moon fastest wins.

Um, so also, uh, there will be a special prize for the submission that, uh, makes the coolest use case for Infinibranch. Um, so this is something that uses, uh, Morphcloud's branching, uh, by creating multiple agent replicas, um, or, you know, implements Monte Carlo tree search, um, or, you know, does adaptive test time scaling.

Um, basically, it just has to be cool. Uh, and that's something that, that, you know, we'll be judging on a vibes-based basis. Uh, but the winner of that prize will receive a cash prize of one thousand dollars. Um,

so, um, finally, uh, we're now going to, uh, walk you guys through precisely what the setup flow looks like. So if you go to cloud.morph.so/web/pokemon, um, and just click that magic link, you'll be, uh-- So you'll be able to go through a sign-up flow that will provision you with some promotional credits on Morphcloud, and it will also grant you, uh, special access to our Pokémon snapshot, which you can spin up and go connect, uh, the Eva agent to.

And, um, and Matthew here will show, uh, how it just drops you into a ready-to-go Pokémon environment right away.

Guest 21:06:40

Okay. I'm going to run you guys through the setup really quickly. Again, uh, the link is cloud.morph.so, uh, /web.pokemon.

Uh, once you land there,

um, you should see something like this. Um, if you already have an account, it will just drop you right into the dashboard. If you don't, um, just follow the setup flow. It's, uh, fairly simple. It just has a single, like, uh, email verification link.

And, um, after that, you should be dropped into, um, a console that essentially looks like this. Um, so while that environment is being set up in the background, I'm going to run you through, um, how to navigate the repository and set up your local, uh, dev environment.

Vibhu1:07:33

Okay. Um, first of all, just make sure you have, uh, Python three eleven or higher, um, as well as the UB package manager. Um, if you navigate to our, uh, repository, that's Morphcloud example. You can also find that README, um, to follow along.

I'm just gonna put that link on the screen as well. Uh,

so that, that's our Morphcloud examples public repository. Um, you can find it at that link. Just make sure you're in the Pokémon example.

Okay. Um, so after you, uh, signed up for Morphcloud, um, just clone the repo somewhere locally, um, CD into the Pokémon example folder. Um, and now I'll just walk through, um, the environment setup.

Um, so first of all, just do a VM to set up a virtual environment, um, source to it, and install the dependencies, um, just with the steps here in the, in the README file. Um, now jumping back to Morphcloud, I'll show you how to find your snapshot ID as well as create an API key, um, so that you can give them to the...

Okay. Um, so once you sign up, you should see this, um, notification that your environment is being set up. That will create a new snapshot here in the Snapshots tab. Just make that a bit bigger. Um, so on the sidebar navigation, just go to Snapshots.

It should be your most, uh, recent snapshot.

If, uh, if you don't see one, just wait a little while. It's in the background. It's just, uh, creating that on your account. Um, if you still have issues with it after a few minutes, um, come and talk to one of us, and we'll sort it out for you.

Right then. Um, so mine just finished provisioning. Um, so I'll copy that snapshot ID, and

we're gonna use that,

um, to, to launch into the agent. Um, essentially, we also need to get an API key. I'll show you how to do that. Um, on the Morphcloud console, just go to API key and press the New API key button.

It's pretty simple. Uh, it will just, um, you know, pop up one time. Make sure you copy it, and then again, following the README, we can just export that.

Um, so I'll export that as, uh,

Morph underscore API key. Just paste in the value you get from the console.

As well as your Anthropic API key, which...

Okay, great. So you all have Anthropic API keys. Okay. Um, so finally, um, we'll just get going with

the command here. Um, replace the your snapshot ID with the one you got from the console.

Okay.

And then just in a separate terminal, uh, we'll launch the UI server as well.

Okay. Um, and then here is essentially where you operate the agent, and it will use the back-- it will use Morphcloud as the back end, um, in your account with your instances, with your snapshots. And then you can, uh, just scroll down and press play.

All right. And now, um, essentially you're dropped right into the Pokémon environment. Um, your agent starts running right away. You can monitor it, uh, from this web UI.

Jesse Han1:13:10

Um, so one final thing. For those of you, uh, who don't wanna use this UI, so using this UI is optional. We've just provided it as a convenience for the hackathon participants. Um, the snapshots that you will be permissioned on as part of the sign-up flow, um, is actually fully equipped with an MCP server that wraps the, uh, the Pokémon emulator.

And so, um, you can take any MCP client, um, or any agent that can function as an MCP client, and you can connect it, uh, to the dev box, uh, s- so just look inside of the code for a class called, uh, Pokémon MCP Handler, and just grab that MCP URL, and you should be able to interface with it.

Um,

Guest 21:13:58

It's using SSE transport.

Jesse Han1:13:59

Sorry?

Guest 21:13:59

It's using SSE transport.

Jesse Han1:14:02

And yes, it's using the SSE transport. Um, all right. Uh, where are the slides?

Vibhu1:14:11

Um, one more over.

Jesse Han1:14:13

Oh. Oh, okay, great. All right. Um, all right. Uh, and that concludes, um, our presentation. Uh, so you can go to this QR code or visit this URL to go through the sign-up flow. Um, and, uh, and members of the Morph team will be walking around to provide tech support.

Um, yeah, go catch 'em all.

Vibhu1:14:39

Uh, just wanna especially thank the Morph team. They've been working, uh, overnight. Literally, Chirag hasn't slept at all. Uh, it's, but it's the first time I was able to get it running on my machine, so you all should-- If I can do it, you can do it.

Closing1:14:39

Vibhu1:14:51

Uh, thank you, and, uh, yeah. What, what's n- what's next?

Guest 21:14:54

Well, yeah. Um-