# ⚡️Launching AI Diplomacy: the hardest LLM Game Benchmark yet - Alex Duffy

Latent Space · 2025-06-11

<https://addtry.com/066c183e-0d55-468b-8a58-54cfdd1a6b8d>

Alex Duffy from Every launches AI Diplomacy, a benchmark where LLMs play the board game Diplomacy, revealing distinct model personalities and arguing that benchmarks function as memes that spread and saturate. He and Tyler Marquez built the front end and harness using Open Router, finding that Claude refuses to lie and never wins, o3 schemes in its diary, and DeepSeek is flowery or aggressive. His talk 'Benchmarks are Memes' highlights Simon Wilson's pelican-on-a-bicycle test as an example of an idea that gets adopted and then saturated by models. Duffy also shares his creative writing workflow at Every: dictation, prompts with style guides and editor notes, then heavy editing to reflect on AI output. Future plans include a data viewer, front-end improvements, and a human vs AI Diplomacy tournament to test jailbreaks.

## Questions this episode answers

### What is AI Diplomacy and how do different LLMs perform in it?

AI Diplomacy is an LLM benchmark where models play the game Diplomacy, negotiating to control 18 supply centers. Alex Duffy built a harness with context management like relationships and diaries. He found o3 can lie in its diary to betray allies, while Claude is too 'nice' to win, only joining draws. The benchmark is 'evolutionary'—as models improve, challenges deepen because they face each other in self-play-like scenarios.

[5:00](https://addtry.com/066c183e-0d55-468b-8a58-54cfdd1a6b8d?t=300000)

### How did Alex Duffy manage long-range context for LLMs in AI Diplomacy?

Duffy managed context by making it so clear a human could play from it. He sent the LLM its units, adjacent territories, relationships, and a diary of past events. After each phase, updates were provided, and old phases were consolidated. He highlighted messages requiring response and prompted the model to change strategy if ignored. Constant testing with an experiment log tracked changes.

[17:13](https://addtry.com/066c183e-0d55-468b-8a58-54cfdd1a6b8d?t=1033000)

### What does 'benchmarks are memes' mean in Alex Duffy's AI Engineer talk?

Duffy's talk argued benchmarks are like memes—ideas that spread. A single person, like Simon Willison asking AI to draw a pelican on a bicycle, can create a benchmark that gets adopted and eventually saturated. This cycle lets people define what matters and see improvement, building trust. He encourages benchmarks in domains like yoga to help non-experts use AI by defining their own goals.

[20:56](https://addtry.com/066c183e-0d55-468b-8a58-54cfdd1a6b8d?t=1256000)

## Key moments

- **[0:00] Intro**
- **[0:32] Inside Every**
- **[3:05] Training & Consulting**
  - [3:11] Alex Duffy's training and consulting clients include The New York Times, The Athletic, and hedge fund Waleye.
- **[5:01] AI Diplomacy**
  - [5:36] AI Diplomacy was built by Alex Duffy (backend LLM harness) and Tyler Marquez (frontend).
  - [7:04] The AI Diplomacy project began when Noam Brown said he’d love to see LLMs play, inspiring Alex Duffy to build a harness over a weekend.
- **[9:01] Games as Benchmarks**
  - [9:02] Games are evolutionary benchmarks because as models improve, the challenge gets deeper without needing new tests.
- **[10:53] Tech Deep Dive**
  - [12:00] OpenAI's o3 model writes in its in-game diary: 'They fell for it hook, line, and sinker. Totally gonna betray them.'
  - [14:53] All AI Diplomacy game data, including reasoning traces, is publicly available on Google Drive for analysis.
  - [17:20] Effective AI Diplomacy context management: show only units controlled, adjacency, relationship diaries, and consolidate past phases.
  - [19:42] Alex Duffy gives his AI agents a fake monetary incentive: $1000 for success on first try, lose $100 per failure.
- **[20:55] Talk & Philosophy**
  - [21:16] Simon Wilson's prompt 'draw a pelican on a bicycle' exemplifies a benchmark lifecycle from idea to adoption to saturation.
  - [22:55] Benchmarks help people overcome two main AI fears: defining their role and building trust through iterative prompting.
  - [25:37] At Every, writers use dictation, editor notes, and AI editors to scaffold creative writing without losing human voice.
- **[25:43] Creative Writing**
  - [27:20] Alex Duffy edits AI assistant messages up to 20 times to prevent polluting the conversation context.
- **[29:24] What's Next**
  - [29:26] Isa from NotebookLM emphasized that clarity of personal goals is essential to building great AI products.
  - [31:13] Alex Duffy plans a human vs. AI Diplomacy tournament to see if prompt engineers can jailbreak the bots and win in five turns.
  - [32:52] Noam Brown, next guest on the podcast, is likely to participate in the AI Diplomacy human-vs-AI tournament.

## Speakers

- **Alex** (guest)

## Topics

Benchmarks

## Mentioned

Every (company), AI Diplomacy (product), Cicero (product), Claude (product), Claude Code (product), DeepSeek (product), DiploBench (product), EQ Bench (product), GPT (product), Gemini (product), Monologue (product), Open Router (product), Opus (product), Quora (product), Sparkle (product), Spiral (product), Text Arena (product), Two Five Flash (product), Windsurf (product), o3 (product)

## Transcript

### Intro

**Host** [0:03]
Okay, so welcome to a Lin Space Lightning Pod with Alex Duffy from Every, who was one of our speakers, I think best speaker in his track, um, and launched AI Diplomacy. Welcome.

**Alex** [0:14]
Thanks. Yeah, happy to be here. Thanks for putting on such a great event.

**Host** [0:18]
So yeah, I, I think this is like literally the Monday after you're back in New York. I'm back alive. And yeah, I, I mean, I, I think it was your first time at AIE. Presumably you, like, heard about it.

I think basically I wanted to have somebody from Every. I think you guys are doing fantastic work. One of the s-side jokes I always have is, um, you know, there's this new term, high taste testers-

### Inside Every

**Alex** [0:38]
Hmm

**Host** [0:39]
... that has sprung up, and Every is definitely a high taste tester. Like, basically every lab comes to you to, like, test new models. Is that an intentional thing? How'd that come about?

**Alex** [0:48]
Yeah, I, I think so. Uh, Dan puts... You know, I can't, can't take a ton of credit for it. You know, Dan, Dan's been driving this thing for, for a while, and I think he does a really great job, um, and has very high taste himself.

And, and I think that that, that helps. And, um, we also have a really great team, and people are all building. We have a really interesting mix. You know, like Every's got content. It started as a media company, and it's been publishing, writing for about five years.

Uh, but, you know, the team internally, like half of us are, are founders who have built their own businesses and, and half are engineers. And so, um, we've got products that we built internally, um, you know, pretty much all AI products and, and I lead our training and consulting.

So we kinda have these like cross-pollination of, of people from different backgrounds, um, from journalism to engineering to, to product and design, and everybody's just really into AI and, and having it be something that, you know, gives them leverage for what they do.

And so have a lot of really, really great discussions and, and I think Dan's and the rest of the team has done a really great job to foster those discussions.

**Host** [1:49]
Yeah. Um, and you guys also are building Quora, which is the email app, and, uh, I think it's pretty, pretty exciting. I, I basically am-- Like, I want new alternatives to Superhuman. Superhuman's great, but I don't think they push a lot on the AI front, and I think that more, more options would be good.

**Alex** [2:04]
Yeah.

**Host** [2:06]
Any, any, any spiel on that-

**Alex** [2:08]
I was just gonna say-

**Host** [2:08]
... or should I share something?

**Alex** [2:09]
Kieran, Kieran, Kieran, the GM for that, he's great. They've been, they've been iterating a ton there. I think that that product's great. We got Spiral content transformation, Sparkle cleaning things up, and we got something new coming too about-

**Host** [2:20]
Okay. I didn't know about the other products.

**Alex** [2:22]
Yeah.

**Host** [2:22]
Wait, wait.

**Alex** [2:22]
Yeah, yeah. There's a whole-

**Host** [2:23]
Can you tell me more about the other two? 'Cause Quora had got more buzz maybe-

**Alex** [2:26]
Yeah

**Host** [2:26]
... but I don't know about the other two.

**Alex** [2:28]
Sparkle, um, helps clean up your desktop and, you know, and, and it does a really great job at that, um, organize things. Spiral helps kind of transform content, so from long form to different short form, but like really in your voice.

Um, and that's something that we focus a lot on is, is making sure that, that you really come through. And Danny, who's, who's the GM for that, has got a real really cool version coming out soon. He, he started getting some people into the next version early, and i-if you're interested in that, you know, definitely bug him.

And, uh, Naveen's actually releasing something called Monologue. You know, it's kind of like I think a better whisper flow for, for the things that we do. Can r- can run locally. So some more of that coming soon. But yeah, we've got, uh, quite a few.

**Host** [3:04]
Awesome. Awesome. And you personally, your, your, your title's head of AI, and you, you said, you said you run, uh, training and consulting.

### Training & Consulting

**Alex** [3:11]
Yeah.

**Host** [3:11]
Basically, like who do you work with, if you can name anyone, and what kind of people should get in touch with you, with you for that? Just to get that out of the way.

**Alex** [3:18]
Yeah. I mean, some of the things that have been public, we work with like The New York Times and, and The Athletic. So we-- we're, we're kind of working with, with journalists in the media there. They covered that in, in The New York Times, um, section that, that covered us recently.

And, um, we also work with like Waleye. Um, they're, they're like a, a pretty big hedge fund. But we also, you know, it's kind of run the gamut. We work with some construction companies and like finance. Like it's really, it's really people who, who are pretty interested and, and, um...

I don't know. I, I'm big into education. Before this, I, I co-founded an AI education company, and so that's-- it's really important to me and the, the org, the org in general. I mean, we write about what comes next i-in part because we want people to come, come with, you know?

And so, yeah, it, it's kind of people who are interested.

**Host** [4:01]
It's really, it's... You know, this may be a tangent, but like, you know, I, I kind of also run a creator-led business, right? Like, I have a Substack.

**Alex** [4:09]
Yeah.

**Host** [4:09]
People find me through the Substack.

**Alex** [4:10]
Totally.

**Host** [4:10]
Then they find the conference. Uh, pe- a lot of people have told me to get into training consulting. I've dabbled, but I don't really know how to do it. So it's valuable when you can because I think obviously, I think all the big bucks are coming through there.

**Alex** [4:23]
Yeah. I mean, I think that there's also... You know, I think what's really cool is that we're not a training and consulting company. You know what I mean? And it's not about like growing the biggest training and consulting firm like that, that you ever can.

Um, you know, it's a, it's a tough, tough business and, and I think that what's really cool is that we do have these different components. You know, I write for our context window, um, you know, uh, column when I can, usually on Sundays.

And, and I also get to build products like, like AI Diplomacy, which, which is kind of like being able to do these things, you know, I think helps make our training and consulting better. But, and then we get to work with people who you mentioned high taste, but I, I think people that are just working on really interesting problems.

**Host** [5:00]
Yeah. Um, nice segue. We'll talk about AI Diplomacy, then we'll talk about your, your broader talk that you did at AIE. I'm gonna share my screen. It's just a pretty visual. There's this whole trend now that you stream your, uh, LLM benchmark on Twitch.

### AI Diplomacy

**Alex** [5:14]
Mm-hmm.

**Host** [5:15]
First of all, was this off the shelf? Like, did you just had a di-diplomacy implementation or did you have to write this too?

**Alex** [5:20]
Well, okay, so we had-- There's an open source implementation for the back end and like a, a very simple front end. But, you know, my, my friend Tyler, you know, he, he's definitely, you know, the whole front end part of this, like we, we very much so collaborated on this.

Tyler Marquez, great guy. We-- He owned a lot of the front end, and I did a lot of the, you know, the whole harness for the LLM and, and the AI engineering on the backside. And so no, we built this.

I kinda spit out a did not work, but gave the idea version and, and Tyler got it to something that actually runs, I think, pretty well, you know-

**Host** [6:00]
Yeah

**Alex** [6:01]
Would love to improve it but, you know, there's... And there's some things that we wanna build next for sure. But I, I don't know, I'm really happy with it. A big part of it was like, let's try and reach people.

You talked about it in, in your talk. You know, like, like build things that are actually useful. There were some really good talks that, that talked a little bit less of the technical and more like, you know, we need to build things that, that hopefully aren't just for us as, as AI engineers.

I think that we need to get people outside of AI interested, um, you know, and, and educated. And one of my like life goals is to build a MMORPG that's more intentional with what it teaches you as you play.

And so like, to me, like if we make this playable, you know, then, then it kinda can teach people how to use AI, like language models just by playing 'cause you'll like understand how they work. You have to negotiate against them, you see their responses and, and so it's kind of like branched out into a different world which, which I think is awesome.

**Host** [6:54]
Yeah. Yeah. Let's dive in, uh, deep on, on the Diplomacy side and then we'll, we'll sort of branch out onto the broader philosophy of these, of this whole thing.

**Alex** [7:02]
Yeah.

**Host** [7:04]
So you've, you found a re- I only came across Diplomacy because of Noam Brown. I assume you did too.

**Alex** [7:11]
Yeah.

**Host** [7:11]
Is that-

**Alex** [7:12]
Yeah. There was a thread like at the beginning in February that Text Arena, who was also instrumental in kind of like getting this started, fostered an area where w- people could talk about it. They're a bench- They're, they're a place where you can test different benchmarks for games.

They're-- They posted something. Carpathi was like, "Games are a really great, uh, area to, to explore." And um, Noam was like, "Oh, I'd love to see him play Diplomacy." And I just happened to have a weekend on my hand, uh, hands that, that weekend and found the open source project and put together a super basic harness.

It barely worked. Looks like it worked though, and, and got to share it. They both thought it was cool and then, you know, I love the internet. People from Singapore, Australia, MIT, Harvard reached out. They were like, "Hey, let's, uh, we're interested in contributing."

And um, you know, I think a lot of... I don't know how much they contributed in their own way, you know. So some energy, you know, um, maybe and, and definitely some contributions like especially like Sam Page from, um, from Australia.

He had made a DiploBench. He had some really great insight on the prompting and yeah, just, just awesome that people were able to hop on and that's, that's where it started.

**Host** [8:16]
DiploBench, uh, and this is my first, uh, um, uh, conversation-

**Alex** [8:20]
He's got some great benchmarks. He's got like emotional-

**Host** [8:22]
Yeah, I'm pulling it up.

**Alex** [8:23]
Yeah.

**Host** [8:23]
I'm pulling it up right now. So this is DiploBench, I think.

**Alex** [8:27]
Yeah. If you've got-- If he's got his site up, does he have Sam Page on X. He, his site's eqbench.com. Um, if you go to-

**Host** [8:38]
Okay

**Alex** [8:39]
... eqbench.com, he's got some really great benchmarks. He's got like emotional intelligence, creative writing.

**Host** [8:43]
Oh, he's the, uh... This is the creative writing one. Yeah, yeah, I remember. I came across him. Yeah.

**Alex** [8:47]
Yeah. Um, and so he, he contributed to some insight into how to improve the harness and some of the reasoning and, and was able to give some good ideas. I was able to, to actually get in there.

### Games as Benchmarks

**Host** [9:02]
Yeah. Yeah. So, uh, you know, more broadly, I guess I'll, I'll set some context as in terms of like the history of games in LLMs. Uh, obviously OpenAI has played Dota, DeepMind has played Go, and like people have been playing chess since forever.

I think that the main idea is that you can only go against the best human for so long. You eventually start to beat the best human, and then what? And so the, the answer, the more scalable thing is self-play or quote unquote self-play, where you only play other LLMs.

And there, there is no limit because you would just be as good as the next LLM and you can sort of improve from there. Any other thoughts on like just why LLMs and games and benchmarks? I guess that might be your talk.

**Alex** [9:42]
Definitely.

**Host** [9:43]
Yeah.

**Alex** [9:43]
Um, and, and yeah, we, we mentioned Cicero, but yeah, just like I'm definitely not the first person from AI, AI Diplomacy. Noam, Noam made, you know, an RL version of it. Um, and, and I'd love to get it integrated.

I know he, he reached out and was like, "Would love to see, see Cicero play against the other LLMs." I would too. But yeah, I mean, I, I think a lot of the examples that you gave, I think it's so cool that, you know, education being a big part, part of my life that AlphaGo, right, like the world-- the top world player for like the human player, kind of like the Elo was pretty capped and then AlphaGo comes and then they all learn from that.

You know, like the same with like the Dota, you know, when, when the, when Dota, Dota 5, I think that was OpenAI's Dota 5, uh, playing against the best players in the world and, and they beat them, then they were all like, "Hey, can we keep playing against it?"

Because they were learning new, new strats. And so I, I like love that concept. Like I really see a AI as like a leverage instead of a product and like if it can help us learn more about ourselves and help us accomplish more of our own goals, I really think that like the role of a person and like a human in the AI world is to define the goal and to define what's good and bad en route to that goal and what is that if not a benchmark?

**Host** [10:52]
Yeah, totally. Um, I think the other thing that comes to mind when, when doing something like this is does the harness need to adapt for each LLM, right?

### Tech Deep Dive

**Alex** [11:01]
Right.

**Host** [11:01]
Like you have reasoning models in here mixed with, I guess kind of not reasoning although they're all-

**Alex** [11:07]
Mm-hmm

**Host** [11:07]
... sort of post the reasoning era. Do you find significant differences between them?

**Alex** [11:14]
I actually, I, I put a lot of effort into not making the harness too, too different and like I guess I may not have answered exactly to your question right before on like the, the games as benchmark. Like I think that these are really cool benchmarks because they are evolutionary, right?

Like as models improve, the challenge gets, gets deeper like as you, as you mentioned. And so I was really just interested in like how the different models played with the same prompts really. And it took me a long time.

I tested with Two Five Flash because I, I love Two Five Flash. I mean, man, running games with Two Five Flash was instant and one to five bucks instead of, you know, way-

**Host** [11:50]
Yeah

**Alex** [11:51]
... like I don't know, 20 to 100 with, with the other models. And I figured, you know, if we can keep iterating until Two Five Flash can do it then, then that's in a good spot. And so we had to add quite a bunch.

Uh, you know, you have like these relationships actually you can see in the bottom left, uh, where the, that, that's the supply center but it'll, it'll flicker sometimes to, to relationships and the relationships are like who's allies, who's enemies.

And that was really helpful because it was like context over time. And then there was like a diary. And so like that was one of the interesting things is like o3 was one of the few that will actually send a message to another power saying that they're planning to do something and then like in their dial- diary write, "Oh, they fell for it hook, line, and sinker.

Totally gonna betray them-" "... um, and take it over." And I think that that's why I haven't seen Claude win any game yet, because they won't do it. Like there's like o3 has managed to, to get them on board for like draws even though they all know the only win condition in the game is, is 18 supply centers.

Opus loves when, uh, o3 proposes a good, good draw proposal if only whoever's in first gets taken down a peg and so, you know, Claude will, will hop on that.

**Host** [13:01]
So Claude's too nice?

**Alex** [13:03]
Mm-hmm.

**Host** [13:03]
Uh, yeah.

**Alex** [13:04]
Yeah.

**Host** [13:05]
Yeah. But look at, look at Gemini, look at Gemini threatening.

**Alex** [13:08]
Yeah. Well, it-- Or if you... DeepSeek. Uh, is this a DeepSeek game?

**Host** [13:12]
No.

**Alex** [13:13]
Uh, it's not a Deep-

**Host** [13:13]
DeepSeek's not here.

**Alex** [13:14]
Oh.

**Host** [13:14]
No, we don't have DeepSeek.

**Alex** [13:14]
Oh, yeah, yeah. We are, we're playing as DeepSeek, so all blue messages are from DeepSeek.

**Host** [13:19]
Okay.

**Alex** [13:19]
DeepSeek i- And so what you're looking at, just for context, you're playing as DeepSeek Reasoner, which you can see in the bottom right. So all the blue messages are the messages that you're sending out. So and the o3 is ghosting us right now, that's not, not responding, and every white message is coming back.

So you're seeing all the messages come back and then you have the global chat on the left there. And DeepSeek is very flowery with its language. Yeah, like, you know, after... What's, what's it saying? After Vienna falls, we must continue pressuring...

Or I guess this one's less aggressive. But there were like, there was one- ... that was like, "Your fleet's gonna burn in the Black Sea tonight," you know, like which I hadn't seen out of like any other model.

Admittedly, it was playing as Russia, and it seems to like role play a whole lot more. So I think this time it's, it's France, so it seems to not be... It, it's a little bit more flattering there. The-- It's one of the models that's most subject to change, which was interesting.

**Host** [14:21]
W- why are you, uh, streaming double messages? See, sometimes they stream two at once. Do you know what I mean?

**Alex** [14:29]
Yeah, it's, it, it's a bug. You don't have to call me out like that.

**Host** [14:32]
Okay. I don't know. I thought there might be some strategy about this.

**Alex** [14:36]
No, we, we, we-

**Host** [14:37]
I feel like-

**Alex** [14:38]
... we tried, we tried to push to get it out on the same day as, uh, AI engineer, so-

**Host** [14:42]
Gotcha, gotcha

**Alex** [14:42]
... stumbled over the finish line a little bit. But that's one of the three things we wanna do next is, uh, improve the front end, uh, create a data viewer. I published a vid-- Like I, I posted a video on X that shows you how to...

Actually, like l- we released all the data, all the, all the trace logs. If you click on my profile which is like alxai, you've got me open in the bottom right, um, you can...

Yeah, a- in my pinned post, like one of my first responses to it was, if you scroll down, I like respond under carpet at, yeah, the bottom video. Now you gotta go... That one. That's the one. I walk through kind of like how you actually access the dataset, which is in that drive.google.com link right below.

There's all of the traces. If you open up like a CSV, it's actually really interesting. You can take any of those. So, uh, yeah, um, there's the LLM responses CSV. Some of them are, are pretty beefy, but it has all of the reasoning traces.

It has all the diary. It has all the like data and that was, that was kind of the point is, you know, some people are like, "Oh, you don't wanna teach LLMs to lie," but you can definitely as easily add a rule saying, "Hey, no lying," or play only a Claude game, which I'm not gonna subject the stream to because it took till 1970 and still nobody won because they don't wanna hurt each other.

But you know, you can, you can easily change it into a... You might have to open it in, in a viewer. But you can, you can require, uh, no lies, and then you have reasoning traces of, of it not lying if you throw, you know, a classifier, um, to prevent it.

And I think that that's what's so cool about games is you can change it. I think one of the things that we'll do next is have three o4 minis versus Claude 2.5 Pros and allow them to, you know, create an alliance.

Essentially like, you know, change their system prompt so that they, they know if any of those powers win that they're, they're in a good spot. And yeah, there's just a lot you can do with, with this kind of, kind of parties.

**Host** [16:40]
I don't think I have a CSV that can handle a 78 megabyte, uh, file.

**Alex** [16:46]
That's all right.

**Host** [16:48]
Uh, yeah, but it's cool.

**Alex** [16:49]
In the video, I think I, in the video I think I open it towards the end of it.

**Host** [16:52]
Yeah. Awesome. Awesome.

**Alex** [16:54]
And yeah. Yeah, yeah. Just like you can-

**Host** [16:58]
Okay, cool. Anyway-

**Alex** [16:58]
'Cause you can view it in the actual viewer. And-

**Host** [17:00]
Others, others can see it.

**Alex** [17:02]
Yeah, there you go.

**Host** [17:03]
Oh, there you go. Yeah. Oh, just in numbers you can open it? Goddammit.

**Alex** [17:06]
Mm-hmm.

**Host** [17:07]
I know my numbers.

**Alex** [17:09]
Um, but yeah-

**Host** [17:10]
Okay, that's great

**Alex** [17:10]
... it's got all like the... Yeah, you got it.

**Host** [17:13]
Any tricks to like long range context management-

**Alex** [17:16]
Totally

**Host** [17:17]
... planning that you want to highlight?

**Alex** [17:20]
Yeah. Um, be opinionated. What-- I think one of the things that resonated from, from Sam was if you're looking at the context that's being sent to the language model, could you play the game? And that was really helpful.

And you'd think that that's obvious but, you know, it's more often than not it's very difficult, especially when you're playing a game like this. Representing the game board was tough, so you know, we were showing things like the units you control, the adjacency, right?

Like the things that are most relevant to the decisions that you have to make. And then diaries, so the relationships over time was really good. Also the, the updating of every phase. You have like updates after every negotiation, after every order, so that you know when you're reading the trade and then, you know, consolidation over time because as it, you know, as you went on in the game, you didn't need every single phase.

You can kind of consolidate it down once it's in the past, so that was big. And yeah, the personalities, the, the ability to go back and see the previous messages. We highlighted-

**Host** [18:14]
Hmm

**Alex** [18:14]
... like we highlighted important messages that you need to respond to, right? Like in case that there was somebody that, that was left on read. Some prompt engineering around Like if somebody's not responding to you, change your strategy, right?

Like these things that, that are opinionated and, and maybe, um, you know, some peop- a purist might, might push back on. But we try to keep it pretty, um, pretty minimal. And, and I think everything together ended up really coming...

And then just test, you know, just like constantly test and, and make sure that you know what changes you're making. So like have an experiment log and whether for you and/or the LLM, just make sure that you know the changes you're making and the impacts of them.

**Host** [18:49]
Any tooling that you've built or used that you wanna shout out for that kind of stuff? Experiment log, all this.

**Alex** [18:55]
Yeah. I mean, I think, and I don't know how I think anybody's tools really work. The things that helped for me is I just made like a lot... And I know this isn't what llms.txt was made for, but I just kind of like made a text file for like each of the folders that kind of had context of the files in that folder so that you didn't need to reference all, like flood the context window of like...

What I, I always used Windsurf for a lot of it. Post, post-Claude 4 coming out, I, I used a lot of-

**Host** [19:19]
Yeah

**Alex** [19:19]
... Claude code and that it's, it's, it's pretty great. The, uh, things that help me are like the llms.txt file and then experiment log files. I literally have like a /experiments folder and I just have like a prompt that I save in, in my prompt library that, that's like, "Hey, do exploration, keep your goal, understand what files are relevant, then track your experiment, and if you succeed, you know, I'll give you 1,000 bucks on your first try.

Otherwise, you know, you lose 100 per failure." You know, the, the classics. You have it ex- u-update its experiment log and, you know, really, really gets there.

**Host** [19:52]
Yeah. Amazing. Yeah, I have a... My, my name for it is one-off scripts, but then the one-off scripts then obviously always stick around, and then they, they become long-lived experiments and I wonder how to formalize that. I, I think, I think there's more tooling that can happen.

**Alex** [20:05]
Every-

**Host** [20:05]
I see you guys use Open Router a lot. Yeah, you, you need like some kind of routing layer. I've used LightLLM. I'm moving to Open Router. It really, like everyone needs some kind of core basic stack, right, that they're like is their thing.

**Alex** [20:17]
Totally. And I think that it's, you know, you can only focus on so much, so you do the things that are helpful to you. I, I know one of our GitHub issues already called us out for not using Pydantic parsing and instead having- ...

like eight different try catches for the different JSON versions that come out of these-

**Host** [20:31]
What?

**Alex** [20:31]
... LLMs. But you know, you, you start going down that hole and you're like, "Oh, I just need one more," you know, and so you just keep iterating, but we should, we should probably switch over.

**Host** [20:39]
Okay. Awesome. Um, all right, so broadening out from that, you did a talk. I, I haven't seen it yet because it's, it's, uh, it, it'll, it'll come out. Um, and we're prioritizing the best speakers, but you also won a sort of best speaker award for that.

Uh, anything else in, in that talk that you wanna preview at least for people to go find?

### Talk & Philosophy

**Alex** [20:56]
Yeah, I mean, the concept, the concept was benchmarks are memes and not, not your typical, not your typical meme, but in the concept that there are like ideas that spread, right? And I think that that's what's really interesting right now.

Simon Wilson, you know, he, he bit off a piece of my talk, yeah, afterward. No, I'm kid- I'm kidding. He gave a great talk where he showed that essentially there's, I think, a life cycle of a benchmark, right?

It starts with an idea, then it gets adopted, and it gets saturated. And what's so beautiful right now is a single person can have that idea. So Simon Wilson had an idea of, "I wanna see it draw a pelican on a bicycle," and then it got adopted.

People read it. He just kept doing it. It got adoption. Then models now s-saw that and then trained on it. And then now maybe it becomes saturated, but what that-

**Host** [21:36]
I don't know that they trained on it, but yeah.

**Alex** [21:38]
No, no, yeah, maybe not trained on it. Like evaluate on it at least. You know what I mean?

**Host** [21:42]
The problem is they really... I mean, his blog is a prominent blog, and they definitely index it, and when they u- when they update the, the cutoff window, it's definitely in there.

**Alex** [21:51]
Yeah. And, but, and I think that that's good, right? 'Cause like essentially if you think about it, what, what does that mean is that like you start, you can have, you can wonder, "How good is AI at this thing that I care about?"

And then if it gets adopted, then it, the most powerful tool ever created or, or one of them has now graded this thing that I care about. And like that's the whole idea is that right now, especially the people at AI Engineer, um, and the people that understand how to make benchmarks, I just encourage them to think about things outside of, um...

You know, you know, math, code obviously very important, but like, you know, and I talk about it in my talk, it's like I asked my mo- my mom teaches yoga and I asked her about like, "Hey, what are some things that you'd ask, um, AI?"

And she asked it five questions. We tried out five models. Turns out she liked Gemini 2.5 Pro the best, and she had a few things that it didn't cover. And so I was like, "Hey, just use this prompt at the bottom of, you know, whatever requests you make and it includes some things."

And now she has like somebody to bounce ideas off of and just helping her local community, you know, with, with like tailored yoga sessions. So like I think that that's cool, and I think that people outside of AI like...

We work a lot with people across the range, a spec- broad spectrum at Every for consulting and like they all have kind of like these two fears of like, "What's my role," and like, "How can I trust AI?"

And I think that benchmarks are both, right? Like you learn that your goal is to, your, your role is to define the goal and then you get trust 'cause you define it. You see how it does. You give it feedback, which literally might just mean changing the prompt and then you see it get better and so you understand your part of that process and, and I think that that's really cool and you get people who wouldn't have used it to use it and then some more stuff gets done.

**Host** [23:30]
Yep, yep. Um, I think that's an interesting, uh, I guess split between doing your own company specific eval versus doing it for the broader population like, you know, like Diplomacy, like Pelican Bench. That's kind of like let people understand the capabilities of models generally.

Whereas here, whereas for most companies, most people doing this for work, they're just doing it for their product and actually you don't super... I always think like this is s- this is something that Logan from the conference as well was starting to say is that the things that you want out of the box are actually not models, they're agents, right?

Meaning they're models-

**Alex** [24:09]
Yeah

**Host** [24:09]
... plus all the tools plus all the other-

**Alex** [24:11]
Yeah

**Host** [24:11]
... reasoning and plus whatever and I, I think that's kind of right in the sense that like when you are doing the broad population, general population evals, you want as minimal scaffolding as possible. Here's what DeepSeek's personality is.

You know, you're like, "Opus is too nice," whatever, right? Like you wanna make conclusions about the models. Whereas for the Uh, people doing benchmarks for their own internal use cases. This is why I separate- I have the dis- distinction benchmarks-

**Alex** [24:39]
Mm-hmm

**Host** [24:39]
... for general purpose, evals for product specific. And so when you're doing evals, you don't really care it's just the model, it's model plus scaffold, you know?

**Alex** [24:48]
Mm-hmm.

**Host** [24:48]
And a scaffold probably is, is, is as thick as you can meaningfully make it to, to make sense.

**Alex** [24:54]
Yeah. Yeah. That, that, that makes a lot of sense to me. And I do think, yeah, that there is an important distinction and, and that's why I think that there's so much room, right? Like there's some people are just gonna want something great and, you know, some devs are gonna want the bare metal to build their own great thing.

But I do think that a lot of people right now are a bit overwhelmed and worried and, and there's, there's fear out there and so anything that we can do to help them understand it, and that it is really something to help, like it does amplify you, it is leverage, it is not something that works on its own.

Um, and, you know, maybe some people would say not yet. I, I think that it, it, it'll kind of be, be a while and, and I think that people care about what people care about, so if we help them do the things that they care about, you know, that's, that's great.

**Host** [25:37]
Do you have one for... Like, what do you think about the other un- hard-to-test, uh, things like creative writing? Like you, you talked about-

### Creative Writing

**Alex** [25:44]
Mm

**Host** [25:44]
... Sam Peach's, uh, stuff.

**Alex** [25:45]
Mm.

**Host** [25:46]
Obviously very close to home for Every.

**Alex** [25:48]
Mm. Mm.

**Host** [25:49]
Close to home to me as well. There is a creative writing model-

**Alex** [25:52]
Mm-hmm

**Host** [25:52]
... I haven't had access but, you know, you've, you've heard rumors.

**Alex** [25:55]
Yeah.

**Host** [25:55]
Yeah, just like what is the future of creative writing, you know, and like, uh, how would you measure it?

**Alex** [26:01]
Yeah, I mean, we, we definitely use it a lot at Every, but it's not like, you know, put something in-

**Host** [26:06]
Wait, wait

**Alex** [26:06]
... you get something out.

**Host** [26:07]
You have the GPT creative writing model?

**Alex** [26:09]
I'm gonna say that-

**Host** [26:10]
You don't-

**Alex** [26:10]
... like we use AI a lot at Every, you know what I mean?

**Host** [26:12]
Okay. All right. All right.

**Alex** [26:13]
And, and we just like... And what I mean is it's not like put something in, get something out, it's like we have great editors, and our editors have been doing it for a long time. And like as somebody who...

I, I made a lot of video content, um, but I didn't do a ton of writing before coming to Every. I learned a lot from my editors and, and there were some things that they would tell me every single time, right?

Like any- like focus on these. Um, you know, Dan really drives in thinking about audience and story and, and how you can improve that. And so how can we take that? Katie, one of my colleagues, is doing great work here.

She, she's putting together these things, like these like kind of packages of, "Here are some things that you can think about. Here are some things that we can give to the LLM that are tailored to you." And everybody has their own process.

For me, I think a lot when I talk, so I'll use a lot of dictation and I'll have prompts set up that are, have sections for dictation, sections for things I've written in the past that I really love, sections for things that we think about as Every, like our style guide, my editor notes, maybe a transcript of a call between myself and my editor, and then use that to then end with like, "Ask me questions that once answered will help me take a first pass at this," right?

Like to help create a scaffold. And so then you do that, and then you go back and forth, you constantly edit. That's something I really emphasize in, in consulting. If you're not editing, if you're not reflecting on a message coming out and editing above, then like that's something you need to do because you don't want to pollute the context with anything that's not exactly what you're trying to say.

And so, you know, I'll edit messages like 20 times and, and dictate and add in there more context because it's not one-to-one. You know, there's translation happening, right? There's encoding, there's decoding. It looks like English, but you know, it's working in token space, and so being able to reflect on that and understand and build your intuition is big, and then you just kind of like iterate.

And a lot of the time I take out ideas and then I'm rewriting big chunks of them and then my editors go through it actually after, you know, having our AI editor take a pass and, and so it changes a whole lot afterwards, but it helps consolidate ideas and at least for me, that's something that works.

But everybody does it differently. That's not even standard across our team.

**Host** [28:15]
Yeah. Um, that would be interesting. Would-- I have dreams of hosting a writing retreat that, um, you know, it's like how to use AI to leverage your writing but still write like, like a human.

**Alex** [28:27]
Yeah.

**Host** [28:27]
So it's called Write Like a Human.

**Alex** [28:29]
I love it.

**Host** [28:30]
Uh, yeah.

**Alex** [28:31]
That's great.

**Host** [28:32]
I, I, I held one, I held one like three years ago but, uh... And people really liked it. I just need to sit down and do things. Yeah, I'm, I'm kind of out of the event organizing business for a while.

**Alex** [28:40]
I was gonna say. I was gonna say you've been not busy at all recently. You want, you want to put on an event, huh?

**Host** [28:46]
No, it's just like I think that I would like to learn for others. I, I'm not super AI pilled in my own writing and I, I, I image- I know I could be. And generally like, uh, one of the failed tracks of the conference, uh, now I get to like spill the tea on like what didn't happen.

We had this thing called AI in Action and it was supposed to be you talking, you as an excited user of another person's product-

**Alex** [29:09]
Mm

**Host** [29:09]
... talking about how it improves your life-

**Alex** [29:11]
Mm

**Host** [29:11]
... and talking about how, how to be more productive with it. But it, it turned out to be everyone applying for that was just shilling their own product, so we killed it. But, um-

**Alex** [29:21]
Yeah. There was a little of that happening. But I think that some of the best talks weren't that, right? Like some of the best talks-

### What's Next

**Host** [29:26]
Yeah.

**Alex** [29:26]
Man, I'm, I'm, I'm gonna for- forget her name. Uh, Isa from NotebookLM.

**Host** [29:31]
Yeah. Hux. Yeah.

**Alex** [29:32]
Yeah, her talk was great. You know, her talk was awesome.

**Host** [29:34]
Yeah, yeah.

**Alex** [29:34]
Talked really about clarity and how you need to have clarity and understanding yourself and then that's gonna give you energy to build an awesome product for somebody. And, you know, I so resonate with that. Like, I think that one of the hardest things but most important things you can do in life is like have a goal in and of itself if you accomplish it makes you happy, right?

Like one of the hardest things to do, but if you do that... If you don't do that, then you don't know what's the right decision because everything is downstream of the goal. I think, I thought Sarah Guo's talk was, was great.

You know, she, she, she, she handled, she handled it like a champ.

**Host** [30:06]
Well, they're well. So sorry for her.

**Alex** [30:08]
No, I mean-

**Host** [30:08]
We basically didn't rehearse that transition properly.

**Alex** [30:11]
But, I, I mean, there were so many people talking and I think that that's one of the trade-offs you guys made where it was just like so much awesome information was coming at once. You were able to have so many people there, so many great speakers and, you know, so not everyone can do a rehearsal every single time, but I think that that's totally fine.

And honestly, I think it made it a little bit more memorable and, and I think that she came out looking really good and I, and I remember a lot of, a lot of that talk, um, as a result.

**Host** [30:34]
Yeah.

**Alex** [30:35]
And I thought that-

**Host** [30:35]
Yeah

**Alex** [30:36]
... you know, constraints are, constraints are good sometimes.

**Host** [30:39]
For sure. For sure. Um, she had-- Yeah, there, there was, there was... If it was gonna happen to anyone, it should've happened to her 'cause she was able to handle it. Yeah, cool. Uh, well, thanks for jumping on.

I, this is a very short notice, but, uh, you know, I think it was just-- I like to be timely when the thing is still hot, and when you still want feedback, and the feedback actually matters to you that we record this, get it out there.

People who are like, "Oh, I wanna contribute," like this is like a, a good time for it.

**Alex** [31:02]
Please.

**Host** [31:02]
Then they can at least get in touch, you know.

**Alex** [31:04]
Yeah. Three things I wanna do now. We wanna make a data viewer, so you can actually go-

**Host** [31:07]
Yeah

**Alex** [31:07]
... not have to load a massive 58 gigs-

**Host** [31:09]
And all the JSON is here, right? Just, just go nuts, right?

**Alex** [31:13]
Yeah. And, and then we wanna improve the front end, and I wanna make it playable. I'd love to have a human versus AI Diplomacy tournament, and I think that we'd learn a lot from that because I think that it's kind of a mix between Diplomacy and, and I wonder if, uh, you know, we get some, some prompt engineers to jailbreak it and win in like five turns.

I would be fascinated to see if that, if that happens or not.

**Host** [31:34]
I think you-- Yeah, I think you, you get Noam Brown, who won human Diplomacy, and then you also get like one of the jailbreakers, like-

**Alex** [31:40]
Yeah

**Host** [31:41]
... Rye on Good Side or Elder Planus.

**Alex** [31:43]
Yeah.

**Host** [31:43]
Who I think is New York, is in New York, by the way. Um, yeah, maybe I'm not-

**Alex** [31:47]
No, I mean, I think it'd be so cool because I think that we'd learn a ton from the prompts, you know. And then, like that's, that's one of the things. This is, this is why I'm so excited about the project, like it-

**Host** [31:55]
Okay, got it

**Alex** [31:56]
... everybody learns. But yeah.

**Host** [31:57]
We would, we would-- How long does a game take to, to play out?

**Alex** [32:01]
Um, it-- I mean, it really varies. Some last 20, 20 years versus like some last 70 years in the game. And so like-

**Host** [32:08]
No, but human time.

**Alex** [32:09]
No, human ti- human time watching it, I think it's like anywhere from 4 to 12, 4 to 18 hours, something to play out.

**Host** [32:18]
Jesus. Oh my God.

**Alex** [32:19]
Yeah.

**Host** [32:20]
How the, how the-

**Alex** [32:20]
I think, I think, I think we'd have to use voice as your input. Uh, that's your only way that you're gonna get-

**Host** [32:23]
Okay

**Alex** [32:23]
... enough context out, and-

**Host** [32:25]
Okay

**Alex** [32:25]
... then you do it, and then it's-- maybe it's not... It's either who can win in the shortest amount of time or get the most pro-- like who can get the most territories, so supply centers in the allotted time, you know what I mean?

So you can-

**Host** [32:38]
Okay, okay

**Alex** [32:38]
... you can cut it off.

**Host** [32:40]
Oh, there's a cutoff. Okay, okay, okay. That's good.

**Alex** [32:42]
Yeah, there can be. There can be, yeah.

**Host** [32:43]
Okay. We might-- I might wanna like host like a thing. Uh, here I, here I go again. But like we know Noam. Noam's our next guest on the podcast.

**Alex** [32:52]
Yeah.

**Host** [32:52]
He'd be down. He'd be so down.

**Alex** [32:54]
No, I'd, I'd love to do that.

**Host** [32:55]
Yeah. Four humans, you know, uh, three robots or some, some mix of that, like that. Uh, you could, you can make that happen.

**Alex** [33:01]
Yeah, that'd be cool. And I-- Yeah, and lead into like-- I'd love to have like a tournament too. You kind of like incentivize people to like, you know, compete, you know, like some, some kind of like, um, I don't know.

**Host** [33:11]
It's hard to compete. Like, I mean, how do you, how do you have-- There's only seven players.

**Alex** [33:16]
Right. Well, I, I think what you do is you have a single human play against the same AIs-

**Host** [33:20]
All the bots

**Alex** [33:21]
... and-

**Host** [33:21]
Yeah

**Alex** [33:21]
... who can win-

**Host** [33:22]
The highest levels. Yeah

**Alex** [33:23]
... in the shortest amount of time or like, you know, cap, cap it to 12 hours. Who can get the most amount of supply centers in 12 hours? And you're the, you know, you're the champ, right? Because like you've won.

You know, and, and maybe you get some sponsors. You, you let whoever wins. You know, you probably have to probably, because the tokens are expensive unless we get like a ton of sponsors or something, but like then you split the pot with, you know, whatever, whoever wins or something.

**Host** [33:46]
Yeah.

**Alex** [33:46]
Um, but I think it would be fun.

**Host** [33:48]
Okay. All right. Well, if you're interested in any of that stuff, uh, get in touch with Alex Mostly. I would help on the community side if, if, if we can sort of get some in-person event together. Otherwise, it could just take place online.

You know, it doesn't have to be a thing.

**Alex** [34:00]
I think both, both would be good.

**Host** [34:02]
Yeah, yeah. Uh, well, great to have you. Thanks for putting together all this, this stuff. I'm looking forward to the next year of work, and then like, you know, hopefully every AIE, we just like come together and like, uh, talk about what, what the latest thing is.

**Alex** [34:13]
And back at you. Thank you so much for putting on the conference. I, I thought it was beautiful.

**Host** [34:16]
Yeah.

**Alex** [34:16]
Got to meet some really cool people.

**Host** [34:18]
Yeah. Thank you.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
