# Beating OpenAI and Anthropic by Looking At Data: the new #1 on SWE-Bench w/ W&B CTO Shawn Lewis

Latent Space · 2025-01-28

<https://addtry.com/e9853e1f-b6ad-449d-8244-deb0b9004c6e>

Shawn Lewis, CTO of Weights & Biases, built a coding agent using OpenAI's o1 model that achieved 64.6% on SWE-bench verified, the top score at the time. He developed a TypeScript agent framework called PhaseShift and an Eval Studio tool backed by Weave, emphasizing that rigorous data inspection and tooling are critical for agent development. Lewis explains his process of running parallel rollouts and a crosscheck method to select the best trajectory, which contributed a 6% improvement. He also discusses the challenges of adapting reasoning models like o1 for agentic tasks and shares plans to open-source PhaseShift. The episode highlights Lewis's dogfooding approach, using Weights & Biases' own tools to build and evaluate the agent.

## Questions this episode answers

### How did Weights & Biases achieve the top score on SWE-bench verified?

Shawn Lewis, CTO of Weights & Biases, built a SWE-bench verified agent achieving 64.6% by using OpenAI’s o1 model, an Eval Studio for analyzing experiment data, and a TypeScript framework called PhaseShift. He also used a Crosscheck technique that runs multiple trajectories and selects the best one, yielding a 6% improvement. The work was done partly to dogfood W&B’s own tools like Weave.

[0:56](https://addtry.com/e9853e1f-b6ad-449d-8244-deb0b9004c6e?t=56000)

### What is Crosscheck in AI coding agents and why is it controversial?

Crosscheck is a method where multiple agent trajectories are generated for a problem, then a selection step picks the best result. Shawn Lewis’s implementation gave a 6% boost on SWE-bench verified. Some competitors dismiss it as expensive rejection sampling, but Shawn argues all top submissions use similar strategies, and it’s necessary to tackle the hardest problems at the top of the leaderboard.

[28:04](https://addtry.com/e9853e1f-b6ad-449d-8244-deb0b9004c6e?t=1684000)

### How does o1 compare to GPT-4 for building coding agents?

Shawn Lewis found that o1 is extremely good at programming and follows instructions precisely, sometimes to a fault with “malicious compliance.” However, o1 was harder to make agentic compared to GPT-4, which is better at long sequences of steps and reasoning over prior actions. For his agent, he tuned o1 to handle both the step-by-step logic and code generation, achieving leading results.

[13:30](https://addtry.com/e9853e1f-b6ad-449d-8244-deb0b9004c6e?t=810000)

## Key moments

- **[0:00] Intro**
  - [0:45] Shawn Lewis achieved 64.6% on SWE-Bench verified and published the result just days before his son's birth.
- **[1:44] SWE-bench Motiv**
  - [1:44] Weights & Biases CTO Shawn Lewis built the agent to dogfood W&B tools and explore AI applications.
- **[4:57] Tooling**
  - [4:57] Q: What custom tools did Shawn Lewis build for SWE-Bench? A: Eval Studio for comparing evals and PhaseShift, a TypeScript agent framework.
- **[6:54] Headline Results**
  - [7:54] Shawn Lewis spent months myopically focused on SWE-Bench, which led to building Eval Studio and other supporting tools.
  - [10:32] Shawn Lewis: 'Everything in spreadsheets,' as he initially tracked evals in Google Sheets despite being W&B CTO.
  - [12:00] Plotting reasoning tokens vs cumulative resolved instances in Eval Studio revealed distinct shapes that guide prompt improvements.
- **[13:08] o1 & PhaseShift**
  - [13:46] Shawn Lewis's agent uses OpenAI o1 for both reasoning and coding, achieving 64% on SWE-Bench versus OpenAI's own 49%.
- **[17:02] Data-driven dev**
  - [17:19] Identifying instances where new evals failed but past evals succeeded—especially those universally solved—accelerated progress.
  - [18:47] Shawn Lewis added a difficulty column counting public leaderboard solves; many SWE-Bench problems remain unsolved by any published agent.
  - [21:47] Shawn Lewis: 'You basically always have to look really closely at data; you cannot avoid manually inspecting these traces.'
- **[22:24] Debugging**
  - [23:33] Weave tracks code and configuration, letting Shawn diff eval runs to pinpoint prompt changes causing regressions.
- **[25:04] PhaseShift Release**
- **[25:59] Real-world Use**
  - [25:59] Q: Can PhaseShift, the SWE-Bench-topping agent, rewrite its own code? A: No—it can currently only operate on SWE-Bench problems.
- **[27:29] Crosscheck**
  - [28:18] SWE-Bench requires trajectory submissions, but the data goes into a private S3 bucket, making it hard for others to access.
  - [30:14] Skeptics vs Shawn Lewis: Is Crosscheck just expensive rejection sampling, or a novel trajectory selection method?
- **[32:45] Future Outlook**
  - [32:54] Shawn Lewis predicts that autonomous AI programmers will work really well within one to two years.
  - [33:16] Shawn Lewis says if there is a moat in AI coding, it's in interfaces for humans and businesses, citing Devin's UI lead.
- **[34:10] Outro**

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **Shawn Lewis** (guest)

## Topics

Agent Platforms

## Mentioned

Anthropic (company), Weights & Biases (company), Aider (product), Cursor (product), Devin (product), Eval Studio (product), GPT-4 (product), PhaseShift (product), SWE-Bench (product), Sonnet (product), Weave (product), o1 (product), r1 (product)

## Transcript

### Intro

**Alessio** [0:00]
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO at Decibel Partners, and I'm joined by my co-host Swyx, founder of Smol AI.

**Swyx** [0:09]
Hey, and today we are doing another quick episode on breaking news. We have a new SWE-bench king, and he, part of his, uh, part of his, uh, appeal is also that he has a great name. It's, uh, Shawn Lewis from Weights & Biases.

Welcome.

**Shawn Lewis** [0:22]
S-H-A-W-N Shawn. It's rare to meet another S-H-A-W-N Shawn. Thanks for, thanks for having me.

**Swyx** [0:27]
Yeah. Yeah. Uh, Weights & Biases are very, very good friends. We've, uh, collaborated on a bunch of things over, over time. And so imagine my surprise when apparently during your paternity leave or just before, you release this top, uh, coding agent in, on SWE-bench verified.

**Shawn Lewis** [0:45]
Yeah. That, I mean, it's also somewhat to my surprise. I, I've, I've been working on this for a few months and knowing that I sort of had this deadline of having a new child join the world. So it really was down to the wire.

Like, you don't exactly know when a baby's gonna be born. And when you're doing work like this, it's very experimental, right? It's not like I had that 64.6% result even like a few days before I published it. So my son was born, I think the day after, two days after we, we published that result.

**Swyx** [1:13]
Oh, amazing. Okay. Awesome. Uh, so like, yeah, w- a hell of a way to go on to paternity leave. Congrats on your son. That is always a better achievement, uh, reproducing humans than with AI. But we're here to talk about, talk about AI.

Uh, yeah, so like, you know, I, I guess my general view is that I did not expect this coming out of Weights & Biases. Like, Weights, you know, Weights & Biases is very well known. ML ops tooling, you know, recently you have Weave and, uh, you're doing a lot in the sort of LLM evals and ops space.

And then suddenly, like what, what got you to start working on something like this?

**Shawn Lewis** [1:44]
Yeah. Um, well, so I mean, first of all, I think, I think that the best way to build great tools is to be a user of them yourself. Our original experiment tracking tools that, you know, is still what we're known for the most today, to some extent, we were able to, to have success there because both Lucas and I had done lots of machine learning training ourselves.

### SWE-bench Motiv

**Shawn Lewis** [2:04]
So Lucas, um, and Chris are my other two co-founders. Years and years ago, I got really interested in this problem of gaze tracking, which is, you know, from a webcam, where is the user looking on the screen? And this was like before deep learning.

Um, I basically just wanted to have like, you know, somebody-- I was working at Google, and people would come up and shoulder tap me as I was like doing work, and I would look away. And I just wanted my screen to tell me where was I last reading in the document.

And still today, computers don't do this. So I think like somebody could like make an app that does that, and it would be useful because it, you would be able to start reading, um, where you left off faster.

**Swyx** [2:37]
Yeah, I think that-

**Shawn Lewis** [2:37]
So I worked on that for a long time, and it was very hard prior to deep learning. Um, but I'm like naturally a tool building kind of person. So like in the course of doing that, it was like I had spreadsheets and like, you know, you're like iterating on code that does all these different algorithms, and then you're losing track of your like prior code that actually had some success.

And so basically, I was building sort of experiment tracking tools as I did that. And then eventually when we started Weights & Biases, those concepts were like first and foremost in my mind. So I hadn't had a chance to really deeply do that with like newer, like just purely by using LLMs.

Um, of course, I use LLMs all the time, both for programming and, you know, like everything like everybody does now. Um, but I hadn't really had a chance to try to really build something. And I want, I want a challenge, you know, I want something that like, that I might be uniquely suited to do, and, you know, I'm a programmer, so that, so trying to automate programming is, is a good place to start.

So just really like a fascinating problem and yeah, like a good hard challenge. So partially it's like, let's dogfood something. Let's like really, let's really try to build applications using AIs in our own tool, using AI in our own tools.

At the same time, like, you know, as um, uh, as AI improves, so we sell tools for, first of all, for people who are training AI models, so experiment tracking. Um, as AI improves, like fewer people probably need to train models from scratch.

It gets concentrated more and more in, in different, like in the big labs. And so that's a really great business for us and like they, you know, they continue to like expand their usage of our tools and there's a lot more that we could build there.

But the, the, you know, our mission is to like make the best tools for AI generally. Now you can get so much value from AI just by calling APIs. So we switched into like building Weave and I, you know, I played a, I, I've done a lot of work on that myself also, um, especially in the beginning of it.

And then, you know, eventually it's like AI keeps improving, so where is it, like at some point, do you even need to, how much do people need to customize like AI using APIs? Probably, I hope a lot in the future.

I hope that we don't like, you know, I hope that we're not like, uh, totally automated out of jobs in the next five years or 10 years, but who knows? But you know, there's this level beyond that, which is applications.

And so the other kind of, uh, the other reason to do this for us is, you know, should we dabble in that world of applications and, um, you know, the most fascinating one and interesting one to me is AI programming.

So there's like potential that we might, you know, sell products in that space, but that's unknown so far.

### Tooling

**Alessio** [4:57]
Right. So we did an episode with Anthropic about their SWE-agent work and like the SWE-bench verified results that they had. One of the big focuses they had was tooling. So you mentioned tooling obviously. I saw in your X thread about the, the agent that you obviously used Weave and some of the stuff you already had, and then for example, you built a eval studio that like you didn't have before.

You built a TypeScript agent framework called Playshift. Like, yeah, maybe talk a bit about-

**Swyx** [5:21]
Can we, can we pull it up on screen? Yeah. We, we should show, try to show, not tell.

**Alessio** [5:25]
Yeah. It's-

**Shawn Lewis** [5:25]
Yeah. Sure.

**Swyx** [5:25]
I think Shawn has it, so we might as well just give Shawn the screen-

**Alessio** [5:28]
Yeah. Yeah. Yeah

**Swyx** [5:28]
... because otherwise we'll be switching back and forth.

**Shawn Lewis** [5:30]
Let's see here.

**Swyx** [5:31]
But yeah, I mean, generally while he sort of pulls it up, I'll buy time, which is like, yeah, total dogfooding, right? At the same time, like you guys are, are very good at experiment tracking and, and, and also like running evals has always been a part of the workflow of training models, right?

This is nothing new in a, in a way, but it, it is new in a sense that it is a coding agent, which is not something typically that we've really had to do end-to-end evals for.

**Shawn Lewis** [5:53]
Totally. Yeah. I mean, b- building anything with AI is this, is a very experimental process, just like training models. So like, you know, the way that I was able to do this basically is run a ton of experiments.

Um, and people use our Lots of people actually use Weights & Biases prior to the advent of, you know, um, more like LLM application tools like Weave and, and others for, for evals. And a lot of our customers today still continue to use the Weights & Biases workspace for, for tracking just evals and not the process of training.

The workspace, the W&B workspace is a really powerful tool for digging around and analyzing the results of evaluations. So yeah, I mean, do you guys wanna see, do you want me to like walk through some of the stuff that I made here?

**Swyx** [6:38]
Sure. Can we, can we do the, the headline first? Like, uh, what should people know? Uh, 'cause th- there's, there's the h- top level thing, and then the, the headri- headline results. 'Cause I, I feel like we might be a bit too in the weeds already.

**Shawn Lewis** [6:50]
Sure. Um, let's see. Headline results.

### Headline Results

**Swyx** [6:54]
Yeah. 'Cause you mentioned in, in process. Yeah, go ahead.

**Shawn Lewis** [7:01]
Do you mean the headline results like, like what did I make here, like and why, basically?

**Swyx** [7:05]
Yeah. Yeah, yeah.

**Shawn Lewis** [7:07]
Sure. Let's see. So in the course of trying to basically like build an AI programmer, you know, first of all... At first, I just built this AI programmer. I've been working on that problem separately for a long time, and there's lots of cool stuff along that path.

But, you know, to really know if something's working, you have to evaluate it at scale. These are like stochastic processes, right? So if you make it work on your own code base for a little while, like i- it might be good at doing some narrow slice of problems.

But, um, uh, evaluating on like lots of problems in parallel is how you kinda like make sure that you're building something general. So SWE-Bench is, you know, the, the best eval that we have for AI programming today, and the most publicly known one.

And so I, I basically decided to work on that, and I spent lots of time on it. The, the work that I've done over the last few months is really all around like just improving results on SWE- SWE-Bench, almost like myopically focused on that problem.

But in the course of doing that, I built a lot of, you know, other tools to support that, that effort. So let's see. There are...

Sorry, I'm trying to think of how to tell this story here. Yeah, I guess the, the like this process of running evals, I guess I, I'm not sure how you-- I'm pretty sure a lot of your listeners will be familiar with this, but it's like, it's this very experimental, like scientific process, right?

So I try like a bunch of prompts to like take a step in an agent, and I, I generate a bunch of results and data. Then I need to kinda like dig into that data and decide, did it work?

What should I do next, right? That's the, the loop of research. And so the, the more that you can like see into the data as you, as you work, the more visibility you have into the actual things that the agents are doing, like at a fine grain level, the better decisions you can make along the way.

So, you know, I didn't really have that visibility, um, in the beginning of doing this. I was using our tools Weave, which are really good at like tracing and, and actually running evals for you. But the views that we have in there aren't really quite as tailored to like agent workflows.

Um, and it's, you know, it's missing some stuff. So, so I'll show this sort of like front end. This is what I was calling Eval Studio. And, uh, this, this lets me do a few different things. So basically, um, what we're looking at here is the, the different, uh, colors.

The different colored lines are three different evals that I ran on, um, a subset of SWE-Bench verified here. So I've like loaded up three evals in this, and this kinda looks like the W&B workspace. You can do a lot of the things that you can in the workspace.

So I'm like filtering down to the resolve chart here. Each points on the chart is like a agent result on a single instance within SWE-Bench, um, an instance being like a GitHub bug in a GitHub repo that the agent tried to solve.

So each point is the agent result on that instance and the scores for the instance. Um, and so this chart here is like the final score that we're looking for on SWE-Bench, whether or not, this is whether or not like the problem was solved.

Um, and so before I... Oh, I didn't pull up the spreadsheets.

Everybody's gone through any of this work without tools, like spreadsheets are such a great tool, right?

**Swyx** [10:31]
Spreadsheets are all you need.

**Shawn Lewis** [10:32]
Everything in spreadsheets. And so I've done lots of work in spreadsheets despite being the CTO of Weights & Biases. Let's see here. This is, uh-

**Swyx** [10:43]
What is this?

**Shawn Lewis** [10:45]
So this is, uh, this is the tool that I was using prior to having this kind of like workspace tool to plot as these evals were running, how were they doing against prior evals. So a SWE-Bench eval for me takes about an hour to two hours to run on like a subset of 100 problems.

And, you know, you kinda wanna know as it's going, like what's happening. Partially like to decide whether or not you should actually continue running that run. You know, also just 'cause you're excited to see what the result will be.

So this spreadsheet is what I had some other tools where I would have like columns of data for each eval, and then I would have to go to the other tool and copy a column out and paste it in here, um, and, you know, make like these little plots.

Eventually, that got replaced by this, this workspace where I can just load up data. I should say all the data in this is backed by Weave, so that's our tools for, for like applying AI. Everything from my FaceShift framework, which is a, a new TypeScript framework for, for building agents, is logged to Weave automatically.

Uh, so there's basically no backend for this new Eval Studio tool that I have here. It's all backed by Weave, um, data.

So let's see. So yeah, um, this, this view is really useful for looking at charts as things run, and then comparing runs at kind of a high level to understand what happened. So another really useful view to look at is if you plot reasoning tokens on the x-axis- And then like a cumulative sum of what instances were resolved.

Over time, you can sort of see different shapes in these charts. So this blue line here kind of has a kink in it. If you look at how it looks like it's going up really steeply, and then it shoots out to the right.

These problems to the right of that kink are problems that, uh, that use significantly more reasoning problems than all the ones running up that-- reasoning tokens than all the ones running up that slope. You know, there's a question here.

I can't, like, answer this without digging into the data, but it's like, why were there some problems that seem to be, like, very easy in terms of the amount of reasoning, the amount of thinking that you have to apply to solve them, um, versus others that took much more reasoning?

So it's really useful to be able to just sort of dream up something to plot and be able to actually plot it and see it, um, in a UI like this.

**Swyx** [13:08]
Good. Do we want to run through also the models that you built on? Because I think the, the headline is kind of o1 driven, but then I know you use a bunch of sub things. And maybe also talk about how you kind of wrangled them together into this and talk about phase shift and, and that.

### o1 & PhaseShift

**Shawn Lewis** [13:24]
Yeah. Yep. Yeah, maybe that's the headline that you're looking for. So,

yeah, so I, I guess, uh, I think for me, like, using the latest and greatest is, is like one of the major reasons, the other major reasons to do this. Um, reasoning models, so I, I, this, this-- the agent that I made here is, is purely based on OpenAI's o1 model.

Um, and I think it was the first, you know, published result on SWE-Bench that, and, and maybe the only one, you know, still as of a week later, that uses o1 to both drive the agent logic, so the step-by-step logic of what should we do next, and like all the programming that happens.

I think, you know, it's also just incredibly valuable to us at Weights & Biases to understand how the new models work, right? And, and applying o1 is, is very different than, than using GPT-4. It does different things. Like, it's very good at s- doing exactly what you say, and sometimes almost like, you know, doing m- malicious compliance where you, you like had something in your prompt that told it to do something, and it really complies to what you said, um, despite there being like maybe a better alternative.

Where it feels like to me something like GPT-4 will sort of ignore stuff and just like keep plodding down the same path. So, uh, this agent is, uh, purely based on o1 and, you know, there was a lot of like experimentation and learning to do to, to, to like make o1 work in this way.

I think OpenAI, they published their SWE-Bench verified o1 result when they, they released o1, and I think they solved something like forty-nine percent of, uh, problems. So, so this solves, uh, sixty-four percent of problems. It solves like something like fifty-seven percent of problems with a single rollout and then using parallel rollouts and selecting the best one.

With other techniques, we get something like sixty-four percent.

**Swyx** [15:14]
Right. So I think did you compare this with the other agent approaches that the other folks adopted, or were you just kind of doing it first principles?

**Shawn Lewis** [15:23]
It's all done from first principles, and I really would like to... I, I didn't get a chance to actually run, um, Sonnet through this all the way, so I don't know what the result would be if I just dropped Sonnet in here.

Just dropping a model in is not really like ever possible. You kind of have to like tune things. You know, like I would have to do a bunch of experimentation to give it a fair try. I actually just didn't have like a, um, high enough rate limit on Sonnet at the time I was doing this, and that's why I like didn't end up running it a lot.

**Swyx** [15:50]
Yeah.

**Shawn Lewis** [15:50]
Yeah, this is... Go ahead.

**Swyx** [15:52]
Well, uh, the, uh, the, the sort of apparent meta game from Paul Gauthier on Aider is that you use R1 as an architect and Sonnet as a code model, and that's apparently the best combination that, that beats o1, uh, for him.

**Shawn Lewis** [16:05]
Yeah, I'm surprised about that actually because, um, I think that it's actually, I think, you know, o1 is very good at programming, but it's kind of the agent part was the harder part to get it to do here.

I think it's, it's like less trained to take the next step in like an agentic task. Whereas GPT-4 for like the last two years has been really, you know, pretty decent at like taking a, a long sequence of steps to solve a problem and being able to refer back to like prior steps in the right order and stuff like that.

Whereas o-, it feels like it, it kind of gets confused. So yeah, I'm surprised. Like, I would actually flip that around if I was gonna try, try that, you know, to start. But you know, it's like, again, you have to experiment a ton and, and that's probably where he ended up.

**Swyx** [16:47]
Well, it's also a different benchmark. He's, he's using it for sort of code editing, whereas this is much more, quote-unquote, agentic, for whatever that word means. But yeah, it's, uh, it's a, it's a really good result. Any, any other sort of, uh, approaches that you found out meaningful?

So it's like wh- when you were developing, you had this little proto framework inside of Google Sheets. What's one thing you noticed that sort of changed your mind as to like your, your agent development? I s- basically I'm looking for example where basically kind of looking at your data as it evolved actually affected the way you developed the agent.

### Data-driven dev

**Shawn Lewis** [17:19]
Sure. Yeah, I mean, I can give you kind of like the moment that, that really maybe accelerated my progress when I got these-

**Swyx** [17:25]
Yeah

**Shawn Lewis** [17:25]
... soles working well. Oops. So let's go back to this tab. So, uh, this is like a simple table where along the rows we have each of the SWE-Bench problems in the dataset that we're looking at, and we can load up other columns from the three evals that we've selected on the side.

So what I typically do, and there's all kinds of ways to use this, but what I typically do is I just load the, the final score values for each of the three evals next to each other, and then sort these columns like this so that I can say, for example, "Let me see cases where this new eval," let's say I just ran the one on the far right, "failed on problems, but where the, the prior evals that I've done succeeded."

And sometimes I'll load like ten other evals here, and you'll have like everybody was getting it right, right, right, right, right, and then, you know, our new one got it wrong. So this is an interesting problem to look into because Because it may be something that was very easy until we changed something in the prompter, in the framework, and now it's broken.

The other important thing to note here before I, like, dive into the traces is this column was also very useful to me. Everybody who's ever working on this should do what I did here. So basically, I took all the public SWE-SWE-Bench leaderboard data, and I just counted how many times every instance was solved by somebody on the public leaderboard.

So that basically gives you a difficulty score for every instance. So if we go down to the bottom and see like all these zeros here are SWE-Bench problems that have never been publicly solved on a leaderboard, and there's actually quite a lot of those, um, problems remaining.

You know, like, so if we look at this problem here, this is one where basically everybody on the leaderboard solved it. My two prior evals solved it, um, and yet this eval didn't solve it. So, like, what's going on there?

And once I had the tool, like really able to load these up and, um, and look at that view of just like at the example level, like what, what things am I doing poorly on that other prior runs have done well on, I started making a ton more progress.

So from there, I can really easily dig into the, uh, the agent trace.

And I've picked a run. This is-- I might need to load another one for you. Well, let's just look at the other one as an example.

**Alessio** [19:42]
And so the idea is that you run all these through Weave and Weave stored it all for you, so it's easy to pull up, right?

**Shawn Lewis** [19:48]
Yep. All this data-

**Alessio** [19:48]
It's the same backend. Yeah

**Shawn Lewis** [19:49]
... is like permanently kind of stored in Weave here-

**Alessio** [19:52]
Yeah, yeah

**Shawn Lewis** [19:52]
... in W&B so.

**Alessio** [19:53]
So it's kind of like a different view of Weave that you've built, you know.

**Shawn Lewis** [19:56]
Exactly, yep.

**Alessio** [19:57]
Like a different, different front end. Yeah. Okay.

**Shawn Lewis** [19:59]
Yeah. So here in Weave, we see like these are the raw eval logs, um, that I've done. You can see in the course of this, I did something like a thousand, um, evals. And there's all this data in here, and Weave has lots of tools for digging into traces and like looking at a fine grain level at this stuff.

But, uh, this tool gives me an easier way to pivot the data and then dig in to get a more agentic, call it, view of what's happening. So maybe we can define that a little bit as we, like, look at this.

What this shows me is the linear sequence of steps that my agent took to solve one of these SWE-Bench problems. And so at each step, I can see what the LLM, the complete LLM output was, what the tool calls were that were made in that step, and then what the state of the current, um, observation is here on the right.

And for me, uh, the way that this, this agent works is, um, it has an editor that, um, it can see like code that it's opened here in this observation. Um, and so as we scroll through here, you can see like here, the agent decided to expand a function in the editor.

So in this editor, everything starts off collapsed, and we could turn on line numbers to see that. So the agent actually sees every line in this Python file prefixed with line numbers that look like this. And when there is a range, it, it indicates that that function was collapsed, kind of like, you know, code collapsing that you would have in, in VS Code.

Um, so the agent, you know, expanded this function so that it could see what was in there. Let's turn that back off. Then as we go to the right, we can just sort of flip through this trace and see what it's doing.

So here it's writing a test script, running the test script, making an edit in the, the code base, and so on. And so

the, the sort of process that, you know, now I can like really easily follow using this tool is to be do a new run, look for things that went wrong, especially things that, that everybody got right and now we're getting wrong, and then dig into this trace.

And really, you cannot avoid this last part. Like, I've tried a lot to, um, automate parts of this by, um... I've tried a lot to automate, to remove the need for me to actually like manually inspect all of these traces.

And I'm here to tell you, like today, that is still impossible. You basically always have to look really closely at data for anything that you're building.

**Alessio** [22:24]
And Shawn, what's the-- what would be your next step from here, right? So you're building this agent, you can see the traces, you see there's some sort of behavior that regressed. How do you go back and figure out, is it a prompt change issue?

### Debugging

**Alessio** [22:37]
Is it a tool description change issue? Like, uh, how do you debug it?

**Shawn Lewis** [22:42]
Yeah. I think so, let's see. My process, um, is there's still other like non, uh, like nice like new tools involved in this process, right? So it's really like I'll flip through these traces and kind of like, um, for each one, I'll write down notes about like what I thought, um, went wrong there.

And I'll do that for like, say, twenty or so. And then I kind of go, "Okay, what's the biggest problem that we encountered in these like twenty failures? Um, let's go try to fix that problem." So, uh, in the case where we just ran a new run and we're doing that, you know, we know what we changed from the prior run to this run, right?

So there's, there's like a lot of signal and like we edited the prompt and now we have these new failures. And so that, that can indicate a problem. I think one interesting thing about Weave that maybe other tools don't do is it's really good at tracking both your code and your kind of configuration altogether.

So the PhaseShift framework makes the best use of Weave possible. Weave is fairly low level, and it's a little bit, you know, tricky to set up this kind of tracking. So PhaseShift makes this really easy on top of Weave.

Um, but what we're looking at here is diffing like two of those evaluation runs, the configurations for those two evaluation runs. And if we, if we drill into... So the question that I might have is like, what did I change from the prior run that, that worked on a problem and this new run, um, that's failing?

And we can drill into like the exact configuration of the PhaseShift code here to see what changed. So everything with a yellow highlight here on the left is something that changed. If we go into this environment, I can see that I like configured the environment, which has all my tool call code a little bit differently here.

So, you know, there's a couple configuration tweaks. But then also in this prompt,

down in the side of the format function. So now here we're looking at a function, not like a fixed string, because these prompts are complex and have, you know, all this concatenation and stuff going on. We can see exactly what changes I made, what the delta was from one to the next.

So, you know, the process is like we know that we made this change here, where we said take the best step instead of the next assistant step to complete the task. And like, how does that, how does the LLM think about that and interpret it differently, and why would it result in these like different outcomes?

### PhaseShift Release

**Alessio** [25:04]
And what's the plan with Phase Shift? Because I think in the blog you mentioned you wanted to polish it up and then release it. Um, do you have a timeline or any, any other stuff outstanding that you really want to build before you release it?

**Shawn Lewis** [25:17]
Yeah. I would love to get it out there. I think, you know, I, I built this in the course of like really trying to solve that SWE-Bench or do really well in SWE-Bench first and foremost. So it's kind of like tuned to Shawn today.

So it needs some like good, like cleanup and, you know, all the stuff you would do to really open source something and make it great. But I do find it, you know, incredibly useful myself and, and it was able to get this great result.

So I'd like love to share it with the world. I think, you know, it lets me kind of compose together any ideas I have really quickly now and then test them out. So yeah, I guess I can't give an exact timeline yet, but it seems like that's next on my list, especially because people have been asking me a lot about that.

So it's nice to see the, the interest.

**Alessio** [25:59]
Can, can you use your agent to polish it up? Or I, I guess like that's another good question is like, you know, there's kind of this SWE-Bench verified benchmark, and then there's like, you know, the more programmatic, how am I gonna use this agent?

### Real-world Use

**Shawn Lewis** [26:11]
Totally.

**Alessio** [26:11]
Uh, what are maybe, yeah, some things that you've been using it for that have been great, or are there some use cases that don't really translate from like benchmark to, to real world?

**Shawn Lewis** [26:20]
Yeah. Great question. Um, well, I'll tell you something funny. Everything in this, uh, in the Phase Shift UI, this or this Eval Studio UI that I showed you was, was written by AI. So this entire UI was written by AI, but it was not written by-

**Alessio** [26:34]
Oh my God

**Shawn Lewis** [26:35]
... Phase Shift, it was written by Cursor. And it's not it. Because as I was doing that, it's really the interactive development process that I want of, of like looking at the UI and saying, "Oh, add this button here, move this over there, give me some charts here and there."

So Phase Shift is really designed for like longer running agentic tasks where you look away for a long time and come back and you get some result. The other thing is I, I just haven't made Phase Shift the framework actually be able to run on my local file system yet.

It's sort of an evolution of a prior thing that I made called Programmer that I've used a lot on my local file system to, to like write code, to write analysis tools, all this stuff. But Phase Shift itself in TypeScript basically can only run on SWE-Bench Verified today.

So, you know, other folks have been suggesting on Twitter, "Well, maybe, maybe fix it up to like edit itself and then see if you can clean the framework up."

**Swyx** [27:29]
Cool. Uh, you, you, there are, there are a number of other things that you mentioned you wanted to write about Crosscheck. Uh, it seems like you're gonna have a blog post about that. Uh, and then also there's like just the general submission process for SWE-Bench Verified.

### Crosscheck

**Swyx** [27:41]
They're asking for your traces. I'm sure there's nothing super proprietary there, but like, yeah, just anything else that you are going to release, uh, you know, in, in the coming days.

**Shawn Lewis** [27:51]
Yeah, totally. Well, let's, let's combine that into a single question. I think, I think SWE-Bench is this like it's an amazing benchmark. It is, like it's really like a bellwether of like can agents independently solve real world programming tests.

They're kind of small problems typically, and they're like tractable and known to be in different ways. But, but it's, it's a great benchmark, and the folks who made it, I think deserve like tons and tons of credit for giving that to the world.

Um, and one of the really important and good things that they did is they, they require trajectories to be submitted, um-

**Swyx** [28:24]
Yeah

**Shawn Lewis** [28:25]
... along with-

**Swyx** [28:25]
We talked about this with, uh, Cosine and Anthropic as well when we had their interviews with them for the same benchmark that you're, you're now on top of.

**Shawn Lewis** [28:34]
Yeah. I think it's like, I think like if they had chosen to say, "Okay, just give us a result, we'll publish it," it would, it would just be much, it would be much less useful to everybody who was like, you know, following along and, you know, harder to verify even what they're doing.

Um, so when I submitted this, I put together like all the trajectories that include the parallel rollouts and the results of the final Crosscheck step, which is a step that can, that can take different trajectories and combine them to figure out like which one was correct.

So I submitted the trajectories from all of that and a write-up in SWE-Bench Verified submissions about how that step works. So I don't think anybody realizes that that's there because unfortunately, the trajectories directory when you submit ends up getting uploaded into S3, and it's actually kind of hard to get it out.

It's in a private bucket. But if you go to my fork of SWE-Bench Verified where I submitted this of the-

**Swyx** [29:31]
Yeah. Trash

**Shawn Lewis** [29:31]
... experiments repo.

**Swyx** [29:32]
Yeah. A lot of people do this.

**Shawn Lewis** [29:34]
You can see the-

**Swyx** [29:34]
Yeah

**Shawn Lewis** [29:34]
... the full submission. What's that?

**Swyx** [29:37]
Yeah. A lot of people do this with like the fork and then, you know, you have the, uh, the JSON files there.

**Shawn Lewis** [29:41]
Yeah. So like if you want like alpha about how to make AI agents, like go figure out the crazy process to get your, uh, AWS credentials and download all these trajectories and what people have written. I think Google has a write-up of how their submission works, like buried in this directory too.

So I wrote up how, uh, Crosscheck works here, and you can go read this, you know, and, and like figure it out. I think I can talk about this if you want to. It's, uh, I thought it was interesting when I made it.

**Swyx** [30:14]
Yeah. Well, I mean, well, you know, one, one thing on the Crosscheck I think, um, uh, is, is something that people are interested in. I think the other agent builders, I won't name names, who are sort of, you know, although slightly salty that you came out ahead of all of them, are, are kind of flagging this as like their sort of like, you know, you basically sent a, like a modified rejection sampling.

Uh, I think you have a different view. I think you... The, the way that you explained it is kind of like choosing between trajectories. Uh, I think that's a good, uh, that's an interesting algorithm that no one has, uh, articulated before.

So maybe a, a brief take on that.

**Shawn Lewis** [30:49]
Sorry, so can you, can you explain the saltiness? Like, what is, what's the-

**Swyx** [30:53]
There, okay. Again, I, I won't, I won't say who this is from, but I did talk to one of the other people who, you know, is in the arena on, on SWE-Bench Verified, and they were like, "Oh, what do you think about, you know, this Weights & Biases one?"

And he was like, "Yeah, it's great, but like, you know, they, they just ran o1 five times and then chose the best one." I'm like, "I could do that too. It's just, like, super expensive." But, you know, that's the dismissive take.

**Shawn Lewis** [31:12]
Well, yeah, I think, um, I think all the top submissions, so I can't... I don't know about, um, the, uh... There's one called, like, Black Box that was right below my submission. I don't know what they did, um, because it's Black Box, I guess.

The one below that from, um, Ade from Code Story, Ade did the same thing. Like, they, they run multiple trajectories. Then they, you know, have a strategy for choosing the best. The Google submission down below I think ran thousands of trajectories and then has a strategy for choosing the best.

You know, y- if you wanna get to the top of the leaderboard at this point, you probably need a strategy like that somewhere and be willing to, like, pay the cost, whatever it is. But it's also, like, it's really important.

Like, to be able to move up 6% by doing something like that, like the problems, remember, get exponentially harder as you go up, um, the leaderboard. So if you can solve 6% more problems on SWE-Bench, you're, you're doing much better than you were at, like, you know, 60%.

Um, so, so basically everybody's doing that. And there's a submission I think that's on, that will top the leaderboard whenever they accept it that has done that even more extremely than mine, so.

**Swyx** [32:12]
Yeah, Isoform. Yeah.

**Shawn Lewis** [32:13]
Yep. Yeah, exactly.

**Swyx** [32:16]
Always, always fun games.

**Shawn Lewis** [32:17]
I think the strategy that we have here for choosing them is different than what other, other people have done. You know, like this, I think this works really well and there's a lot of potential using this kind of, kind of way that I did it to, like, get much better results.

**Swyx** [32:31]
Yeah. Awesome. Um, you know, I, I'm sure, I'm sure, you know, the roadmap for 2025 is still, uh, very up in the air. But if we are gonna see more agents work from Weights & Biases, I think the world will really welcome it.

### Future Outlook

**Shawn Lewis** [32:45]
Yeah. We'd love to. We'd love to keep going down this path.

**Alessio** [32:48]
Awesome. Any parting thoughts, Shawn? Anything we forgot? Any-

**Swyx** [32:52]
Yeah. Call to action.

**Shawn Lewis** [32:54]
Call to action. Let's see. Parting thoughts. You know, I don't know where all this is going. I think, I think that we will have autonomous AI programmers working really well for us, you know, within the next year or two.

I think-

**Swyx** [33:09]
Have you guys tried Devin or Devin Lite?

**Shawn Lewis** [33:12]
Yeah, yeah.

**Swyx** [33:13]
For-

**Shawn Lewis** [33:13]
We try to use everything internally.

**Swyx** [33:15]
Yeah.

**Shawn Lewis** [33:16]
I, uh, I've used Devin. I think it's really cool. I think the UI is incredible, like all the work that they've done around that. I think that may be where most of the, like... If there's any mode here, it's gonna be around, like, the interfaces into humans and their businesses.

Um, and Devin, like, has a major, major lead on, on making that work really well. And it's like, you know, there's a, there's a post floating around right now that says, "Oh, this didn't, like, get us the results that we want."

Well, just wait like another six months until they plop the next great models in there, and you'll get good results and have a great UI on top of it. So, you know, I think spending time focusing on UI tools, um, you know, the business aspects of things is like the way to kind of succeed and win here.

**Swyx** [33:57]
Yeah. Awesome. All right. Well, thank you so much. That was a really quick run-through. Um, thank you for coming on even though you're on paternity leave. Uh, but also I'm sure your mind is still buzzing with all, all these ideas on, you know, what to do and how to, how to compete.

### Outro

**Swyx** [34:10]
I mean, it's a very competitive arena, and, uh, congrats for coming out on top.

**Shawn Lewis** [34:15]
Yeah. Thanks so much. Yeah, the next question is, "Does R1 do as well?" So we'll find out.

**Swyx** [34:20]
Run it. I'm sure you have all the-

**Alessio** [34:22]
All the-

**Swyx** [34:22]
... all the setup you need. Yeah.

**Alessio** [34:23]
Maybe MB can bounce back if you fix it. Um, all right. Thank you so much, Shawn.

**Swyx** [34:29]
Thank you.

**Shawn Lewis** [34:30]
Hey, have a good one.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
