LALatent SpaceAug 22, 2024· 1:01:16

Is finetuning GPT4o worth it?

Ali Pullen, founder of Cosine, discusses Genie, an AI software engineering colleague that achieved 30% on SWE-Bench and 43.8% on SWE-Bench Verified via fine-tuning GPT-4o on billions of tokens of synthetic data. Cosine built Genie by training on the process of software engineering—not just working code—using synthetic runtime errors and a self-improvement loop. The agent follows a four-stage workflow: code retrieval (66% accuracy), planning, writing diffs, and running CI tests. Pullen reveals that model performance degrades linearly beyond 60K tokens in the context window, and that OpenAI dynamically sizes LoRA adapters to handle Cosine's massive dataset. He explains why Genie isn't on the SWE-Bench leaderboard (refusing to publish trajectories to prevent distillation) and outlines plans to scale the dataset and fine-tune on customer codebases for personalized performance.

  1. 0:00Intro
  2. 1:17Origin
  3. 7:39Pivot
  4. 10:59Scaling Up
  5. 16:18Data Mix
  6. 21:41Code Retrieval
  7. 31:20Planning & Code
  8. 39:56Tests & Fine-tuning
  9. 48:35Benchmarks & Future
  10. 58:04Outro

Powered by PodHood

Transcript

Intro0:00

Alessio0:04

Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host Swyx, founder of Small AI.

Swyx0:13

Hey, and today we're back in the studio in person-

Alessio0:17

We are

Swyx0:17

... uh, after about three to four months in visa jail, and travels, and all other fun stuff that we talked about in a previous episode. Uh, but today with special guest, Ali Pullen from Cosine. Welcome.

Ali Pullen0:28

Hi, thanks for having me.

Swyx0:29

We're very lucky to have you because you're on a two-day trip to San Francisco.

Ali Pullen0:33

Yeah, I wouldn't recommend that. I would not recommend it. Don't fly from London to San Francisco for two days.

Swyx0:37

And you launched Genie on a plane-

Ali Pullen0:40

Yes

Swyx0:40

... on plane Wi-Fi, um, claiming state-of-the-art in SWE-bench, which, which we're all gonna talk about.

Ali Pullen0:45

Mm-hmm.

Swyx0:46

I'm excited to dive in into, into your whole journey because it has been a journey. I've been lucky to be a small angel in, in part of that journey, and it's exciting to see that you're launching to such a, such a, a claim and, and you know, such results.

Um, so I'll, I'll go over your brief background, and then you can sort of fill in the blanks as on, on, you know, what else people should know about you. You did your bachelor's in computer science in Exeter.

Ali Pullen1:07

Yep.

Swyx1:07

And then you worked at a startup that got acquired into Gopuff.

Ali Pullen1:10

Mm-hmm.

Swyx1:10

And round about 2022, you started working on a stealth startup that became a YC startup. What, what's that overall story?

Ali Pullen1:17

Yeah, so basically when I left university, I, I met my now co-founder, Sam. At the time we were both mobile devs. He was an Android developer, I was an iOS developer. And whilst at university we built this sort of small consultancy, sort of we'd, um, be approached to build projects for people, and we would just take them up and start with our student projects.

Origin1:17

Ali Pullen1:39

They weren't, they weren't anything crazy or anything big. We started with those, and over time we started d- doing larger and larger projects, more interesting things. And, um, actually when we left university, we just kept doing that. We didn't really get jobs, traditional jobs.

It was also, like, in the middle of COVID, middle of lockdown.

Swyx1:55

Yeah.

Ali Pullen1:55

So we were like, "This is a pretty good gig. We'll just keep, like, writing code in our bedrooms." And we did that for a while, and then a friend of ours that we went to Exeter with started a YC startup during COVID, and it was one of these fast grocery delivery companies.

At the time, I was living in the deepest, darkest countryside in England where fast grocery companies are still not a thing. So he, he sort of pitched me this idea and was like, "Listen, like, I need a, an iOS dev.

Do you fancy coming along?" And I thought, "Absolutely." It was a chance to get out my parents' house, chance to move to London, you know, do interesting things. And at the time, truthfully, I had no idea what YC was.

I, I had no idea. I wasn't in the startup space. I knew I liked coding and building apps and stuff, but I'd never, never really done anything in that area. So I said, "Yes, absolutely." I moved to London just sort of as COVID was ending, and yeah, worked at what was Fancy for about a year and a half.

Then we brought Sam along as well. So we, Sam and I were the two engineers at Fancy for basically its entire life, and we built literally everything. So, like, the, the front-- the client mobile apps, the, the backends, the internal, like, stock management system, the driver routing algorithms, all those things.

Literally, like, everything. It was my first... You know, both of us were super inexperienced. We didn't have, like, proper engineering experience. There are definitely decisions we'd do differently now. We'd definitely buy a lot of stuff off the shelf, stuff like that.

But it was the initial dip of the toe into, like, the world of startups and we were both, like, hooked immediately. We were like, "This is so cool. This sounds so much better than all our friends who are, like, consultants and doing, like, normal jobs," right?

We did that, and it ran its course, and after, I wanna say, 18 months or so, Gopuff came and acquired us, and there was obviously a transitionary period, an integration period, like with all acquisitions. And we did that, and as soon as we'd vested what we wanted to vest, and as soon as we thought, "Okay, this chapter is sort of done," uh, in about 2022, we left, and we knew that we wanted to go it alone and try something.

Like, we'd had this taste. Now we knew, we'd, we'd seen how a, like a YC startup was managed, like, up close, and we knew that we wanted to do something similar ourselves. We had no idea what it was at the time.

We just knew we wanted to do something. So we, we tried a small, um, some small projects in various different areas. But then Sam talked to me about GPT-3. He'd seen it on Reddit, and I-

Swyx4:11

The source of all knowledge.

Ali Pullen4:13

The source of all knowledge. Absolutely. Yeah. Sam loves Reddit. I'd actually heard of GPT-2, and obviously had, like, loosely followed what OpenAI had done with... What was the game they trained a model to play? I can't remember.

Swyx4:23

Dota.

Alessio4:24

Dota.

Ali Pullen4:24

Was it Dota? Yeah. So I'd, I'd followed that, and I knew loosely what GPT-2 was. I knew what BERT was. So I was like, "Okay, this GPT-3 thing sounds interesting." And he just mentioned it to me on a walk, and I then went home and, like, Googled GPT-3, and there was the playground.

It was the-- And the model was DaVinci 2 at the time, and it was just the, the old school playground completions, nothing crazy, no chat, no nothing. Uh-

Swyx4:47

I miss completion still.

Ali Pullen4:48

Yeah. Oh, completion. Honestly, I had this conversation in OpenAI-

Swyx4:51

Yeah

Ali Pullen4:51

... obviously yesterday. I was like, I just wish-

Alessio4:52

Completion

Ali Pullen4:52

... I know. But yeah, so we, we, um, I, I started playing around with the, the playground, and the first thing I ever wrote into it was like, "Hello, world," and it gave me some sort of, like, fairly generic response back, and I was like, "Okay, that looks pretty cool."

The next thing was I looked through the docs, um, or the s- they had a lot of example prompts, 'cause I had no idea. I didn't know if the, if you could put anything in. I didn't know if you had to structure it in a certain way or whatever, and I, and I saw that it could start writing, like, tables and JSON and stuff like that.

So I was like, "Okay, can you write me something in JSON?" And it did. And I was like, "Oh wow, this is, this is pretty cool. Um, can it, can it g- can you just write arbitrary JSON for me?"

And, um, immediately, as soon as I realized that, my mind was racing and I, like, got Sam in, and we just started messing around in the playground, like, fairly innocently to start with. And then of course, both being mobile devs and also seeing, at that point we'd learnt about what the Codex model was.

It was like this thing's trained to write code. Sounds awesome. And Copilot was start- I think, I d- I can't actually remember if Copilot had come out yet. Yeah, it, it might have done.

Swyx5:54

It's round about the same time as Codex.

Ali Pullen5:55

Round about the same time, yeah. And we were like, "Okay, as mobile devs, let's see what we can do." So the initial thing was like, "Okay, let's see if we can get this AI to build us a mobile app from scratch."

We eventually built the world's most flimsy system, which was back in the day were like 4,000 token context windows, like chaining prompts, trying to keep as much context from one to the other, all these different things where- Essentially, you'd put in an app idea in a box, and then we'd do, like, very high-level stuff, figuring out what the stack should be, figuring out, um, what the front end should be written and back end should be written, and all these different things.

And then we'd go through, like, for each thing, more and more levels of detail until the point that you actually got Codex to write the code for each thing. And we didn't do any templating or anything. We were like, "No, we're gonna write all the code from scratch every time," which is basically why it barely worked.

But there were, like, occasions where you could put in something and it would build something that did actually run, the back end would run, the database would work, and we were like, "Oh my God, this is insane. This is so cool."

And that's what we showed to our co-founder Yang. I met my co-founder Yang through, through Fancy 'cause his wife was their first employee. And, um, we showed him, and he was like, "You've discovered fire. What is this? Like, this is insane."

He has a lot more startup experience historically. He's had a few exits in the past and has been through all different industries. He's like our dad. He's a bit older. He hates me saying that, but he's a bit older.

He's your COO now? He's our COO, yeah. And, uh, we showed him, and he was like, "This is absolutely amazing. Let's just do something." 'Cause he w- he, at the time, um, was just about to have a child, so he didn't have anything going on either.

So we, we applied to YC, got an interview. The interview was, as most YC interviews are, short, curt, and pretty brutal. They told us they hated the idea, and they didn't think it would work, and that's when we started brainstorming.

It was l- almost like the interview was like an office hours kind of thing. And we were like, "Okay, given what you know about the space now and how to build things with, with these LLMs, like, what can you bring out of what you've learnt in building that thing into something that might be a bit more useful to people on the daily?"

Pivot7:39

Ali Pullen7:53

And also YC obviously likes B2B startups a little bit more, at least at the time they did back then. So we were like, "Okay, maybe we could build something that helps you with existing code bases, like can sort of automate development stuff with existing code bases," not knowing at all what that would look like or how you would build it or any of these things.

And they were like, "Yeah, that sounds interesting. You should probably go ahead and do that. You're in. You've got two weeks to build us an MVP." And we were like, "Okay. Okay." We did our best. The MVP was absolutely horrendous.

It was a CLI tool. It sucked. And, um, at the time we were like, we, we don't even know how to build what we want to build, and we didn't really know what we wanted to build, to be honest.

Like, we knew we wanted to try to help automate dev work, but back then we just didn't know enough about how LLM apps were built, the intricacies and all those things. And also, like, the LLMs themselves, like 4,000 tokens, you're not going very far.

They're extremely expensive. So we ended up building a, uh, a code-based retrieval tool originally. Our thought process originally was we want to build something that can do our jobs for us. That is, like, the gold star. We know that.

We've seen, like, there are glimpses of it happening with our initial demo that we did, but we don't see the path of how to do that at the moment. Like, the tech just wasn't there. So we were like, "Well, there are gonna be some things that you need to build this when the tech does catch up," so retrieval being one of the most important things, like the model's gonna have to be able to, like, pull code out of a code base somehow.

So we were like, "Well, let's just build the tooling around it, and eventually when the tech comes, then we'll be able to just, like, plug it into our, our tooling, and then it should work basically." And to be fair, that's basically what we've done, and that's basically what's happened, which is very fortunate.

But in the meantime, whilst we were waiting for everything to sort of become available, we built this code-based retrieval tool. That was the first thing we ever launched when we were in YC, like, and it didn't work. It was really frustrating for us 'cause it was just me and Sam, like, working, like, all hours trying to get this thing to work.

It was quite a big task in and of itself trying to get, like, a good semantic search engine working that could run locally on your machine. We were trying to avoid sending code to the cloud as much as possible.

And then for very large code bases, you're like, you know, millions of lines of code. You're trying to do some sort of, like, local HNSW thing that runs inside your VS Code instance that, like, eats all your RAM as you've seen in the past, all those different things.

Yep. Yeah. My first call with you, I think I had trouble- You were like, "Yeah, Ali, it sucks, man." Yeah. I was like, "Yeah, I know, I know. I know it sucks. I'm sorry." Um, but building all that stuff was essentially the first six to eight months of, of what at the time was Bilt.

Which by the way, Bilt- Bilt. Yeah, it was a terrible, terrible name. That was the worst le- Yes ... um, part of trying to think about whether I would invest is whether or not people could pronounce it. Pronounce the name.

No, in, in, when we, so when we went on our first ever YC, like, retreat- ... no one got the name right. They were like, "Bilt, Bilt, Bilt, what?" Um, and then, yeah, we actually changed the name to Cosine, like, although some people would spell it co, as in like, as if you're cosigning for an apartment or something.

Like, that's like, can't win. Yeah, that was what Bilt was back then. But the ambition, and I, and I did a talk on this back in the end of 2022, the ambition to, like, build something that essentially automated our jobs was still very much, like, core to what we were doing.

But for a very long time it was just never apparent to us, like, how would you go about doing these things? Even when, like, you had 3.5 16K, 16K suddenly felt huge 'cause you've gone from 4 to 16, but even then 16K is like, a lot of Python files are longer than 16K.

Scaling Up10:59

Ali Pullen11:13

So you can't, you know, b- before you even start doing a completion. Even then we were like, "Eh, yeah, it looks like we're still waiting." And then, like, towards the end of last year, you then start, you see 32K.

32K was really smart. It was really expensive, but also, like, you could fit a decent amount of stuff in it. 32K felt enormous. And then finally 128K came along, and we were like, "Right, this is like, this is what we can actually deal with."

Because fundamentally, to build a product like this, you need to get as much information in front of the model as possible and make sure that everything it ever writes in output can be traced back to something in the context window so it's not hallucinating it.

As soon as that model existed, I was like, "Okay, I know that this, this is now gonna be feasible in some way." We'd done early sort of dev work on Genie using 3.5 16K, and that was a very, very, like, crude way of proving that this loop that we were after and the way we were generating the data actually had signal and worked and, and could do something.

But the model itself was not useful because you couldn't ever fit enough information into it for it to be able to do the task competently, and also the base intelligence of the model. I mean, 3.5, anyone who's used 3.5 knows the base intelligence of the model is Is lacking, especially when you're asking it to, like, do software engineering.

This is quite, quite involved. So we saw the 128K context model, and, um, uh, at that point, we'd been in touch with OpenAI about our ambitions and, like, how we wanted to build it. We essentially, uh, I just took a punt and I was like, "I'm just gonna ask to see can we, like, train this thing?"

'Cause at the time, 4 Turbo had just come out, and back then there was still a decent amount of lag time between, like, OpenAI releasing a model and then allowing you to fine-tune it in some way. They've got much better about that recently.

Like, 4o fine-tuning came out either, I think, a day... 4o Mini fine-tuning came out, like, a day after the model did, and I know that's something they're definitely, like, optimizing for super heavily inside, which is great to see.

Swyx13:08

Which is a little bit, you know, what, for a year or so, YC companies had, like, a direct Slack channel to OpenAI.

Ali Pullen13:14

We still do.

Swyx13:15

Yeah.

Ali Pullen13:15

Yeah. I mean-

Swyx13:16

So it's a little bit of a diminishing of the YC advantage there-

Ali Pullen13:19

Yeah

Swyx13:19

... if, if they're releasing this fine-tuning ability, like, a day after.

Ali Pullen13:22

Yeah, no, no, absolutely. But, like- ... you can't build a startup on the YC advantage. It's obviously nice, and it makes you feel warm-

Swyx13:27

Yeah, yeah.

Ali Pullen13:27

... and fuzzy inside, but, like, at the end of the day, it, it, it, it's not that that's gonna make you win.

Swyx13:31

Yeah.

Ali Pullen13:31

But yeah, no, so, like, we'd, we'd spoken to Shamul, their, their DevRel guy. I'm sure you, you know him. Um-

Swyx13:36

I think he's solution- head of solutions or something.

Ali Pullen13:38

He is in their applied team. Yeah. We'd been talking to him from the very beginning going to YC, and he's been absolutely fantastic throughout. I basically had pitched him this idea back when we were doing it on 3.5 16K, and I was like, "This is my, this is my crazy thesis.

I wanna see if this can work." And as soon as, like, that 128K model came out, I was... I started, like, laying the groundwork. I was like, "I know this definitely isn't possible 'cause you released it, like, yesterday, but know that I want it."

And in the interim, like, GPT-4, like, 8K fine-tuning came out. We tried that. It's obviously even fewer tokens, but the intelligence helped. And I was like, "If we can marry the intelligence and the context window length, then we're gonna have something special."

And eventually we were able to get on the experimental access program, and we got access to 4 Turbo fine-tuning. As soon as we did that, because in, in the entire ramp to that, we'd built the data pipeline. We already had all that set up, so we're like, "Right, we have the data.

Now we have the model. Let's put it through and, and iterate," essentially. And that's, that's where, like, Genie as we know it today really was born. I won't pretend like the first version of Genie that we trained was good.

It was a disaster. That's where you realize all the implicit biases in your data set, and you realize that, oh, actually, this decision you made that was fairly arbitrary was the wrong one. You have to do it a different way.

Other subtle things like, you know, how you write Git diffs and, you know, using LLMs and how you can best optimize that to make sure they actually apply and work, and loads of different little edge cases. But as soon as we had access to the underlying tool, we were like, "Right, we can actually do this."

And I was... I breathed a sigh of relief 'cause I di- I didn't know it was, like, it wasn't a done deal, but I knew that we could build something useful, and we, I knew that we could build something that, um, would be measurably good on whatever eval at the time that you wanted to use.

Like, at the time, back then, we weren't actually that familiar with... But once Devin came out and they announced SWE-bench score, I, like, that's when my life took a turn. And, uh-

Swyx15:30

Challenge accepted.

Ali Pullen15:31

That's, yeah, challenge accepted. And that's where, like, yes, that's where my, my friendships have gone, my sleep has gone- ... like, my, my weight, everything gone into SWE-bench. And yeah, we, we, it, it was actually a very useful tool in building Genie, 'cause beforehand it was like, "Let's vibe check this thing and see if it's useful."

And then all of a sudden you have a, an actual measure to, to see, like, can it do software engineering? Not, not the best measure obviously, but, like, it's, uh, it's the best that we've got now. We, we just iterated and, and built, and eventually we got it to the point that where it is now and a little bit beyond since we actually, like, did...

We, we actually got that score a couple of weeks ago. And yeah, it's been a hell of a journey from the beginning all the way now. That was a very rambling answer to your question-

Swyx16:10

Great

Ali Pullen16:10

... about how we got here, but that's essentially the potted, uh-

Swyx16:12

Yeah

Ali Pullen16:12

... yeah, answer of how we got here.

Swyx16:13

Got the full origin story out.

Alessio16:14

Yeah, no, totally. You mentioned bias in the data and some of these things.

Data Mix16:18

Swyx16:18

Mm.

Alessio16:18

In your announcement video, you called Genie the world's first AI software engineering colleague, and you kinda highlighted how the data needed to train it needs to show how a human engineer works. I think maybe you're contrasting that to just putting code in it.

There's kinda, like, a lot more than code that goes into-

Ali Pullen16:36

Yes. Yeah, absolutely

Alessio16:36

... software engineering.

Ali Pullen16:37

That's correct. Yeah.

Alessio16:38

How do you think about the data mixture, you know? And, like, uh, there's this kinda known truth that code makes models better-

Ali Pullen16:45

Correct

Alessio16:45

... when you put in the pre-training data.

Ali Pullen16:46

Mm-hmm.

Alessio16:47

But since we put so much in the pre-training data, what else do you add when you train the Genie in?

Ali Pullen16:51

Yeah, I think, well, there's, the, I think that sort of boils down fundamentally to the difference between a model writing code and a model doing software engineering. Because the, the, the software engineering sort of discipline goes wider. Because if you look at something like a PR, that is obviously a artifact of some thought and some work that has happened and has eventually been squashed into, you know, some diffs, right?

What the, very crudely, what the pre-trained models are, are reading is they're reading those final diffs, and they're emulating that, and they're being able to output it, right? But of course, it's a super lossy thing, a PR. You have no idea why or how, for the most part, unless there are some comments, which, you know, anyone who's worked in a company realizes PR reviews can be a bit dodgy at times.

But you see that you lose so much information at the end and, and that's perfectly fine because PRs aren't designed to be something that perfectly preserves everything that happened. But what we realized was if you want something that's a software engineer, and very crudely we started with, like, something that can do PRs for you, essentially, you need to be able to figure out why those things happened.

Otherwise, you're just gonna rely... You essentially just have a code writing model. You have something that's good at human eval, but, but not very good at SWE-bench, essentially. That realization was, was part of the, the kernel of the idea of, of, of the approach that we took to design the agent that, that is Genie.

The way that we decided we want to try to extract what happened in the past, like, as forensically as possible- Has been and is currently like one of the, the main things that we focus all our time on.

Because doing that as we're getting as much signal out as possible and doing that as well as possible is the biggest thing that, that we've seen that determines how well we do on that benchmark at the end of the day.

Once you've sorted things out, like h- like output structure, how c- how to get it consistently writing diffs, and all the stuff that is sort of ancillary to the model actually figuring out how to solve a problem, the core bit of solving the problem is how did the human solve this problem, and how can we best come up with how the human solved these problems?

So all the effort went in on that, on that pipeline, and the mix that we ended up with was, as you've probably seen in the technical report and so on, all of those different languages and different combinations of different task types, all of that has run through that pipeline, and we've extracted all that information out.

Alessio19:06

How does that differ when you work with customers that have private workflows? Like, do you think-- Is there usually a big delta between what you get in open source and maybe public data-

Ali Pullen19:15

Oh, enormous

Alessio19:16

... versus like-

Ali Pullen19:16

Yeah, yeah, yeah. When you scrape enough of it, most of open source is updating readmes and docs. It's hilarious. Like, we had to filter out so much of that stuff because when we first did the 16-- uh, 3.5 16K model, like the amount of readme updating that went in, we did like no data cleaning, no real, like, we just sort of threw it in and saw what happened.

And it was just like re- it was really good at updating readmes, really good at r- writing some comments, really good at, um, complaining in Git reviews, in, in PR reviews rather. And it would, again, like we didn't clean the data, so you'd like give it some feedback, and it would just like reply and like it would just be quite insubordinate when it was getting back to you.

Like, "No, I don't think you're right," and it would just sort of argue with you. So the process of, of doing all that was super interesting 'cause we realized from the beginning, okay, there's a huge amount of work that needs to go into like cleaning this, getting it aligned with what we want the model to do to be able to get the model to be useful in some way.

Alessio20:09

I'm curious, like, how do you think about the customer willingness to share a lot of this historical data? I've done a lot of developer tools investing-

Ali Pullen20:17

Mm

Alessio20:17

... in my career, and, uh, getting access to the code base is always one of the hard things. Are people getting more cautious about sharing this information? In the past, it was maybe like, you know, you're using static analysis tool, like whatever else-

Ali Pullen20:30

Mm

Alessio20:30

... you need to plug into the code base, fine. Now you're building a model based on it. Like, uh, what's the discussion going into these companies? Are most people comfortable with like letting you see how they work and sharing everything, or?

Ali Pullen20:41

It depends on the sector mostly. We've actually seen, I'd say, people becoming more amenable to the idea over time actually rather than more skeptical 'cause I think they can see the, the upside. If this thing does what they say it does, it's gonna be more help to us than it is a risk to our InfoSec.

Um, and of course, like companies building in this space, we're all gonna end up, you know, complying with the same rules, and there are gonna be new rules that come out to make sure that we're looking at your code, that everything is safe and so on.

So from what we've seen so far, we've spoken to some very large companies that you've definitely heard of, and all of them obviously have stipulations, and many of them want it to be sandboxed to start with and all the like very obvious things that I, you know, I, I would say as well.

But they're all super keen to have a go and see because like despite all those things, if we can genuinely make them go faster, allow them to build more in a given time period and stuff, it's, it's super worth it to them.

Swyx21:34

Okay, I'm gonna dive in a little bit on the process that you have created. You showed the demo on, on your video.

Ali Pullen21:41

Yes.

Swyx21:41

And by the time that we release this, you should be taking people off the waitlist and-

Code Retrieval21:41

Ali Pullen21:44

Right. Yes

Swyx21:45

... sort of, uh, launching people so people can see this themselves. There's four main parts of the workflow, which is finding files, planning action, writing code, and running tests. And controversially, you have set yourself apart from the Devins of the world by saying that things like having access to a browser is not that important for you.

Is that an accurate reading of what you wrote?

Ali Pullen22:07

I don't remember saying that, but at least with what we've seen, the browser is helpful, but it's not as helpful as like writing the correct files.

Swyx22:16

Mm-hmm.

Ali Pullen22:16

If that, if that makes sense. Like, it is still helpful, but obviously there are, there are more fundamental things you have to get right before you get to like, oh, yeah, you can read some docs, or you can read a Stack Overflow article and stuff like that.

Swyx22:27

Yeah. The phrase I was indexing on was the other software tools are wrappers around foundational models with a few additional tools such as a web browser or code interpreter.

Ali Pullen22:35

Oh, I see. No, I mean, no, I'm, I'm not, I'm not, I'm not der- I'm deriding the, the, the approach there, not the, not the tools.

Swyx22:40

Yeah, exactly. So like I would say in my standard model of what-

Ali Pullen22:44

Mm-hmm

Swyx22:44

... a code agent should look like, uh, Devin has been very influential obviously.

Ali Pullen22:47

Yeah, yeah.

Swyx22:48

Because you could just @ the docs of something-

Ali Pullen22:51

Mm-hmm

Swyx22:51

... and like, you know, now I have-- now when I'm installing a new library, I can just @ docs.

Ali Pullen22:55

Yeah.

Swyx22:55

Cursor also does this, right? And then obviously having a code interpreter does help.

Ali Pullen22:59

Yep.

Swyx22:59

I guess you have that in the form of running tests.

Ali Pullen23:02

I mean, uh, the Genie has both of those tools available to it as well. So-

Swyx23:05

Oh, okay.

Ali Pullen23:05

Yeah, yeah. So we have a tool where you can like put in URLs, and it will just read the URLs, and you can... It, it also uses Perplexity's API under the hood as well to be able to actually ask questions if it wants to.

Swyx23:15

Okay.

Ali Pullen23:15

So no, we use both of those tools as well. Like those tools are super important and super key. I, I think obviously the most important tools to these agents are like being able to retrieve code from a code base, being able to read Stack Overflow articles and, and what have you, and just be able to essentially be able to Google like we do is definitely super useful.

Swyx23:35

Yeah. I thought maybe we could just kind of dive into each of those actions. Retr- code retrieval, one of the core problems, you had an indexer that-

Ali Pullen23:42

Yes

Swyx23:42

... you've worked on, uh, ev- even as, as built. What makes it hard? What approach you thought would work, didn't work?

Ali Pullen23:49

Yeah.

Swyx23:49

Anything like that.

Ali Pullen23:50

It's funny, I had a similar conversation to this when I was chatting to the guys from OpenAI yesterday. The thing is that searching for code, specifically semantically, at least to start with, I mean, like keyword search and stuff like that is a, is a solved problem.

It's been around for ages. But at least being able to... Uh, the phrase we always used back in the day was searching for what code does rather than what code is. Like searching for functionality is really hard. Really hard.

The way that we approached that problem was that obviously like a, a very basic and easy approach is, right, let's just embed the code base. We'll chunk it up in some arbitrary way, maybe using an AST, maybe using number of lines, maybe using whatever, like some overlapping.

Just chunk it up and embed it. And once you've done that, I will write a query saying, like, find me some authentication code or something, embed it, and then do the cosine similarity and get the top K, right?

That doesn't work, and I wish it did work, don't get me wrong. It doesn't work well at all because fundamentally, if you think about, like, semantically how code looks is very different to how English looks, and there's, like, not a huge amount of, of signal that's carried between the two.

So what we ended up... The, the first approach we, we took and, and that, and that kind of did well enough for a long time was, okay, let's train a model to be able to take in English code queries and then produce a hypothetical code snippet that might look like the answer, embed that, and then do the cosine similarity.

And that process, although very simple, gets you so much more performance out of the retrieval accuracy, and that was kind of like the start of our, of our engine, as we called it, which is essentially, like, the aggregation of all these different heuristics like semantic, keyword, LSP, and so on to...

And, and then we essentially had, like, a, a model that would, given an input, choose which ones it thought were most appropriate given the type of request you had. So the whole code search thing was a really hard problem, and actually what we ended up doing with Genie is we, um, let the model through self-play figure out how to retrieve code.

So actually we don't use our engine for Genie. So instead of, like, a request coming in and then, like, say, GPT-4 with some JSON output being like, "Well, I think here we should use a keyword with these inputs, and then we should use semantic, and then we should, like, pick these results," it's actually, like, a question comes in and Genie has self-played in its training data to be able to be like, "Okay, this is how I'm going to approach finding this information."

Much more akin to how a devel- a developer would do it, 'cause if I was like, "Sean, go into this new code base you've never seen before and find me the code that does this," you're gonna probably... You might do some keywords.

You're gonna look over the file system. You're gonna try to figure out from the directories and the file names where it might be. You're gonna, like, jump in one, and then once you're in there, you're probably gonna be doing the, you know, go to definition stuff to, like, jump from file to file and try to use the graph to, like, get closer and closer, and that is exactly what Genie does.

Starts on the file system, looks at the file system, picks some candidate files. Is this what I'm looking for, yes or no? If there's something that's interesting like an import or something, it can, it can command click on that thing, go to definition, go to references, and so on, and it can traverse the code base that way.

Swyx26:59

Are you using the VS Code, uh, LSP or-

Ali Pullen27:02

No. That's... No. We're not allowing... We're not doing this in VS Code. We're just using the language servers running.

Swyx27:06

Okay.

Ali Pullen27:07

But we really wanted to try to mimic the way we do it as best as possible, and we did that during the self-play process when we were generating the dataset. So although we did all that work originally, and, and although, like, it...

Genie still has access to these tools, so it can do keyword searches, and it can do, you know, basic semantic searches, and it can use the graph. It uses them through this process and, and figures out, okay, I've learnt from data how to find stuff in code bases, and I think in our technical report, I can't remember the exact number, but I think it was around 65 or 66% retrieval accuracy overall measured on we know what lines we need for these tasks to find for the task to actually be able to be completed, and we found about 66% of all those lines, which is one of the biggest areas of free performance that we can get a hold of.

Because when we were building Genie, truthfully, like, a lot more focus went on assuming you've found the right information, you've been able to reproduce the issue. Assuming that's true, how do you then go about solving it? And the bulk of the work we did was on the solving.

But when you go higher up the funnel, obviously, like, the funnel looks like, have you found everything you need for the task? Are you able to reproduce the problem that's seen in the issue? Are you then able to solve it?

And the funnel gets narrower as you go down. And at the top of the funnel, of course, is rank. So I'm actually quite happy with that score. I think it's still pretty impressive considering the size of some of the code bases we're doing, we're using for this.

But as soon as that n- If that number becomes 80, think how many more tasks we'd get right. That's one of the key areas we're gonna focus on when we continue working on Genie.

Swyx28:35

Be interesting to break out a benchmark just for that, um-

Ali Pullen28:38

Yeah. I mean, it's super easy

Swyx28:38

... just to, just to try to-

Ali Pullen28:39

Like-

Swyx28:39

'Cause I don't know what state-of-the-art is.

Ali Pullen28:41

Yeah. I mean, like, for a, um... It, it's super easy 'cause, like, for a given PR, you know what lines were edited.

Swyx28:46

Oh, okay.

Ali Pullen28:47

Yeah. You know what lines were edited.

Swyx28:47

So you can just, you can source it from SWE-Bench, actually.

Ali Pullen28:49

Yeah. You can do it, you can do it with SWE-Bench super easily.

Swyx28:51

But there might be-

Ali Pullen28:51

And that's how we got that figure out at the other end. Um, for us, being able to see it against, um, our historic models was super useful, so we could see if we were, you know, actually helping ourselves or not.

Swyx29:00

Yeah.

Ali Pullen29:00

And initially, one of the biggest performance gains that we saw when we were work- when we did work on the rag a bit was giving it the ability to use the LSP to, like, go to definition, and really try to get it to, uh, emulate how we do that.

Because I'm sure when you go into an editor without... where, where, like, the LSP's not working or whatever, you suddenly feel really, like, disarmed and naked. You're like, "Oh my God. I didn't realize how much I actually use this to get about rather than just find stuff."

So we really tried to get it to do that, and that gave us a, a big jump in performance. So we, we went from, like, 54% up to, like, the 60s, but just by adding, focusing on that.

Swyx29:31

Just one weird trick.

Ali Pullen29:32

Yes.

Swyx29:33

Um, I'll, I'll briefly comment here. So this is the standard approach I would say most, uh, code tooling startups are pursuing.

Ali Pullen29:40

Mm-hmm.

Swyx29:41

The one company that's not doing this is Magic.dev.

Ali Pullen29:44

Yes.

Swyx29:44

So would you do things differently if you have a 10 million token context window?

Ali Pullen29:49

If I had a 10 million context window and hundreds of millions of dollars, I wouldn't have gone and built, uh... It, it's an LTM. It's not a transformer, right, that they're using-

Swyx29:59

Right

Ali Pullen30:00

... if I'm not mistaken. I believe it's not a transformer.

Swyx30:02

Yeah. Eric's gonna come on at some point anyway.

Ali Pullen30:04

I'm just... Listen, I, l- They obviously know a lot more about their product than I do. I don't know a great deal about how Magic works.

Swyx30:08

Nobody knows anything yet.

Ali Pullen30:09

Yeah. I don't know. I don't... I'm not gonna, I'm... So I'm not gonna, I'm not gonna speculate. Would I do it the same way as them? I like the way we've done it because fundamentally, like, we focus on the- ...

act of software engineering and what that looks like and showing models how to do that. Fundamentally, the underlying model that we use is kind of null to us. Like, so long as it's the best one, I don't mind.

And the context windows, we've already seen how you can get transformers to have, like, million, m- m- one and a half million token context windows, and that works perfectly well. So, like, as soon as you can fine-tune Gemini 1.5, then you best be sure that Genie will ha- will work, will run on Gemini 1.5 and, like, will probably get very good performance out of that.

I like our approach 'cause we can be super agile and be like, "Oh, well, Anthropic have just released whatever," uh, you know, and it might have half a million tokens and it might be really smart, and I can just immediately take my JSONL file and just dump it in there, and suddenly Genie works on there and it can do all the, the new things.

Swyx31:04

Does, uh, Anthropic have the same fine-tuning support as OpenAI?

Ali Pullen31:08

Um-

Swyx31:08

I actually haven't heard any-

Ali Pullen31:09

They are-

Swyx31:09

... anyone do it

Ali Pullen31:09

... working on it. They are partnered, they are partnered with AWS, and it's gonna be in Bedrock.

Swyx31:13

Okay.

Ali Pullen31:13

As far as I, as far as I know. I think I'm, I th- I think, I think that's true. Um, yeah.

Swyx31:18

Cool. We have to keep moving on to, uh, the other segments.

Ali Pullen31:19

Sure.

Swyx31:20

Uh, planning, the second piece of your four-step grandmaster plan. That is the frontier right now. You know, a lot of people are talking about Strawberry, Q*, whatever that is.

Planning & Code31:20

Ali Pullen31:29

Yeah.

Swyx31:29

Monte Carlo tree search. Is current state-of-the-art planning good enough? What prompts have worked? I don't even know what questions to ask. Like, what is the state of planning?

Ali Pullen31:38

I think it's fairly obvious that with the foundational models, like, you can ask them to think by step by step and ask them to plan and stuff, but that isn't enough because if you look at how those models score on these benchmarks, and then they're not, they're not even close to state-of-the-art.

Swyx31:49

Which ones are you referencing? Benchmarks?

Ali Pullen31:50

So, like, just, uh, like, SWE-Bench and-

Swyx31:52

Okay, yeah

Ali Pullen31:52

... and so on, right? And, like, even the things that get really good scores on HumanEval are agents as well 'cause they have these loops, right? Obviously, these things can reason, quote unquote, but the reasoning is the model-- Like, it's constrained by the model's intelligence, I'd say, very crudely.

And what we essentially wanted to do was we still thought, like, obviously reasoning is super important. We need it to get the performance we have. But we wanted the reasoning to emulate how we think about problems when we're solving them as opposed to how a model thinks about a problem when we're solving it, and that was, that's obviously part of, like, the derivation pipeline that we have when we, when we, when we design our data.

But the reasoning that the models do right now and, and who knows what Q*, whatever it ends up being called- ... looks like. But certainly what I'm excited-- On a, on a small tangent to that, like, what I'm really excited about is when models like that come out, obviously the signal in my data when I regenerate it goes up, and then I can then train that model that's already better at reasoning with improved reasoning data and just, like, I can keep bootstrapping and keep leapfrogging every single time, and that is, like, super exciting to me 'cause I don't-- I welcome, like, new models so much because immediately it just floats me up without having to do much work, which is always nice.

But at the state of reasoning generally, I don't see it going away anytime soon. I mean, that's like an autoregressive model doesn't think per se. And in the absence of having any thought, maybe a, an energy-based model or something like that, maybe that's what Q* is, who knows?

Some sort of like high level abstract space where thought happens before tokens get produced. In the absence of that, for the moment, I think it's, it's all we have, and it's gonna have to be the way it works.

For what happens in the future, we'll have to see, but I think certainly it's never going to hinder performance to do it. And certainly the reasoning that we see Genie do when you compare it to, like, if you ask GPT-4 to break down step by step an approach for the same problem, at least just on a vibe check alone, looks far better.

Swyx33:43

Two elements that I like that I didn't see in your initial video, we'll, we'll see when, you know, this, um, Genie launches-

Ali Pullen33:50

Mm-hmm

Swyx33:50

... is a planner chat, which is I can modify the plan at-

Ali Pullen33:53

Yes

Swyx33:53

... while it's executing.

Ali Pullen33:54

Yes.

Swyx33:54

And then the other thing is playbooks, which also from Devin-

Ali Pullen33:57

Mm-hmm

Swyx33:57

... where here's how I like to do a thing, and I'll use Markdown to specify how I do it. I'm just curious if, if, like, you know, those things help.

Ali Pullen34:06

Yeah, no, absolutely. We're, we're 100%-- We want everything to be editable, not least because it's really frustrating when it's not.

Swyx34:11

Yeah.

Ali Pullen34:11

Like, if you're ever, if you're ever in a situation where you're like, "This is, this is the one thing I just wish I could..." Then you'd be right if that one thing was right, and you can't change it.

So we're gonna make everything editable, including the code it writes. Like, you can-- If it makes a small error in a patch, you can just change it yourself and let it continue, and it will be fine.

Swyx34:25

Yeah.

Ali Pullen34:26

So yeah, like, those things are super important. We'll be doing those two.

Alessio34:29

I'm curious, once you get to writing code-

Ali Pullen34:31

Mm-hmm

Alessio34:31

... is most of the job done? I feel like the models are so good at writing code when they're, like, in small chunks that are, like, very well instructed. What's kind of the drop-off in the funnel? Like, once you get to, like, you got the right files and you got the right plan.

Ali Pullen34:44

That's a great question because by the time this is out, there'll be another blo- there'll be another blog post.

Alessio34:48

Yeah.

Ali Pullen34:49

There'll be another blog post which, uh, contains all the informa- all the learnings that I delivered to OpenAI's fine-tuning team when we finally got the score.

Alessio34:56

Oh, that's okay.

Ali Pullen34:56

Um, and-

Swyx34:57

Go for it. It's already out.

Ali Pullen34:58

And, um, yeah, yeah. I d- I don't have it on my phone.

Swyx35:00

Ah, okay.

Ali Pullen35:00

But basically, I, um, broke down the log probs. I basically got the average log prob for a token at every token position in the context window. So imagine an X-axis from zero to 128K, and then the average log prob for each index in there.

As we discussed, like, the way Genie works normally is, you know, at the beginning you do your RAG, and then you do your planning, and then you do your coding, and that sort of cycle continues. The certainty of code writing is so much more certain than every other aspect of Genie's loop.

So whatever's going on under the hood, the model is really comfortable with writing code. There is no doubt, and it's, like, in the, in the token probabilities. One slightly different thing, I think, to how most of these models work is, at least for the most part, if you ask GPT-4 in ChatGPT to, to, to edit some code for you, it's gonna rewrite the entire snippet for you with the changes in place.

We train Genie to write diffs and, you know, essentially patches, right? Because it's more token efficient, and that is also fundamentally we don't write patches as humans, but it's like what the result of what we do is a patch, right?

When Genie writes code, I don't know how much it's leaning on the pre-training, like, code writing corpus, 'cause obviously it's just read code files there. It's obviously probably read a lot of patches, but I would wager it's probably read more code files than it has patches.

So it's probably leaning on a different part of its brain is my speculation. I have no proof for this. So I think the discipline of writing code is slightly different, but certainly it is its most comfortable state when it's writing code.

Once-- So once you get to that point, so long as you're not too deep into the context window, another thing that I'll bring up in that, in that blog post is, um- Performance of Genie over the length of the context window degrades fairly linearly.

So actually, I actually broke it down by probability of solving a SWE-Bench issue given the number of tokens of the context window. At 60K, it's basically 0.5. So if you go over 60K in context length, you are l- more likely to fail than you are to succeed just based on the amount of tokens you have on the context window.

And when I presented that to the fine-tuning team at OpenAI, that, that was super interesting to them as well, and that is more of a foundational model attribute than it is an us attribute. However, the attention mechanism works in, in GPT-4 or however, you know, they, they deal with the context window at that point is, you know, i- influencing how Genie is able to form.

Even though obviously all our, all our training data is perfect, right? So even if, like, stuff is being solved in 110,000 tokens, sort of that area, the training data still shows it being solved there. But it's just in practice the model is finding it much harder to solve stuff down that end of the context window.

Alessio37:32

That's the scale with the context, so for a 200K context size is 100K tokens, like the 0.5 or-

Ali Pullen37:39

Well, I, I don't, I don't know. Um-

Alessio37:39

Yeah, yeah.

Ali Pullen37:40

Yeah. But I, I, um, hope not. I hope you don't just take the context length and halve it and then say, "Oh, this is the usable context length." But what's been interesting is knowing that, actually really digging into the data, looking at the log probs, looking at how it performs over, over the entire window, it's influenced the short-term improvements we've made to Genie since we did the...

that got that score. So we actually made some small optimizations to try to make sure, as best we can without, like, overdoing it, trying to make sure that we can artificially make sort of stuff sit within that sort of range because we know that's our sort of zone.

And if we go outside of that, we're starting to push the limits and we're more likely to fail. So just doing that sort of analysis has been super useful without actually messing with anything, um, like more structural in, in getting more performance out of it.

Alessio38:26

What about, um, different languages? So in your technical report-

Ali Pullen38:30

Mm-hmm

Alessio38:30

... the data mix is 21% JavaScript, 21% Python, 14% TypeScript, 14% TSX. Um, your-

Ali Pullen38:37

Which is JavaScript, JavaScript, JavaScript.

Alessio38:39

Yeah, yeah.

Ali Pullen38:39

Yes. Yeah, yeah.

Alessio38:40

It's like 49% JavaScript.

Ali Pullen38:41

That's true. That's true. Although TypeScript is so much superior, but anyway.

Alessio38:44

Do you see how good is it at just, like, generalizing? You know, if you're writing Rust or C++ or whatever else, it's quite different.

Ali Pullen38:52

It's pretty good at generalizing. Um, there obviously, there, I think there's 15 languages in that technical report I think that we've, that we've covered. The ones that we picked in the highest mix were, uh, the ones that selfishly we internally use-

Alessio39:04

Mm-hmm

Ali Pullen39:05

... the most, and also that are, I'd argue, some of the most popular ones. When we have more resource as a company and more time and, you know, once all the craziness that has just happened sort of dies down a bit, we are going to, you know, work on that mix.

I'd love to see everything ideally be represented in a similar level as it is. If you, if you took GitHub as a data set-

Alessio39:25

Uh-huh

Ali Pullen39:25

... if you took, like, how are the languages broken down in terms of popularity, that would be my ideal data mix to start. It's just that it's, it's not cheap doing all this.

Alessio39:33

Yeah, yeah.

Ali Pullen39:33

So, um, yeah, trying to have an equal amount of, of Ruby and, and Rust and, and all these different things is just at the mo- at, at our current state is, is, is not really what we're looking for.

Alessio39:43

There's a lot of good Ruby in my GitHub profile. You can have it all.

Ali Pullen39:46

Well, okay. Perfect. We'll just train on that.

Alessio39:48

For running tests, it sounds easy, but it isn't, especially when you're working in enterprise code bases-

Ali Pullen39:54

Yes

Alessio39:54

... that are kind of, like, very hard to spin up.

Ali Pullen39:56

Yes.

Tests & Fine-tuning39:56

Alessio39:56

How do you set that up as, like, how do you make a model actually understand how to run a code base, which is different than writing code for the code base?

Ali Pullen40:04

The model itself is not in charge of, like, setting up the code base and running it. So Genie sits on top of GitHub, and if you have CI running GitHub, you have GitHub Actions and stuff like that, then Genie essentially makes a call out to that, runs your CI, sees the outputs, and then, like, moves on.

Making a model itself set up a repo wasn't scoped in what we wanted Genie to be able to do because, for the most part, like, at least most enterprises have some sort of CI pipeline running, and, like, a lot of...

If you're doing some even, like, a lot of hobbyist software development has some sort of, like, basic CI running as well. And that was, like, the lowest hanging fruit approach that we took. So when, when Genie ships, like, the way it will run its own code is it will basically run your CI, and it will, like, take the, um...

I'm not in charge of writing this. The rest of the team is. But I think it's the checks API on GitHub allows you to, like, grab that information, then throw it in the context window.

Alessio40:54

What's the handoff like with the person? So Genie, you give it a task.

Ali Pullen40:58

Mm-hmm.

Alessio40:59

And then how long are you supposed to supervise it for? Or are you just waiting for, like, the checks to eventually run, and then you see how it goes? Like, uh-

Ali Pullen41:07

Yeah, so it's-

Alessio41:08

What does it feel like?

Ali Pullen41:08

There are a couple of modes that it can run in. Essentially, it can run in, like, fully headless autonomous mode. So say you assign it a ticket in Linear or something, then it won't ask you for anything. It will just go ahead and try.

Or if you're in, like, the GUI on the website and you're using it, then you can give it a task and it, it might choose to ask you a clarifying question. So, like, if you ask it something super broad, it might just come back to you and say, "Hmm, what does that actually mean?"

Or, "Can you point me in the right direction for this?" Because, like, our decision internally was it's gonna piss people off way more if it just goes off and has, and makes a completely, like, ruined attempt at it because it just, like, from day one got the wrong idea.

So it, it can ask you for a lot of questions. And once it's going, much like a regular PR, you can leave review comments, issue comments, all these different things, and it, because, you know, it's been trained to be a software engineering colleague, responds in actually a better way than a real colleague would because it's less snarky and less high and mighty.

And also the amount of filtering it has to do for LGTM.

Alessio42:09

Yeah, yeah.

Ali Pullen42:09

When you train a model to, like, be a software engineer, essentially it's like you can just do anything. It's like, "Yeah, looks good to me, bro. Let's ship it."

Swyx42:16

I just wanted to dive in a little bit more on your experience with the fine-tuning team. John Ellard was publicly sort of very commentary supportive and, you know, was, was part of it. Like, what is it like working with him?

I also picked up that you initially started to fine-tune what was publicly available, the 16 to 32K range. You got access to do more than that.

Ali Pullen42:35

Mm-hmm.

Swyx42:35

You've also trained on billions of tokens instead of the usual millions range. Just, like, t- take us through that fine-tuning journey and any advice that you may have.

Ali Pullen42:44

It's been so cool, and this will be public by the time this goes out. Like, OpenAI themselves have said, "We are pushing the boundaries of what is possible with fine-tuning." Like, we are right on the edge, and like we are working, genuinely working with them in figuring out how stuff works, what works, what doesn't work because no one's doing...

No one else is doing what, what we're doing. They have found what we've been working on super interesting, which is why they- they've allowed us to do so much, like, interesting stuff. Working with John, I mean, I had a really good conversation with John yesterday.

We, we had a little brainstorm after the video we shot, and one of the thing- You mentioned the billions of tokens. One of the things we've noticed, and it's actually a very interesting problem for them as well when you're building out like a self-serve fine-tuning API, they have to decide how big your PEFT adapter, your LoRA adapter's gonna be in some way.

And like figuring that out is actually a really interesting problem because if you make it too big, and because they support datasets that are so small, you can put like 20 examples through or something like that. Like, if you had a really sparse large adapter, you're not gonna get any signal in that at all.

So they have to dynamically size these things, and there is an upper bound, and actually we use models that are larger than what's publicly available. Well, it's not even publicly available yet, but w- at the, when this goes out, it will be.

But we have larger LoRA adapters available to us just 'cause of the amount of data that we're pumping through it, and at that point you start seeing really interesting other things like you have to change your learning rate schedule and do all these different things that you don't have to do when you're on the smaller end of things.

So working with that team is such a privilege 'cause obviously they're like at the top of their field in, you know, in the fine-tuning space. So we'll- As we learn stuff, they're learning stuff, and one of the things that I think really catalyzed this relationship is when we first started working on Genie, like I delivered them a presentation which will eventually become the blog post that you'll love to read soon.

The information I gave them there I think is what showed them like, "Oh, wow, okay, these guys are really like pushing the boundaries of what we can do here." And truthfully, our dataset, we view our dataset right now as very small.

It's like the minimum that we're able to afford, literally afford right now to be able to produce a product like this, and it's only gonna get bigger. So yesterday while I was in their offices, I was basically so...

We were planning. We were like, "Okay, how... This is where we're going in the next six to 12 months." Like, we're putting our foot on the gas here 'cause this clearly works. Like, I've demonstrated this is a good, you know, the best approach so far, and I wanna see where it can go.

I wanna see what the scaling LoRA's like for the data, and at the moment, like it's hard to figure that out because you don't know when you're running into like saturating a PEFT adapter as opposed to actually like is this the model's limit?

Like, where is that? So finding all that stuff out is the work we're actively doing with them. And yeah, it's, it's gonna get more and more collaborative over the next few weeks as we, as we explore like larger adapters, pre-training extension, different things like that.

Swyx45:25

Awesome. I also wanted to talk briefly about the synthetic data process.

Ali Pullen45:28

Mm-hmm.

Swyx45:29

Um, one of your core insights was that the vast majority of the time the code that is published by a human is, is in a working state, and actually you need to fine-tune on non-working code.

Ali Pullen45:38

Yes.

Swyx45:39

So just, yeah, take us through that inspiration. How many rounds, uh, did you, did you do?

Ali Pullen45:44

Yeah, I mean-

Swyx45:44

Like, what's useful?

Ali Pullen45:44

Uh, it might, it might be generous to say that the vast majority of code is in a working state. I don't know if I can-

Swyx45:48

Yeah, I know. I was like, "That's very nice of you to say that my code works."

Ali Pullen45:52

Certainly it's not true for me. Um, no, I think that... So, so yeah, no, but it w- it was... You're right. It's an interesting problem and, and what we saw was when we didn't do that, obviously we'll just ho- You have to basically like one-shot the answer 'cause after that it's like, well, I've never seen iteration before.

How am I supposed to figure out how this works? So, so what, what the, um, what you're alluding to there is like the self-improvement loop that we started working on, and that was in sort of two parts. We, we synthetically generated runtime errors where we would intentionally mess with the AST to make stuff not work or index out of bounds or refer to a variable that doesn't exist or errors that the foundational models just make sometimes that you can't really avoid.

You can't expect it to be perfect. So we threw some of those in with a, with a, with a probability of happening, and on the self-improvement side, uh, I spoke about this in the, in the blog post. Essentially the idea is that you generate your data in sort of batches.

First batch is like perfect, like one exam- Like, here's the problem, here's the answer. Go train the model on it. And then for the second batch you then take the model that you trained before that can look like one commit into the future, and then you let it have the first attempt at solving the problem.

And hopefully it gets it wrong, and if it gets it wrong, then you have like, okay, now the code base is in this incorrect state. But I know what the correct state is, so I can do some diffing essentially to figure out how do I get the state that it's in now to the state that I want it in, and then you can train the model to then produce that diff next and so on and so on and so on.

So the model can then learn and also reason as to why it needs to make these changes to be able to learn how to like learn, like solve problems iteratively and learn from its mistakes and stuff like that.

Alessio47:32

And you pick the size of the dataset just based on how much money you can spend generating it. Maybe you think you could just make more and get better results.

Ali Pullen47:39

How, yes, how, what multiple of my monthly burn do I want to spend doing this? Yeah, basic- It was, it was very much related to, yeah, just like capital.

Alessio47:46

Yeah, yeah.

Ali Pullen47:46

And, um, yes, with any luck that will, that will be alleviated soon.

Swyx47:50

Very soon.

Ali Pullen47:51

Yeah.

Swyx47:51

Yeah, I like drawing references to other things that are happening in, in the, in the wild so 'cause we only get to release this podcast once a week.

Ali Pullen47:57

Mm-hmm.

Swyx47:57

The Llama 3 paper also had some really interesting, uh, thoughts on synthetic data for code.

Ali Pullen48:02

Mm-hmm.

Swyx48:02

I don't know if you have-

Ali Pullen48:04

I haven't had the chance to read it

Swyx48:04

... uh, reviewed that. Uh, but I'll highlight the, the back translation section because one of your dataset focuses is updating documentation. I think that translation between natural language English versus code-

Ali Pullen48:15

Yeah

Swyx48:16

... and back and forth I think is actually, actually a really ripe source of synthetic data.

Ali Pullen48:20

Mm-hmm.

Swyx48:20

And Llama 3 specifically called out that, that they trained on that.

Ali Pullen48:23

Yeah.

Swyx48:23

Uh, we should've gone more into that in our podcast with them, but- ... we, uh, we didn't, we didn't know. But, uh, there's a lot of interesting work on synthetic data stuff. We do have to wrap up soon, but I'm going to briefly touch on the submission process for SWE-Bench.

Ali Pullen48:35

Mm-hmm.

Benchmarks & Future48:35

Swyx48:35

So you have a 30% state-of-the-art SWE-Bench result-

Ali Pullen48:39

Yeah

Swyx48:39

... but it's not on the leaderboard because of submission issues. I don't know if you want to comment on, on like that stuff versus, uh, you know, we also have like a, we also wanna talk about SWE-Bench Verified.

Ali Pullen48:48

Mm-hmm.

Swyx48:49

Um, yeah, just anything on the benchmarking side.

Ali Pullen48:51

The potted history of this is, is, is quite simple actually. SWE-Bench up until, I want to say two weeks ago, but it might be less than that or more than that. But I think two weeks ago suddenly started mandating what they call trajectories when you submit.

So b- prior to this, essentially when you run SWE-Bench, you run it through their harness and out the other end you get a report.json, which is like, "Here's how many I resolved, here's how many I didn't resolve. These are the IDs of the ones I did, these are the ones, the IDs I didn't," and it gives you any ones that like might have errored or something like that.

And what you would submit would be all of your model patches that you outputted, and that report, and then you would like PR that into the SWE-Bench repo, and that would be it. That was the, still the case when we made our submission on whatever day it was.

They look at them every Monday. We submitted it at some point during the week. I wanna say it was four, four days before that. And, um, I sort of like sat back and waited. I assumed it would be fine.

When it came to Monday, um, they then said, "Actually, no, we want model trajectories." And I was like, "Okay, let me see what this is," and, and so on. I sort of dug into it. And like model trajectories are essentially the context window or like the reasoning process of like show your working, how did you get here?

If you do a math exam, show me your working. Whereas before they were like, "Just give me the final answer," now they wanna see the working, which I, I completely understand why they wanna see that. Like, SWE-Bench fundamentally is an academic research project, and it, they want all the stuff to be open source and public so people can learn from each other and improve and so on and on.

That's very good. I completely agree. However, at least for us, and the reason that we're not on the leaderboard is that obviously the model outputs that we generate are sort of a mirror of our training dataset, right? Like you train the model to do a certain thing and output a certain way.

Whatever you output looks like your training data. For the moment, as a closed source company, like fighting for an edge, we've decided not to publish that information for that exact reason. Like I don't want someone basically taking my trajes and then taking a model that's soon gonna be GA and just distilling it immediately and then having Genie for themselves.

And, you know, as, as a business owner, that's the decision I've had to make. The patches are still public, so like the, dare I say, traditional SWE-Bench submission, you can go to a GitHub repo and see it and run them for yourself and verify that the numbers come out correctly.

Like that is all, that is the potted reason as to why.

Swyx51:06

That's the story.

Ali Pullen51:06

That's the story.

Swyx51:07

Uh, SWE-Bench Verified, you have a score.

Ali Pullen51:09

I do have a score. I do have a score, 43.8%. It's one of those things where like there aren't that many people on the leaderboard yet , so you don't know how good or bad that is. Like-

Swyx51:17

And it's a, it's a smaller dataset, right?

Ali Pullen51:20

Oh, it's, it's great. So on a tangent, SWE-Bench, original SWE-Bench was 2,294 instances.

Swyx51:26

Which is expensive. It's like $8,000 to run.

Ali Pullen51:29

Oh, that's, that's cheap.

Swyx51:31

They're cheap when you think about it. Very reasonable.

Ali Pullen51:32

I, I, I don't, I don't, I know, at least for us, I don't, I, I don't even wanna say publicly how much it costs us, how much it costs us to run that thing. Expensive, slow, really like crap for iteration because like, you know, you make a change to your model, how does it do on SWE-Bench?

I guess that's why SWE-Bench Light existed, but SWE-Bench Light was not a-- It was, it was easy stuff, right? It wasn't a, a comprehensive measure of the overall thing. So we actually had the idea a month ago to what we were gonna call SWE-Bench Small, where we were gonna try to map out across SWE-Bench, like what is the distribution of like problem difficulty and all these different things, and try to come up with like 300 examples that sort of map that, where given a score on SWE-Bench Small, you could then predict your SWE-Bench Large score and sort of go from there.

Fortunately, OpenAI did that for us, and probably much better than we would've done. They used some human labelers and as obviously we're working with, with OpenAI quite closely, they talked to us about it and they, um, you know, were able to let us know what the instance ID were, IDs were that were in the, the new SWE-Bench version.

And then, uh, as soon as I had that, I could just take the report from the one that I'd run and just diff them and I was like, "Oh, we got 219 out of 500," which is 43.8%. Which is, to my knowledge, at least right now, state-of-the-art also, which makes sense.

But also GPT-4o gets, I believe, 33%, which is like-

Swyx52:52

Oh.

Ali Pullen52:53

I... Double-check that.

Swyx52:54

Mm-hmm.

Ali Pullen52:55

But I believe-

Swyx52:55

The August one, the, the new one.

Ali Pullen52:57

Yeah. It's in their blog post. I, I can't remember which one it was. I don't know what the model version was. But GPT-4, I believe, gets 33%, which is obviously like significantly better than what it got on the, um, original, like SWE- Bench, SWE-Bench-

Swyx53:10

2%.

Ali Pullen53:11

Yeah, yeah, yeah, exactly. Exactly.

Swyx53:12

Something ridiculously low.

Ali Pullen53:13

But no, SWE-Bench Verified, like it's so good. It's like it's smaller. We know that the problems are solvable. It's not gonna cost me a lot of money to run it. It keeps my iteration time, you know, lower. And there are also some things that we're gonna start to do internally when we run SWE-Bench to have more of an idea of how right our model is.

So one of the things I was talking to John about yesterday was SWE-Bench is a pass or fail, right? Like you, you, you either have solved the problem or you haven't. That is quite sparse. Like it doesn't give you a huge amount of information 'cause your model could have got a lot of it right.

Like looking through when you do a math paper, you could have got the reason, you know, you're working right until like the penultimate step and then you get it wrong. So we're gonna look into ways of measuring, okay, well, your model got it right up to this line, and then it diverged.

Um, and that's super easy to do because obviously you know the correct state of all of those questions. So I think one of the ways we're gonna keep improving Genie is by going more in depth and saying, "Okay, for the ones that failed, was it right at any point?

Where did it go wrong? How did it go wrong?" And then sort of trying to triage those sorts of issues.

Swyx54:17

So future plans, you have mentioned context extending an open source model. But basically, I think, you know, what the Genie is, is basically this like proprietary fine-tune dataset and process and software that, uh, you can add onto any model.

Is that the plan? That's the, that's the, the next year is gonna just be doing that?

Ali Pullen54:31

That, that is... We're gonna, we're gonna get really, we're gonna be the best in the world at doing that, um, and continue being the best in the world at doing that and throwing it as many models as we can, um, seeing what the performance is like and seeing what things improve performance in what places.

Um, and also making the dataset larger is like one of the biggest things that we're gonna be working on.

Swyx54:49

I think one of the decisions before you as, as a CEO is how much you have like the house model be like the one true thing-

Ali Pullen54:56

Mm.

Swyx54:56

-and then how much you spend time working on customer models.

Ali Pullen55:00

That's the thing that really, uh, that gets me so excited genuinely. Like, we have a version of Genie that we named after one of our employees. It's called The John. Uh- ... uh, we have a version of Genie that is fine-tuned on our code base.

So we basically, it's the base- base Genie, and then we run the same data pipeline that we run on, like, all the stuff that we do to generate the main dataset on our repo, and then all of a sudden you have, like, something that is both very good at software engineering, but is also extremely good at your repo, and that is phenomenal to use.

Like, it's really cool.

Swyx55:34

More broadly, outside of Cosine, what are you seeing? What, what trends are you, uh, seeing that you're, you're really excited by? Who's doing great work that you wanna call out?

Ali Pullen55:42

The, one of the, one of the ones that, I mean, it's, it's not an original choice, but Cursor are absolutely killing it. All the employees at Cosine love using it, and it's a really, really good example of, like, just getting, like, UX right, basically.

Like, the, the putting the LLM in the right place and letting it allow you, and getting out of the way when you don't want it there, and making it familiar 'cause it's still VS Code, and all these things.

They've, yeah, they've done an amazing job, and I think they just raised a round, so congrats on that to them.

Swyx56:09

Yeah.

Ali Pullen56:09

So, like, they're, they're doing amazing work.

Swyx56:11

The decision to fork VS Code I think was controversial. You guys started as a VS Code extension.

Ali Pullen56:15

We did, yeah.

Swyx56:16

Many, many, many people did that, and they did the one thing that no one wanted to do.

Ali Pullen56:19

I commend the bravery, honestly. Like, I commend the bravery 'cause, like, in hindsight obviously it's paid off. But at least for me in, in the moment, I was one of those people being like, "Is that gonna... Are people gonna do that?

Are people gonna download that?" And yes, obviously they are. Like, sure. Doing the hard thing, which is having worked on Genie recent- you know, for the past eight months or whatever, as taxing as it's been on us, like, one of the main things I have learnt from this is, like, no matter how small you are or how much resource you have, just, like, try to do the hard thing 'cause it, it, I think it has the biggest payoff.

Swyx56:52

More broadly, just like, uh, lessons that you've learned running your company.

Ali Pullen56:57

Oh.

Swyx56:57

You know, it's, it's been a two, it's been a two-year journey.

Ali Pullen56:59

Two-year journey. Um, I mean, it's better than any real job you could ever get. Like, I feel so lucky to be working in this area, like, especially, you know, it was so validating to hear it from the guys at OpenAI as well telling us, like, "We're on the cutting edge, on the bou- We're pushing the boundaries of what's possible with what we're doing."

Because, like, I get to do, I get to be paid to do this. You know, I, I have briefly, as you heard at the beginning, done real jobs and, and normal stuff, and, like, just being able to do this on the daily is so interesting and so cool.

It's like I pinch myself a lot genuinely about the fact that I can do this, and also, like, not only I can do this, but fortunately being a co-founder of the company, I have a huge amount of say as to where we go next.

And that is a big responsibility, but it's also so exciting to me 'cause I'm like, you know, steering the ship is, has been really interesting so far, and I like to think that we've got it right, you know, in the last, in the last sort of eight months or so.

Uh, and that this is, like, really the starting point of something massive to come.

Swyx57:56

Awesome. Call to action.

Ali Pullen57:57

Mm.

Swyx57:57

Uh, h- I assume you're hiring. I assume you're also looking for customers. What's the ideal customer, ideal employee?

Outro58:04

Ali Pullen58:04

On the customer side, honestly, people who are just willing to try something new. Like, the Genie UX is, is different to a conventional IDE. Give it a chance. Like, the, what we, we really do believe in this whole idea of, like, developers' work is gonna be abstracted, you know, levels higher than just the code.

We still let you touch the code. We still want you to dive into the code if you need to. But fundamentally we think that if you're trying to offload the coding to a model, the model should do the coding and you should be in charge of guiding the model.

So people who are willing to give something new a chance. Size of company, any, hon- honestly... Well, preferably the languages that are the most represented in our, in our training data. So, like, any... If you're, like, doing TypeScript, JavaScript, Python, Java, that sort of thing.

And in terms of size of company, like, so long as you're willing to try it, um, and there aren't any massive, like, infosec things that get in the way, like, it doesn't really matter. Like, code base size can be arbitrary for us.

We can deal with any code base size, and essentially any language, but your mileage may vary. But for the most part, like, anyone who's willing to give it a try is the ideal customer.

Swyx59:04

Mm.

Ali Pullen59:04

And on the employee front, honestly, we just want people who, um, we're, we're gonna be hiring both on, like, what we call, like, tradit- like, the traditional tech side. So, like, building the product essentially, and also hiring really heavily on the AI machine learning, um, dataset side as well.

Swyx59:21

Oh.

Ali Pullen59:21

And in both cases, essentially what we just want are, like, really passionate people who are obsessed with something and are really passionate about something, and are willing to, it sounds so corny, but, like, join us in what we're trying to do.

Like, we have a very big ambition and we're biting off a very large problem here, and people who can look at what we've done so far and been like, "Wow, that's really impressive. I want to do that kind of work.

I want to be pushing the boundaries. I want to be dealing with experimental stuff all the time and... But at the same time be putting it in people's hands and shipping it to people," and so on. So if that sounds, you know, amenable to anyone, that's the kind of person we're looking to apply.

Swyx1:00:00

Excellent. Any last words? Any Trump impressions that you-

Ali Pullen1:00:04

Did you like the Trump impression?

Swyx1:00:05

Yeah, everyone loved the Trump impression.

Ali Pullen1:00:07

Yeah. I mean, it's funny 'cause, like, I, I, I, I have some bloopers. I'll show you the bloopers after we've finished recording. I'll probably tweet them at some point.

Swyx1:00:13

Uh-huh.

Ali Pullen1:00:14

The initial cut of that video had me doing a Trump impression. I sort of sat down into the chair and been like, "Cosine is the most tremendous AI lab in the world. Unbelievable. I walked in here and I said, 'Wow, this is an amazing lab.'"

And, like, we sent it to some of our friends and they were like, "Nah, you can't cold open with Trump, man. You just can't." Like, "No one knows who you are."

Swyx1:00:33

You can end with it.

Ali Pullen1:00:34

But you can end with it. Now that that has gone out, we can now, um- ... we can now post the rest of the bloopers, which are essentially me just, like, fluffing my lines the entire time and screaming at my co-founder out of frustration.

So yeah.

Swyx1:00:46

Well, it was very well executed. Uh, actually, very few people do the content creator that you did. I, I'm, as a sort of developer relations person, I, I'm actually excited by that stuff. But-

Ali Pullen1:00:54

Mm

Swyx1:00:54

... um, well, thank you for coming on. Very, very short notice. I hope you have a safe flight back and, uh, excited to see the, the full launch. Um, I think this is a super fruitful area and, uh, congrats on your launch.

Ali Pullen1:01:04

Thank you so much for having me. Cheers.