LALatent SpaceOct 16, 2025· 1:08:23

Why RL Won — Kyle Corbitt, OpenPipe (acq. CoreWeave)

Kyle Corbitt, co-founder and CEO of OpenPipe (acquired by CoreWeave), explains why reinforcement learning has replaced supervised fine-tuning for training reliable AI agents. He argues GRPO is a dead end due to its requirement for perfectly reproducible parallel rollouts, which is extremely hard in practice. Instead, OpenPipe’s RULER uses relative LLM-as-judge rewards, achieving state-of-the-art performance even with a weak judge. Corbitt reports that 90% of AI projects remain stuck in proof-of-concept due to reliability issues, and that LoRAs are underrated for production while GEPA failed in his tests. He predicts continuous RL from real-world experience can unlock 10x more inference demand.

  1. 0:00YC Roots
  2. 2:29OpenPipe Origins
  3. 7:49Fine-Tuning Wave
  4. 14:51RL Pivot
  5. 21:09GRPO & Sandboxes
  6. 36:00Prompt Wars
  7. 44:33Model Debate
  8. 50:19RULER & Worlds
  9. 1:00:12Acquired & Vision
  10. 1:06:18YC Advice

Powered by PodHood

Transcript

YC Roots0:00

Alessio0:04

Hey everyone, welcome to the Latent Space Podcast. This is Alessio, founder of Kernel Labs, and I'm joined by Swyx, editor of Latent Space.

Swyx0:11

Hello, hello, and we're so excited to have Kyle finally in the studio. Welcome.

Kyle Corbitt0:16

Hey, I'm very excited to be here.

Swyx0:17

Uh, Kyle, you're CEO, founder-

Kyle Corbitt0:19

Uh, yeah

Swyx0:20

... uh, co-founder, founder-

Kyle Corbitt0:20

Co-founder, CEO, yeah

Swyx0:22

... of OpenPipe, which, uh, started two years ago and recently got acquired by CoreWeave. Congrats.

Kyle Corbitt0:28

Thanks.

Swyx0:29

Uh, we're-- I think you might be our first, like, started and exited founder that we've had on the pod, maybe-ish. I don't know.

Kyle Corbitt0:36

Possibly.

Swyx0:36

I'm not, I'm not keeping-

Alessio0:37

Especially on that timeline.

Kyle Corbitt0:38

Well, I don't think I was exited when we... I don't remember if it-- if, uh, we set this up before or after we, um, announced we were getting acquired.

Swyx0:46

I specifically pinged you because, uh, you got a-- I think you got acquired.

Kyle Corbitt0:50

Okay.

Swyx0:50

You've been on my list, uh, to watch. Obviously, you've spoken three times at AIE, and you've been on my list of, like, when is it a good time to have a OpenPipe or fine-tuning RL discussion, and then you got acquired, and I'm like, "Okay, yeah, that's a, that's a good, that's a good time to talk about it."

Also, because I think, like, it gives us a window to talk about acquisitions, consolidation, like, what should be an independent company, what, what maybe doesn't have to be. Anyway, but we-we'll maybe do this chronologically, so we don't, we don't get too far ahead of ourselves.

You were famously director of Startup School.

Kyle Corbitt1:21

Yes.

Swyx1:22

Maybe for people who don't know, like, what, what is Startup School? Did that make you become-- like, fall in love with the color orange?

Kyle Corbitt1:29

Yes. I'm, I'm wearing an orange shirt for those who are listening. A very bright orange shirt. This is, this is my conference shirt, and I felt like, you know, it was appropriate for the, the pod as well. Um, so yes, I was at, I was at Y Combinator for about four and a half years and led the Startup School team there.

So Startup School, it's, it's changed over the years. It meant one thing before I was there. It means another thing now. But during the time I was at YC, Startup School was basically all of the external-facing, a lot of the content, um, certainly all of the tech.

So it was, it was things like we had a, like, like a MOOC effectively, where founders could come in, they could learn about how to start a company. They could get advice from YC founders, YC partners. Um, we had a co-founder matching, uh, service that we built, which actually worked really well.

Um, we got a lot of, of people through. Our total, like, you know... I, I guess technically I can't-- that probably doesn't matter anymore, but a very large fraction of the batches that went through YC while I was there, um, were directly attributable to people that we found and, and ended up recruiting to YC, um, through their experience too at, at Startup School.

So that was kind of what we were working on.

Swyx2:29

Yeah, I was g- I always kinda con-consider it as like the, the scout program for YC.

OpenPipe Origins2:29

Kyle Corbitt2:32

Yeah.

Swyx2:32

Right? Like the YC before the YC. Any notable, like, famous people that, that met as a part of your co-founder match-matching? 'Cause I'm always very negative on those things because, like-

Kyle Corbitt2:41

Yeah

Swyx2:41

... it's like online dating.

Kyle Corbitt2:42

Mm-hmm.

Swyx2:42

Like the, the chances of success is super low.

Kyle Corbitt2:44

Yeah.

Swyx2:44

But when it works, it, it's really nice.

Kyle Corbitt2:46

You know, that's a great question. I left-- So we launched that product probably nine months before I left, and so I don't know what the, the long-term outcomes were of that specifically.

Swyx2:56

Yeah. So you left YC, you spent a year in the kind of the wilderness. You went to YC, uh, S '23.

Kyle Corbitt3:02

Mm-hmm.

Swyx3:02

Uh, what's that journey like? What's the-

Kyle Corbitt3:04

You know, I was very excited about AI things in general. Um, this was-- So I left YC, I guess, uh, beginning of 2022, and I was trying out a bunch of different things. Um, ended up landing on what turned into OpenPipe in early 2023.

This was, uh... Let's see. So I'd been working-- So my, my co-founder is my brother, um, my little brother, which has been fun journey on its own. We were looking at different ideas, and one thing we realized was we actually started the company immediately after the GPT-4 launch.

And what we saw as the opportunity in the market at the time, which has changed since then, was GPT-4 was insanely expensive and extremely powerful. But there was an opportunity to distill like specific workflows from GPT-4 down to much smaller, much cheaper models, and there was like a very clear value prop there, given how expensive GPT-4 was.

Um, it was hard to deploy in production, but you could sort of like take those abilities and deploy them much more cheaply. So, so that was kind of the first thing we built was this kind of very managed, very clean, um, distillation flow.

Alessio4:02

What was that process like in the beginning to like get people to actually care? Because I'm assuming most people are doing experimentation, but like they don't really have these large production workflows that they needed to like distill down.

Kyle Corbitt4:12

Mm-hmm.

Alessio4:13

And then I think maybe once we got there, the models get cheaper and faster. So what was like the initial, you know, six, nine months of the company through the evolution of the models?

Kyle Corbitt4:21

Yeah. So it worked. It, it was great. So I mean, it did take us a while. I guess we formed the company early, may-maybe March of 202-2023. By the time we launched our product, it was August, I wanna say.

There were, there were some like different things we were trying in between, and actually it was not hard, uh, to find people and get them excited. There weren't very many. I mean, this was even late 2023, there weren't very many people in production.

But anyone who did have production workflows, it was extremely painful. Like, you know, they're, they're paying hundreds of thousands of dollars a month to OpenAI, so it was very easy to convince them to try this out. And so we got our first three customers after launching probably within a month, and we were doing significant revenue.

Um, over the next six months, we actually got to a million in ARR, um, over about a eight-month period following that launch, so by the latter part of 2024. So actually, yes, initial traction was, was, was super strong, um, very clear value prop.

Um, but then as you were alluding to kind of like there was just this slow march of like the frontier model token prices just dropping over and over- ... by, you know, three, five X over and over again, which kind of ate away at our, our value prop over time.

Alessio5:23

What was the process of like fine-tuning the model? Because even the open models were not that great, you know? And so what were maybe the bottlenecks, like instead of having three to get to like thirty customers? Did you feel like in the beginning it was like a matter of like just the market growing, like the open source models not being good enough, like the fine-tuning not being simple, efficient enough?

Kyle Corbitt5:41

The pain point, I guess repeating what I said before, was the price was too high on the closed models, but you couldn't just drop in an open model and replace them 'cause like you're saying, the quality was quite bad, especially as you're moving to, to smaller model sizes.

But larger models, open models weren't even available at that time. So, so that's kind of where the value prop was, was like, "Hey, the closed models are too expensive, at least the ones that are performant enough to, to do the job.

The open ones are not good enough. We have like a very clear managed flow." Um, the way the flow worked was, was quite simple. You simply put in our SDK, it's a drop-in replacement for the OpenAI SDK- It's capturing-- You continue to use GPT-4 in production for a period of time.

We're capturing the requests and responses, and then we have just a, a very clean managed flow where it's like, okay, at some point you say, "Hey, I wanna distill this down," and you-

Swyx6:22

Yeah

Kyle Corbitt6:23

... you train on that. And then, you know, we provided an API that was a direct drop-in replacement. You would just change kind of the inference URL, and you were using your own model, and it, it, your app continued working.

Swyx6:33

Yeah. I-I think the market analysis here, because I was also exploring starting a business around that at, at, at the time, and that's why-

Kyle Corbitt6:39

I remember that

Swyx6:40

... I, I ended up-

Kyle Corbitt6:41

Yeah

Swyx6:41

... not investing, was basically you get squeezed between the GPU providers who also wanna do fine-tuning as a service because then that, that makes people more sticky, uh, and the, the labs who keep putting out distilled versions of, like, something, whatever mini versions of their, their models.

What was the analysis on the, on the Neo cloud side? Because you co- you kinda also want to host the inference.

Kyle Corbitt7:03

Yeah. Honestly, we-- So we, we, like I was saying, felt very squeezed from the frontier labs that were putting out just more capable models at lower cost. I did not see the competition ever really materialize from the Neo clouds, from the, the GPU providers.

Um, everybody had an offering in fine-tuning. When we would talk to customers, nobody used them because they just were really hard to use.

Swyx7:22

Mm-hmm.

Kyle Corbitt7:23

Um, so I do think that, like, you know, call it a product thing, I guess.

Swyx7:26

Like it's not their focus.

Kyle Corbitt7:27

Yeah.

Swyx7:28

Who cares? Yeah. Interesting. Developer experience matters.

Kyle Corbitt7:31

It does, yeah. Still does. Did. I don't know. Maybe it doesn't matter anymore. Now we, we just have coding models do everything for us.

Swyx7:37

No, it still does.

Alessio7:37

Oh, experience.

Swyx7:37

Like, when you have, when you have Thinking Machines launching an API and people getting excited about the API, you're like, "Yeah, okay, that's-there's a pure developer experience there."

Kyle Corbitt7:44

That's fair, yeah.

Swyx7:45

Yeah.

Alessio7:46

What's the-- I'm just going through the chronological list here.

Kyle Corbitt7:49

Yeah.

Fine-Tuning Wave7:49

Alessio7:49

Was, like, the Mistral 7B fine-tune kinda like one of the big inflection points, like, in the history of the company? It's like, okay, this is, like, a good open model and, like, the 7B size, or is it just

Swyx8:01

Yeah.

Alessio8:01

... most

Swyx8:01

Mistral and Mixtral. That-

Alessio8:03

Yeah

Swyx8:03

... there was, like, a golden period of fine-tuning-

Alessio8:05

Mm-hmm

Swyx8:05

... startups because Mistral was, like, a credible open source model.

Kyle Corbitt8:09

Yeah. They were really strong models, um, better than the Llama 2 that they were, you know, effectively replacing. Um, and they also had the super open license, uh, which, which I think the licensing has become maybe less of a concern over time at the margin- ...

because people are getting used to maybe. But, um, at the time, that was, like, a pretty big deal that they had this fully open Apache 2 license and, you know... Yeah, maybe, maybe they have their own, like, IP issues with how they train it.

I don't know. I have no inside information there. But at least the guarantee they were making to people using their model is, is-

Swyx8:39

Yeah

Kyle Corbitt8:39

... openness

Swyx8:39

I call this, uh, Mistral washing.

Kyle Corbitt8:41

Yes.

Swyx8:41

As long as it's like-- It's, it's, it's, uh, you know, comes from the sparkling region of France called Mistral- ... it's okay. Don't ask about what goes into it.

Kyle Corbitt8:48

There's, there's, there's plausible deniability-

Swyx8:50

Exactly

Kyle Corbitt8:50

... arm's length connection there, yeah.

Swyx8:51

Okay. There, there was this Mistral period. Uh, Jan 2024, you, you talked about SLora, and that was a, that was a period of time where LoRAs, uh, became more important. I feel like they, they then became less important, and I don't know what-what's, like, the rise and fall of LoRAs for, for you as a, as a business.

Kyle Corbitt9:07

Yeah. So, so LoRAs, um, have really, really... So if you're predicate on the fact that you're doing fine-tuning at all, LoRAs have very, very attractive properties relative to doing a fi- full fine-tune, right? Because if you're doing a LoRA, you can add training time.

It makes-- It helps some. You're using less memory to train. But it really-- where it really helps you out is at inference time because if you're doing LoRAs, then when you deploy for inference, you can multiplex, you know, basically an arbitrarily large number of LoRAs on the same GPU deployment.

That lets you do things like do per token pricing as opposed to GPU hour pricing. Um, it just gives you much more flexibility, uh, at deployment time. I'm actually still a LoRA bull, like, for the record. Uh, you know, you're talking about the rise and fall.

I think, I think LoRAs, you know, their, their, their future is still out there.

Swyx9:48

I mean, they're cool again because of Thinking Machines.

Kyle Corbitt9:49

Yeah. And I felt very vindicated by that blog post- ... for the record. Um, uh, uh, just I guess for listeners, Thinking Machines put out, like, uh, a week or two ago a, a blog post doing, um, quite a lot of research on the, the trade-offs between LoRAs and full fine-tuning in, in various different training regimes.

I think the reason LoRAs were uncool for a while was, was mostly just 'cause, like, fine-tuning was uncool. Like, I think if you're doing fine-tuning anyway, like, LoRAs are still, like, you know, in, in many cases the way you wanna do it.

But not that many people are doing fine-tuning.

Swyx10:16

As a marketing guy, LoRAs had bad marketing.

Kyle Corbitt10:18

Hmm.

Swyx10:18

Like, they, they were just like, "Oh, like, you can't afford fu- full fine-tuning? Here's like-

Kyle Corbitt10:22

Yeah

Swyx10:22

... here's like the Wa- Walmart, like, store brand- ... fine-tuning."

Kyle Corbitt10:26

No, that's, that's fair. There is some of that. I think we didn't have a huge issue. Like, we've had to do some user education, like, "Hey, just try it." I think for the training runs that-- like the types of training runs that we're interested in, where it's like, "Hey, I'm doing a relatively lightweight customization of an existing model for a specific task," there's really no downside to using a LoRA, and there's a lot of, like, upsides from an, like, infra simplicity point of view.

I agree that there's, like, a branding issue around that. Hopefully, the Thinking Machines blog post kind of like-

Swyx10:52

Yeah

Kyle Corbitt10:52

... you know, addressed that.

Swyx10:53

So like, yeah, rank one.

Kyle Corbitt10:54

Uh-huh.

Swyx10:54

I, and like, you know, I think there's, there's different hyperparameters to LoRAs that you can use to, to make yourself happy. The fact that John Schulman was like, "Nope," like, "we're actually banking the company on this"-

Kyle Corbitt11:05

Mm-hmm

Swyx11:05

... at least for now, is a pretty big vote of confidence. I, you know, I, I feel-- it's-- I think it's surprising that no one's done the research prior to, to them.

Kyle Corbitt11:13

Yeah.

Swyx11:13

This, this extensively.

Kyle Corbitt11:13

And I was talking to someone at Thinking Machines prior to their launch who had come from one of the big labs, and, and what that research showed was like, oh, no, everyone doing post-training research inside this big lab uses LoRAs.

I mean, not for like the full run, but, like, when they're doing, like, their experiments, they'll, they'll just use LoRAs on, on a base model to, to run the experiments, and it works fine.

Swyx11:29

For listeners of the pod, that, that was leaked, uh, in, in one of the pods that we released, but it's up to you to find it. Cool. Uh, and then so then we-- it, it was the first Worlds Fair.

You talked about you, you probably don't need fine-tuning as, as a fine-tuning founder. Basically, I think your, your talks are really good. I would recommend people watch all of them. What I pulled out was you had a piece of advice.

So your, your talk title was obviously somewhat intentionally clickbaity, but your, your actual advice on when people should fine-tune is when it's cost, latency, or quality consistency that you, that you really care about.

Kyle Corbitt12:00

Yeah. I mostly stand by that. I, I don't think it's changed. And the biggest one we see today, and this is true for kind of like classical SFT, it's also true for the RL stuff we're doing today. Cross my fingers it's not always the thing.

But the, the main one I see that really drives fine-tuning is if you have to move to a smaller model, and it's typically for latency reasons, and this is usually, like, real-time voice. So if you're sort of forced into a smaller model anyway- Then there's a very high chance that, that doing some tuning on that model is going to get you-- like it will be necessary basically to have a successful deployment.

So we see that a lot coming from customers that again have those latency requirements. There's other reasons as well. Sometimes, for whatever reason, you really have to deploy on a single GPU, you have to deploy within your own cloud, and you want a, you know, you, you basically have to use a smaller model to do that.

So basically, in the case where you're forced to a smaller model anyway, then fine-tuning it is often necessary. I would say for ninety percent of use cases where you aren't forced to a smaller model, then it's still not a good ROI, and you, you, you probably shouldn't invest in it today.

Alessio13:01

How do you quantify these things? So cost, right, could always be lower.

Kyle Corbitt13:05

Mm-hmm.

Alessio13:06

So is there kind of like a threshold of like cost to ROI? Like because it's also hard to figure out how much it's gonna cost to do the fine-tune because you need to get the data and all of that.

Like do you have a mental model of that?

Kyle Corbitt13:17

This is sort of like a, a function of the total amount of overhead required. I'd say there's, there's two parts on the cost side and then, you know, there's, there's one-- multiple parts on, on the benefit side. On the cost side, the main things you have to think about are the upfront effort required to get an actual like training system set up for your task.

And that can be quite variable, but I would say at a minimum, you're gonna have to dedicate a couple of weeks of like a fairly competent engineer's time. And if you have-- and if you have like a very complex system and you're doing RL and you need to set up a whole environment, it could be a lot longer.

It could, it could be, you know, a couple of months of time. So that's just like a fixed cost you have to pay. There's also like an ongoing carrying cost where once you've committed to doing fine-tuning, it does make other parts of your stack less flexible, less nimble, because whenever you're updating your prompt or like you're adding new context or whatever, like now you have to like, you know, spend a few hours training a model, and that's just gonna like slow down your, your iteration cycle, which is a real cost, and in many cases, that's the larger cost.

So you only wanna do that if like the benefits are large enough. The dollar cost, I would say, is basically never a factor. Um, it's just so much less than the time, the amount you're spending with this engineer to, to do the work that it's not-- I mean, it's, you know, each of these runs is, uh, between five and a couple hundred dollars, um, and it's just you, you, you don't have to do that many of them.

Alessio14:39

Yeah, because most of the data is like first party.

Kyle Corbitt14:41

Mm-hmm. Yeah.

Alessio14:42

Right. Okay. When was the switch to RL? Was it when o1-preview came out, you were maybe like, "Okay, it's time to move on from SFT," or?

Kyle Corbitt14:51

Yeah. So that was a big moment for us, um, with, you know, there's all the leaks before that about Strawberry and all this and like, you know, a lot of people talking about, okay, how are they doing it?

RL Pivot14:51

Kyle Corbitt14:58

Um, we realized through that, that like, okay, someone's figured out how to make RL actually work with LLMs, which was not a thing. I mean, it was a thing that like some people had played around with before that, but it wasn't like a thing many people were thinking about.

And so our bet at that point was, yes, let's figure out whether this works for task-specific thing. And the space we, we just-- I think it's important to kind of like tease out different parts of the market. I think with the release of o1, and this has been like proved out many times with releases since then, I think like there's now like a very strong consensus that like, okay, on the frontier model, like general purpose model side, investments in RL are paying off.

I, I think I, I don't think most people would argue with that. You're-- Especially as, as you're getting into these agentic tasks, um, and training them to do that, like it seems very clear. Well, obviously, the big labs are paying like ridiculous amounts of money for these environments and everything, but also like they're actually getting really good results.

The, the, the models coming out, you know, we're seeing it especially on the coding model side, but like in other, in other contexts as well, we're seeing the sort of especially agentic use is working way better because of this.

So I think like even late 2024, it was pretty clear that like RL was gonna work in that context. And then the question in our mind was like, can we apply this in a different segment of the business, which is kind of like task-specific customization?

And so the question is like, does that work well? How much effort does that take? Is it going to be something that ends up being unnecessary because, oh, the, the big labs can just like train on every single task, and the base models are gonna be just good at everything, and so there, there's, you know, no benefit to it.

So those were kind of the open questions in, in our mind, but it seemed like there was like at least a good enough bet that, you know, we wanted to try it out.

Alessio16:32

Yeah. And you had this, um, agent reinforcement training framework-

Kyle Corbitt16:36

Mm-hmm.

Alessio16:36

-and you did the email agent. That's kinda like the first proof of concept. Was that obvious to do email? Was it obvious to call it that way? What, what was like the behind the scene? How should we package this?

Kyle Corbitt16:46

So what I told our team, and this was-- We decided to go all in on RL in January of 2025, and we've been doing some experience before that. We released before that kind of like an RL model that had, you know, would generate like Hacker News titles from, from articles- -um, which is a fun project.

So we'd done a little bit before that, but that was kind of like we're like, "Hey, we're gonna bet the company on," uh, not in a literal sense. Like we-

Alessio17:05

Right.

Kyle Corbitt17:05

-we could've done something else later, but like this is like the thing that we're gonna spend all of our time working on for, for at least a few months. And like what I told our team at that time in, in January of '25 was like, there's probably like a twenty-five percent chance that this is the right direction in the sense that like a year or two years from now, all the companies, you know, everyone doing inference should be doing RL and task-specific training so that like their model's just way, way better at their task.

Is a relatively low chance, but it was sort of like one of those big if true things. Like if that is true, if it turns out that like just doing RL on your task is just like something everyone should be doing, and it's, and it's just, you know, teaching these agents continually, teaching them through experience is just going to be a huge benefit, then like being the first people working on that would be a really, really like awesome position to be in.

So that's how we thought about it, is like, you know, less than fifty percent chance, but really big outcome if not, if so. I think since that time, and I've been very transparent with this, like with our team and like when I'm talking to other people, like I don't think the chance that that is the right approach is a hundred percent yet.

I think that we're still in the process, even after going through this and, and, you know, doing that of like figuring out. But the probabilities in my mind are going in the right direction. Like now I think they're actually-- Like today, I was actually just thinking about this with another conversation.

I think that the chances that like everyone should be, or you know, everyone who's deploying an agent at scale should be doing RL with it, either as part of sort of like a, you know, like pre-deployment or even like continuously as it's deployed, that that's like the pattern that that's gonna get to- I'd say there's like a 55, 60% chance that that's just like the better thing to do, and that's informed by kind of like our experiments working with customers.

So anyway, not 100%, but like going all the way back to your question, like no, it was not obvious. It, it was an informed bet. You know, it's, it's still a bet, but one that I'm, I'm feeling pretty good about right now.

Swyx18:48

One thing I think that is tricky about just as, as you're onboarding onto this space is all the math.

Kyle Corbitt18:53

Hmm.

Swyx18:54

Uh, I remember reading the DPO paper, I think, I think they were at NeurIPS for 2023-

Kyle Corbitt18:59

Hmm

Swyx18:59

... and people were very excited about it. Some of it's, like, just being pretentious for a paper-

Kyle Corbitt19:03

Mm-hmm

Swyx19:03

... but some of it's actually like real complexity. You know, you don't have like a PhD, like a prior sort of, uh, ML background. How do you sort of come to grips with it? Like, what were the best ways to get around it for you?

Kyle Corbitt19:15

I would probably push back on that a little bit. I don't think the math is actually that complicated. I think that like when you, you know, you, you see the PPO equation or something with all the symbols, like if, if that's your first intro to it, then it feels very complicated.

But I think like if you were to show that exact same equation, just like code, not... maybe not PyTorch code 'cause that you also have to like understand. But if you just like did the naive implementation in like Python and like showed someone like, "Hey, this is, this is kind of like how we're, we're computing the loss here," who was like a strong engineer, like I think it's actually like quite grok-able.

So yeah, I mean, like I, I don't think it's like the barrier to entry is that high. I think you just have to like believe you can do it, and then like spend some time staring at it. That would be what I would recommend is like, uh, you know, you can read the papers and look at the equation.

I, I think actually this is one area where, where LMS have been super helpful. If I'm reading a new paper and I look at one of those equations and I'm like, "I don't understand how this new term they introduced intro- like corresponds to like the, these other terms," then I can like dump like all the context around it into, you know, GPT 5 and say like, "Hey, can you like write this out of Python for me and show me what, what they're doing differently?"

And that's super helpful for kind of like my background, I guess.

Swyx20:22

Yep. The way I put it is I wish that all these papers would just publish with pseudocode or Python, just straight up Python-

Kyle Corbitt20:28

Mm-hmm

Swyx20:29

... instead of math.

Kyle Corbitt20:30

Yeah.

Swyx20:30

'Cause like you actually just need to look at the implementation.

Kyle Corbitt20:32

Yeah, totally. I know like Jeremy Howard's been beating this drum for, for, for years-

Swyx20:34

Yeah, I know

Kyle Corbitt20:35

... and I mostly agree with him.

Swyx20:35

Well, I mean, there, there's a, there's a little website called Papers With Code.

Kyle Corbitt20:38

Mm-hmm.

Swyx20:39

And like people just keep not following it.

Kyle Corbitt20:40

Hmm.

Swyx20:41

I remember interviewing the DPO guys when they were, when they were at NeurIPS, and it was just like they were just very obsessed with like proving in principle equivalence to PPO, and like it was just v- it was like it was very hard to follow.

I'll definitely say that. And then I think like now obviously at some point like GRPO kinda took over the, the, the general consensus. It was very strange because I think when DeepSeek first like started talking about it, it was viewed as an optimization.

GRPO & Sandboxes21:09

Swyx21:09

They tend to just generally couch everything as an optimization. But I think like the leader insight, which I think you touched on in one of your blog posts, was that no, actually it, it, it makes comparisons independent rather than global, and like that's, that's actually what unlocks some f- some amount of like sort of self-supervised RL.

Kyle Corbitt21:28

Mm-hmm. Yeah, I mean, it's interesting. There's real pros and cons if you're moving from PPO or, or something similar to it to GRPO. There are some big pros. I mean, one pro is just sort of like operational simplicity, like there's a whole extra model you need for this value model you need for PPO that you can throw away with GRPO, and that just like makes your life easier.

You don't have to train that model, but also like there's like no hyperparameters around that model that you have to f- configure, so, so that, that's nice. Another thing is the benefit that you're talking about, which we've observed.

So the way GRPO works is, is you have to do like, you know, a set of, of different, uh, trajectories or a set of different rollouts, um, all in parallel with the exact same environment, the exact same conditions, and then you score each of them.

And GRPO uses the differences in those scores to promote the trajectories that do better and, and sort of like decrease the probability of the ones that did worse. Because they do it in sort of a group relative way, the only...

it lets you be a little bit looser with how you score them potentially. Like you don't have to necessarily have a globally aware scoring function. You just need some scoring function that is able to distinguish between this small set of things you have in front of you.

And, and that's easier. That's easier for a human. You know? If you, if you tell a human which of these-- choose which of these is better, it's easier for them to do than say like, "Is this one good or bad in absolute terms?"

Swyx22:43

Objectively. Yeah.

Kyle Corbitt22:44

Yeah. So that's nice. The big downside, the huge downside of GRPO, and I think actually the reason why GRPO actually is, is, is likely to be a dead end and we probably will not be continue using it indefinitely, the fact that you need to have these parallel rollouts in order to train on it is actually the...

like that makes the data generation much more complicated because you need a fully reproducible environment to be able to do these sort of parallel rollouts. And it turns out in practice, that's like getting that set up is the hardest challenge today with getting RL working is, is like actually designing this robust, uh, reusable, you know, environment that you can run all of this training in.

Most companies... and, and that's not true. Like sometimes that's easy to do. Like, like there's certain situations where, where you can do that. But for the work we do at least, where we're training agents on real code bases to like operate like, you know, real applications, it turns out it's like really, really hard to sandbox those things in a way that's like totally reproducible.

And PPO, now in practice a lot of times when you're training with PPO, you also will use an environment like that because it lets you do a bunch of runs and, and be more data efficient. But at least in principle, you have the option with PPO, you can actually like purely train on like, say, real production traces of like real people interacting with your app, and so you don't have to have a simulated environment at all, which makes the deployment like much easier.

Swyx24:02

Can you double-click on why it's hard to do the sandboxing? Because in principle you just capture all the inputs.

Kyle Corbitt24:09

Yeah. Well, you don't need to just capture all the inputs. You, you need, you need a system that reacts the same way, uh, your production system does. That's... and, and in many different ways. And, um, so let's say your, your Airbnb, right?

I'm bringing this up because this is like an example of one that like, you know- Companies have gone out and built sandboxes. Like, if you're Airbnb and you're trying to, um, you wanna train an agent to like... Maybe you're not Airbnb, fine.

You're, you're a company like us that's trying to train an agent to, like, do really well at operating Airbnb and booking on your behalf, right? Like, you have to f- build a copy of the Airbnb website-

Alessio24:42

Mm

Kyle Corbitt24:42

... that reacts to you as the user the exact same way that the real one does with the same failure modes, right? 'Cause if you don't include the same failure modes and bugs they have, then, like, one of those bug- when one of those bugs comes up in production, your agent's gonna have no idea what to do with it, and it's just gonna fall over.

You also need to simulate if this is, like, a sort of cooperative agent, right, where it's getting human input as well and kind of, like, working with the human to get something done, which in practice is the way a lot of these are deployed.

You also need to simulate the user. And, I mean, you can do the naive thing and just say, "Oh, we're, we're gonna have a separate LLM that, you know, with a system prompt that is, like, the user simulator," and we do that.

But it's like, okay, but, like, the breadth of ways a user might respond, there's, like, a lot more diversity in that than the actual diversity you'll get in practice when you have this, like, simulated user. And so then it's like, okay, well, is this environment close enough to how a real user would interact that like, you know, if, if a user says something different, that it's gonna know what to do?

And the answer in many cases is no. If, if you're just purely training on kind of like an LLM user simulator, it's gonna have its own idea of, like, what the correct way to answer is, and the breadth of, like, a way a human might respond in the situation is, is wider and, and, and your agent just may not be able to deal with that.

Alessio25:46

Do you feel like it's hard to build the simulations as a company that needs to build the product that lets everybody do it? Or do you feel like even for the individual companies that own the code base that are, like, domain experts in their own product, it's still just, like, a very hard infrastructure problem?

Kyle Corbitt26:01

I think it's still very hard. You know, like, ideally, all companies should have this anyway 'cause they're g- you know, if you're doing end-to-end-

Alessio26:06

Right

Kyle Corbitt26:06

... testing. Like, theoretically, if you're following best practices, you would have one of those set up. When we talk to enterprises almost universally, that's like- ... not something that really exists. So there are some startups-- Like, there's some companies we've talked to that do have it, and, and we can just, like, use that, but it's, it's a very, very small number that, that actually have an environment like that.

Um, and I think it's hard to do. And, and, like, there's lots of, like, weird bugs that don't show up in an environment like that. And, and even if they do have a testing environment, they don't have it populated with, like, full realistic data, which is also, like, important so that the-- it, it understands how to, you know, interact.

So I think in practice it's hard in both cases. Maybe it's easier for the company, but at the same time, depending on, you know, the quality of the company's engineers, it might not be easy for them either.

Alessio26:49

Yeah. How do you classify the types of environments? So you have formal environments like a compiler. You know, you can put in there, so you don't need to do any work, they just work. Then you have this kind of like RL environment startups in a way that are building a bank environment.

They're building these things that are not digital twins or whatever term of like the actual environments, but they're like close to it.

Kyle Corbitt27:11

Mm-hmm.

Alessio27:12

And then on top of it, you have helping people trying to build the exact replica of their thing. There's obviously value in like the formally verified ones. We verified that. Do you think there's value in this, like, RL environment startups that are building, like, somewhat generic but task-specific environments?

And then if none of those work, then what do we do instead of GRPO, I guess- ... is the, the question.

Kyle Corbitt27:35

Yeah. I suspect there is value in that. You know, I think the... You know, the, the folks buying those environments and training on them in the big labs would have the best knowledge on how well they work. I think they probably work okay.

I think they probably also are like... You know, and we'll see maybe with the next generation of models released, like how well they transfer. I would say so far, um, it seems like they don't train well enough. Like, if you, if you use, um, you know, OpenAI's agent interface, it's, like, okay.

Or if you use the computer use products that, that everybody's putting out, they're, like, okay, but, like, not reliable enough to, like, actually, like, let go do something interesting unsupervised in the world. And I think if the eight-- you know, if the environments they were training them in were high enough fidelity, then they would be good enough in the same way that, like, coding agents can go much further because I think that in that case, we do have environments that are much higher fidelity because it's a much simpler environment in a lot of ways.

Like, it's a code base. It's, like, maybe running a web browser. Like, it's, it's, it's much easier to capture the full realistic environment in that context.

Swyx28:36

For those who are interested, when you make a reference to RL environment startups selling to the big labs, they're selling it for a lot of money.

Kyle Corbitt28:45

Yeah.

Swyx28:46

Like at least seven figures, right? Like I, I, I don't know-

Kyle Corbitt28:49

That's my understanding, yeah

Swyx28:50

... I don't know-

Kyle Corbitt28:50

I'm not a buyer, but-

Swyx28:51

Please, please-

Alessio28:52

Yeah, yeah. Mm-hmm

Swyx28:52

... like drop data points because, like, people who are not in Silicon Valley don't know this.

Kyle Corbitt28:55

Mm-hmm.

Swyx28:55

And like, it, it's like probably the current thing in VC, which is, is RL environment startups. Um, anyway, I, I just-

Alessio29:03

A lot of them. A lot of them.

Swyx29:04

There's like-

Alessio29:04

But-

Swyx29:04

... 20 of them, apparently.

Alessio29:05

Yeah.

Kyle Corbitt29:05

Mm-hmm.

Alessio29:06

But it, it's like a small number. I know that, yeah, all the labs are buying ad hoc, but, uh, in a way it's almost like they don't even care. It's not a product. It's like they're basically, like, paying the company to build an environment ad hoc for them-

Kyle Corbitt29:18

It's a very services business at the moment

Alessio29:20

... to build six.

Swyx29:20

Services business.

Alessio29:20

Exactly.

Kyle Corbitt29:21

Mm-hmm.

Alessio29:21

But I mean, if you're spending like a billion dollar in a training run-

Swyx29:24

Yeah, but, like, you can specialize in, like, we are the one that does e-commerce. Like, we are the e-commerce experts, so come to us for e-commerce.

Alessio29:30

Yeah.

Swyx29:30

Go to the other guys for, like, social media. Go to the other guys for like, I don't know-

Alessio29:34

But I'm curious-

Swyx29:35

... whatever

Alessio29:35

... your take is, like, how do you need to get the data out to make it fit in your training run? Especially when you get to, like, these larger labs. I think they have, like, very sophisticated post-training pipelines.

Kyle Corbitt29:46

Mm-hmm.

Alessio29:46

And I don't know if there's, like, a way to just build a company where it's like you just send them a CSV of, like, data. It needs to be very integrated in it. But I'm curious what you've seen working with customers too.

Kyle Corbitt29:57

So for RL, like, the whole way this works is, is, you know, it, it has to sort of be getting feedback from the real environment. So I don't, I don't see a world where it's as simple as like, hey, you can...

You know, there's, there's like a CSV-

Alessio30:09

Yeah

Kyle Corbitt30:09

... type approach. I guess you, you could encode anything as a CSV, but- ... if you try hard enough. Um, for RL to work, you have to be looking at real runs, ideally of your actual agent in its current state across, uh, within an environment as real as possible.

So you have to, like, look at actually, um... And, and, like, the data format's, like, actually super simple. Like, it's just, like, basically a list of, you know, like, chat completion messages. Um, it's, it's effectively whatever-

Alessio30:34

Tool calls.

Kyle Corbitt30:34

Yeah, your a- exactly. Yeah. It's whatever your agent will be seeing and doing when it's running. So the, the getting the data is not hard, but what's hard is, like- When you're doing one of these runs and your agent makes a tool call, okay, now that tool call has to connect, you know.

Le- somehow it's gotta get data back from something, and that data has to look like it will look in, in real usage. So setting up that whole part of the system is, uh, is the challenge.

Swyx30:55

And then for just a reference job for more people, Web Arena is my first instance of this kind of thing where you literally have a Docker container that has, like, a clone of Reddit-

Kyle Corbitt31:04

Hmm

Swyx31:04

... a clone of Wikipedia, a clone of GitLab, clone of CMS, and a clone of a e-commerce place. And I think since then there's, like, MindToWeb maybe. Uh, I don't know if there's an other large, well-known academic, um, environments where people are basically using these as benchmarks, but probably also u- it's pretty useful for training.

Kyle Corbitt31:22

Yeah.

Swyx31:23

Uh, so, so if you wanna check out those things, yeah, you can definitely check there. I think the question for you is, as someone who bet on SFT, then you bet on RLFT, and then now you see these guys making a lot of money, why didn't you go there?

Kyle Corbitt31:36

It seems to me like that definitely is a services-heavy business at the moment as it, as it's presently constituted. I'm sure that these companies are all developing different kinds of secret sauce on, like, how to, how to do this, like, more quickly.

So that's part of it, and I, I, I don't particularly enjoy services businesses. Um, but you know, I also kind of feel like we will move towards a world where either the big labs can... Like, it's one of those businesses where, like, the only customers right now are, like, whatever, four big, maybe, maybe, maybe six big labs that, like, you know, are training these models on environments, and I don't think...

I'm a little skeptical-

Swyx32:07

Right, what's the tab?

Kyle Corbitt32:08

Yeah. Um- But you know, like, look, you, you can say the same about Scale AI and, and all of their competitors that are like, you know, many billion-dollar companies that's have basically the exact same customer set, so, so yeah.

Swyx32:20

Uh-

Kyle Corbitt32:20

It may work out.

Swyx32:21

Yeah. Unless, yeah, I don't know if you wanna do a small shameless plug for Veras.

Alessio32:25

Oh, yeah, I mean, so Veras, one of our portfolio companies, they work with the people building the agents now with the model on, like, their internal tool call loop.

Kyle Corbitt32:33

Mm-hmm.

Alessio32:33

So they can observe all the internal traces and, um, build the data to then have, like, a OpenPipe do the RFT on the thing. I think in the enterprise we've seen a lot of that, especially for chatbots. It's, like, the less sexy use case, but, like, they work with a lot of financial services company where their customers go in there and say, "What's my balance?"

Like, "When did I do this transaction?" And those are all tool calls, you know?

Kyle Corbitt32:56

Mm-hmm.

Alessio32:56

And they need a way to test and improve that behavior.

Kyle Corbitt32:58

Mm-hmm.

Alessio32:59

And the models haven't gotten that much better because these tools are, like, badly documented. They're, like, badly named. I think that's kind of, like, the, the problem with a lot of the agent builders that are not AI native companies is, like, they just put this, like, very generic tools in the thing, and then they expect it to work like magic, and these simulations kind of help them.

Also have the usual compliance things. It's like before shipping this, we tested that it doesn't give financial advice. We test that, you know, there's all these different things. So I'm curious to see how much the companies generalize, you know?

I think, like, Veras has a lot of success in, like, highly regulated environments because of different requirements. But I'm curious if you have a different way to segment the market of, like, when you think about RL, there's, like, environments that are, like, low stakes, there's, like, environment that are, like, high stakes, there's environment that have implicit rules that are made by the SEC or, uh, other government agencies, how you think about it.

Kyle Corbitt33:55

Mm-hmm. Yeah, I don't know that that segmentation is, is necessarily the most relevant. I, I, I'd have to think more about that segmentation, whether, whether it's, um, you know, that there's, like, a strong difference in how useful RL is, uh, w- across those sectors.

Where I see the segmentation is something basically just, like, capabilities-based, where it's like, hey, if I'm trying to do something that's, like, much more advanced, um, and you know, maybe, like, long horizon, then RL can probably give me a much better behavior.

And I might almost think that, like, yeah, those sort of, like, more compliance... Like, I, I feel like in those kind of environments you, you probably don't want your agent doing very much because then-

Alessio34:32

Right

Kyle Corbitt34:32

... like, you can't make any guarantees about what it might do. Um, and so, uh, you're probably not doing these long horizon things, and maybe RL is, is not gonna get you what you want. But I don't know.

Yeah, I haven't thought about it too much.

Alessio34:42

Yeah. I think, like, a lot of the customers don't necessarily end up doing RL anyway.

Kyle Corbitt34:47

Mm-hmm.

Alessio34:47

It's almost like the simulation and the environment-

Kyle Corbitt34:49

Mm

Alessio34:50

... is, like, a way for them to understand the paths-

Kyle Corbitt34:54

Mm-hmm

Alessio34:54

... that the agent can take and less about we need to then use that data to do fine-tuning.

Kyle Corbitt34:58

Mm-hmm.

Alessio34:58

But I think it's like a, it's gonna be a spectrum, you know.

Kyle Corbitt35:00

Yeah.

Swyx35:01

What replaces GRPO?

Kyle Corbitt35:03

Yeah. Uh, it's a good question.

Alessio35:04

We need the alpha.

Kyle Corbitt35:06

Yeah, I mean, I don't know is, is the short answer. I do think... This, this is, like, a fairly high salience question in the research community. I think there's a lot of folks, like, trying to figure that out.

Swyx35:15

Every paper has a variant, like, yeah.

Kyle Corbitt35:17

Um, yeah. But a lot of-- But I think, you know, the, the big question is, like, are we doing, you know, normalization based on grouping or in some other way, right? That's, that's like... I, I would say, like, I would claim we're just gonna keep calling it GRPO as long as the normalization is done within, like, a group, even though, yeah, there's a lot of things that, like, probably should get their own names.

A lot of things that have tried to get their own names and, and have failed on the marketing side.

Swyx35:38

Yeah.

Kyle Corbitt35:39

I think something that, like, doesn't require group-level normalization, which a lot of, you know, older things didn't, probably works, but I think that the older things also are really finicky, so there's, there, there may be other kinds of simplification, and I don't know exactly what, what those will be.

Alessio35:53

Where do you put the prompt optimization thing? We did a Dev Day episode, and we mentioned GEPA, and then everybody came out of the woodwork on, on Twitter-

Prompt Wars36:00

Swyx36:00

DS5 Bros

Alessio36:01

... and complained about it. Yeah, exactly.

Kyle Corbitt36:02

Yeah, no. Okay, tell me, have you or people you talked to tried GEPA? I wanna know, like, what-

Swyx36:06

I read the paper.

Kyle Corbitt36:07

Okay.

Swyx36:07

I, I'm just, like, look, like, the, the prompt layer up- updates are not the same as weights updates, which, which they're just comparing apples and oranges. And I, I, I talked with a few people I respect on, on this, on, on, on the RL side, and they, they kind of validated, like, the way that these grad students market their papers is their thing beats the current hot thing, and the current hot thing is GRPO.

But, like, I, I think they're just not that comparable.

Kyle Corbitt36:31

I disagree with that. Like, I actually think they are comparable in the sense that, like, w- it depends on for what purpose, right? But, like, if I'm a company and trying to, like, get the best performance out of my agent, like, I don't care if you're changing my prompt or if you're changing my weights.

If you get better performance on my agent-

Alessio36:45

Exactly

Kyle Corbitt36:45

... you know, I'm, I'm, I'm happy. On that front, I do think they're comparable, and we've evaluated-- I mean, we evaluated, like, GRPO too.

Swyx36:51

Like, the, um... So, so their answer was, "You are gonna do both. If you really want max performance, you're gonna do both."

Kyle Corbitt36:56

Yeah. We evaluated everything from DSPE, and we, we evaluated GEPA as well, and it's like it just doesn't work. Uh, o-okay, like, okay, now that's gonna be the-

Swyx37:05

Fighting words

Kyle Corbitt37:06

... the quote. Yeah.

Swyx37:06

GEPA doesn't work.

Kyle Corbitt37:08

It didn't work on the problems we tried it on. It just didn't. It got like a minor boost over the sort of like more naive prompt we had and was just like... It, it was like, okay, just kind of like our naive prompt with our model gets maybe like 50% on this benchmark, and like GEPA got to 56, and we do RL and we get to like 96.

I mean, it was just like not even-

Swyx37:26

Yeah

Kyle Corbitt37:26

... comparable. And so maybe we were holding it wrong-

Swyx37:29

Well, well, you see, no, so, so y- both sides are claiming skill issue, right? So what they would say is you probably used it wrong. And then with-

Kyle Corbitt37:34

That's right. Yeah

Swyx37:35

... RL people are saying that probably the GEPA, GEPA guys when they, when they set up the j- the GRPO benchmark they, they... it wasn't a very fair comparison, which is exactly what, uh, my source said. It's hard to tell, you know, everyone has, everyone has, uh, is trying to get to some version of the truth.

Kyle Corbitt37:47

Yeah.

Swyx37:48

Uh-

Kyle Corbitt37:48

But I will... I... What I will say is like we, we want it. I mean, I don't know if I would say, go so far as to say we want it to work, but we certainly want to know if it works.

Like, that's like actually very relevant to like the product we're building, honestly.

Swyx37:57

Yeah, and if it's more efficient to get there-

Kyle Corbitt37:59

Yeah

Swyx38:00

... uh, then you should be able to go do it

Kyle Corbitt38:01

... then we just haven't been able to get it working. That's... Yeah.

Swyx38:02

It's actually kind of more credible now, now that you're like, you know, you, you, you're part of a larger CoreWeave that you're not... Obviously, 'cause I think GEPA maybe is, uh, makes, uh, OpenPipe like less relevant.

Kyle Corbitt38:15

I, I totally would disagree with that-

Swyx38:16

Okay

Kyle Corbitt38:17

... because like the level we are- see ourselves operating at is actually we're not like RL bros trying to figure out like the use case for all RL. We're like, "Hey, we're working with all these enterprises. We have all these big companies we're talking to, and we're trying to figure out like how we make their stuff work better."

And so like I personally am very motivated like if something like GEPA works, like, okay, let's, let's build a product around that. That, that's how, that's how I think about OpenPipe at least.

Swyx38:38

No, I mean that, that's, that's a good clarification to make. Even more so you actually took a sincere look at it and you concluded that there was nothing to do, to do, to nothing to build.

Kyle Corbitt38:46

Well, well, you know, maybe we were holding it wrong.

Swyx38:47

So we had Shenyu on the podcast a while ago and like, I mean, uh, he's been a proponent of automatic prompt optimization and, uh, this idea that like you can do a lot more in the prompts than you can do in the weights.

And, uh, in principle I'm biased inclined to believe that something like a DSPEI, something like a, a GEPA works. Uh, so I'm very surprised to, to hear this.

Kyle Corbitt39:08

Yeah. Like we keep trying it, you know?

Swyx39:10

Yeah.

Kyle Corbitt39:10

Um, we tried the MIPRO V2 stuff that was hyped before that.

Swyx39:13

Yeah. Also, okay, I should not bury the lead on the best argument for this, which is it basically GEPA models how the big labs do their system prompts.

Kyle Corbitt39:21

Mm.

Swyx39:21

It's, uh, genetic evolution, you know, and, and they, and they just sort of incrementally, uh, evolve based on like the overall evals that they have. It's slow because it's done by humans, but GEPA theoretically improves it. I mean, it automates this.

Kyle Corbitt39:36

Okay, hold on. Is the claim of the big labs have something... Uh, this is, this is news to me. This is interesting.

Swyx39:40

No, no, no, no, no. This is philosophically the same.

Kyle Corbitt39:42

Oh, okay.

Swyx39:42

I'm not saying like-

Kyle Corbitt39:43

Oh, sure. But like you're injecting a whole lot of human intuition and kind of like-

Swyx39:48

Yes

Kyle Corbitt39:48

... potentially out-of-band information.

Swyx39:49

We have the best model in the world, which is humanity-

Kyle Corbitt39:51

Yeah

Swyx39:51

... or like smart humans.

Kyle Corbitt39:52

Yeah.

Swyx39:53

Uh, and now we're doing GEPA using dumb LMS.

Kyle Corbitt39:56

Right. But they're also like the humans can bring in out-of-band information that like maybe is not captured in the actual like, you know, the eval. Like they can be like, "Oh, yes, technically this did well on the eval," but it's like not really...

You know, like I, I would suspect that a lot of that ends up getting injected through that human being in the loop.

Swyx40:11

Yeah, yeah. I've always been very surprised at how these guys work on their system prompts, which are tens of thousands of words long, and there's no ablations. They just kind of pick what seems to work and then chuck it in there.

And that is the Claude system prompt.

Kyle Corbitt40:28

Yeah. Can't argue with success.

Alessio40:32

Is GPT-5 the first model that had a prompt optimizer by one of the large labs? I believe so, but I don't remember.

Swyx40:38

Uh, Claude Workbench had this like a year and a half ago, if you see it that way. It just wasn't like fully automated, but it was extremely good for its time. I kept telling people about it and nobody believed me.

Kyle Corbitt40:48

Do we know if they used it internally?

Swyx40:50

Claude Workbench? Yeah.

Kyle Corbitt40:51

Okay.

Swyx40:52

Why, why not?

Kyle Corbitt40:53

Oh, I don't know. Like I-

Swyx40:54

Yeah

Kyle Corbitt40:54

... I just, my experience, you know, knowing a lot of people at these labs is like they launch a lot of products because-

Swyx40:58

Sure

Kyle Corbitt40:58

... like some team is super excited about this product, but that-

Swyx41:01

Yeah

Kyle Corbitt41:01

... I, I wouldn't put that much weight on it just because they launched it.

Swyx41:04

For some measure of use internally, I, I, I am sure I, I'm... talk- the guy- people I talk to are biased. I don't know if you fully explored that, that thread.

Alessio41:11

Yeah, no, I think that's, uh... It's just interesting that now it's been acknowledged that like the LLM can improve your prompt. And so I think like GEPA and always also riding this wave of like, okay, maybe we can do this programmatically.

Swyx41:23

Yeah.

Alessio41:23

But I also think the long tail of people just prompts really badly.

Kyle Corbitt41:26

Hm.

Alessio41:26

And so I think there's some value there versus once you go into RL you already have like a more sophisticated audience, you know?

Swyx41:33

Yeah, yeah.

Alessio41:33

Like who gets to do GRPO? People that are really smart. Who gets to do prompt optimization? Like everybody's trying to do it. So-

Kyle Corbitt41:39

Yeah, that's fair. Maybe, maybe our baseline was, was too high to start with.

Alessio41:41

I know your, your naive prompt is probably like, you know, top 10 percentile of prompts that people put in these LLMs, you know? So.

Kyle Corbitt41:47

Sure. I'll, I'll, I'll take it. Yeah.

Alessio41:49

Yeah. Yeah.

Swyx41:50

And then the other thing that comes to mind as you were talking about things, injecting things out of band and all that, I think it's a, there's a broader tr- trend that I'm tracking for Worlds [26] , which is the move to online evals.

Um, the, the, the way that we do evals today is probably too locked down. You're kind of fighting the war that you already know should be fought, and you're not fighting the wars that you don't know about 'cause you didn't, you didn't plan for it.

Whatever. How can we sort of move more online evals into our GEPA process? And maybe that's, that's part of it.

Kyle Corbitt42:18

That part I'm much more bullish on. And, and we can make the analogy, like we can, we can pull in kind of like RL intuition here, which is if you're doing GEPA on a sort of static data set of like, oh, this is the input, this is like what makes a bad, good or bad output, then like as you're updating your prompt, like your information, the tr- data you're training on becomes less useful, right?

'Cause it's generated by, you know, 'cause, 'cause it's based on kind of like the problems you were running into before. And that's the same problem you have with, with RL where, where you have this concept of being off policy, where it's like as you're doing training you really wanna be training on rollouts that came from the latest version of your model 'cause if you train on some that came from further back then it's like it's sort of stale data and it's like not-- it's no longer representing the current issues with your model.

And so if you try and correct for the issues that existed back then, it, it may not actually be helping you that much. And I think, you know, for either RL or prompt optimization, that's definitely true. I think that, like, one way to apply that in practice is exactly what you're saying, where you're using the actual data from your, your real evals.

You have some way of saying like, "Hey," either people flagging these or an LLM flagging these, or some way of saying like, "This was a good or bad output." I totally agree with you that, like, if you're bringing that into your process, I'm, like, much more optimistic that you're gonna get good results.

Swyx43:28

Yeah. And the pipelines are not set up.

Kyle Corbitt43:30

Yeah.

Swyx43:30

Like, this is like analytics and UX people like trying to... being drawn into the ML process, which they've never been done before. If I had to make a bet as a big theme for next year, this is gonna be it.

Kyle Corbitt43:40

No, I, I, I agree. And I, I mean, I think that like all of the sort of observability people, like platforms see that and like are trying to figure out what the right shape is. I haven't seen the right shape yet, but yes, it, I, it seems like a, like a, a theme for next year.

Swyx43:56

StatSig.

Kyle Corbitt43:57

Maybe. Yeah. I haven't, I haven't used them, but OpenAI seems to like them.

Swyx44:01

Yeah, I mean, like, uh, I do think like buying, you know, an experimentation platform makes sense and like, you know, I think it's sort of... Like I've said before on the podcast, I think that I'm very bullish on model routing as a feature, but less bullish on model routing companies-

Kyle Corbitt44:16

Mm-hmm

Swyx44:16

... because, uh, of exactly stuff like this where, like, it is just gonna get, get absorbed into the model. It's, it's a very big part of building and process. You probably don't wanna... A- and it's not that hard.

Like it's it's not rocket science. It is ju- it's, it... You're just like connecting pipes and making sure things are set up so that it's easy to use that data.

Model Debate44:33

Kyle Corbitt44:33

I have a question for you, a general question.

Swyx44:35

Mm-hmm.

Kyle Corbitt44:36

So what fraction of tokens generated by, say, like the end of 2026, do you think are gonna come from open source models versus proprietary models?

Swyx44:44

Oh. Oh, oh, that's a fun question. So we have an answer from Ankur from Francus, where he was like, "It's 5% and going down."

Kyle Corbitt44:52

Mm-hmm.

Swyx44:53

I think it's going to go up because of the amount of enterprise adoption of open models that I'm seeing. And also-

Kyle Corbitt45:02

'Cause there's a lot of demand. Like there's... The enterprises would much rather be on open models-

Swyx45:07

Of course

Kyle Corbitt45:07

... if they actually could get the performance-

Swyx45:08

Yeah

Kyle Corbitt45:08

... they're looking for.

Swyx45:09

Yeah, for cost, for privacy, all that stuff. A- and I think like ba- basically honestly, it's just literally like we may have hit quote, unquote AGI in a sense of like it is the, the average LLM ca- is capable of the work of the average human, not the best human, but the average human, sure.

Like it's actually pretty decent at customer service. Like it's um, and it's actually pretty decent in like, I don't know, transcribing things from PDFs, whatever. So like, yeah, I mean, uh, totally I think, I think that should rise, but people who believe that it should rise to like 50% are out of their minds.

Kyle Corbitt45:43

Mm-hmm.

Alessio45:43

And I think it's a trick question. We should take coding out. I think once you take coding out, I, I think, yeah, it can be like 15, 20%. But I think with coding it's still gonna be very low-

Kyle Corbitt45:53

Yeah

Alessio45:53

... because like these max plans are like so subsidized and so many tokens are being generated. Like Anthropic is like, you know-

Kyle Corbitt46:00

So-

Alessio46:00

... 50% of their revenue is like-

Kyle Corbitt46:01

So it's your claim, it's your claim that it'll, it'll mostly be, you know, that coding will mostly be closed models because the tokens are subsidized or because the models are just so much better that-

Alessio46:10

I think as long as-

Kyle Corbitt46:11

... people use anyway

Alessio46:11

... I mean, I'm paying 200 bucks a month, and it's like I'm spending thousands of dollars.

Kyle Corbitt46:16

Sure.

Alessio46:16

Like by accident, by accident I pay with like my credit card and I spend like 100 bucks in like an hour and it's like-

Swyx46:22

By the way this-

Alessio46:22

... you think about, "Hmm"

Swyx46:22

... this is like the, the thing that nobody wants to talk about for Anthropic. Like Anthropic went from like 1 billion in revenue to 5 billion and everyone's like, "Ooh, hoo, yay." And then like, "What's the margins?" You have this like goose meme going like, "What's the margins?"

Alessio46:32

Right.

Swyx46:33

Um-

Kyle Corbitt46:34

Yeah

Swyx46:34

... they say it's like 6%. There-- You are part of the 6% that is abusing everything- ... so everyone else-

Alessio46:39

I'm not abusing. I'm just-

Swyx46:40

You are the lost leader

Alessio46:41

... I'm just-- It's not like I'm rotating accounts.

Swyx46:43

Yeah.

Alessio46:43

I'm just using-

Swyx46:44

Yeah, yeah, you're using the product

Alessio46:45

... the one that I paid for.

Kyle Corbitt46:45

You're using the product. Yeah.

Alessio46:45

You know, it's like-

Swyx46:46

Yeah, yeah

Alessio46:46

... I could-

Swyx46:46

But like through you, people like hear about Cloud Code, they pay the $200 a month, and then they don't use it, and they, they s- they pay for your inference.

Alessio46:52

Yeah. Thank you. Thank you, everyone. Keep doing it.

Swyx46:54

Right? So-

Alessio46:55

I don't want that to go away.

Swyx46:56

Uh-

Alessio46:56

But I think like I don't really see... It's hard to see a world in which QwantaCoder or whatever model replaces that becau- between quality and cost. It's like to make-- to generate this amount of tokens for 200 bucks a month, I don't know how anybody can like offer like Together, Fireworks.

They cannot really offer it at that price, and the quality is not as good.

Kyle Corbitt47:18

But the reason they can't offer at that price is, is 'cause of the subsidies-

Alessio47:21

Subsidies

Kyle Corbitt47:21

... right? Which, which is-

Alessio47:22

Yeah, exactly

Kyle Corbitt47:22

... not like the long-term like sustainable dynamic.

Alessio47:25

Well, I, I, I mean, it's interesting because so both Anthropic and OpenAI are building their own infra, right? And like, they're gonna get to a place where they're gonna have idle GPUs that they own.

Kyle Corbitt47:35

Mm-hmm.

Alessio47:35

And so they will also be incentivized to have 100% utilization, and so, you know, they will subsidize some of it.

Kyle Corbitt47:41

Mm-hmm.

Alessio47:41

Just the, the same way, you know, if you go on SM Comp- SF Compute, like you pay a buck 40 for like an H100 instead of like the 220 listed price on AWS. Um, so I think it will continue, but again, it depends on whether or not they actually have the 500 billion like they were saying, which I think they do.

You know, just to be clear-

Kyle Corbitt47:59

Okay

Alessio47:59

... I think Stargate will go online.

Kyle Corbitt48:01

That's fair.

Alessio48:01

But once it goes online, then it's like, well-

Kyle Corbitt48:03

If they figure out how to pay for $500 billion worth of compute, then, then they probably can subsidize-

Alessio48:07

Oh

Kyle Corbitt48:07

... for a while.

Swyx48:07

I think they have the 500 B. They're going bigger. Isn't it obvious?

Kyle Corbitt48:12

Ha- What, what, what do you mean by have? Like a-

Swyx48:14

At the start of this year when they announced Stargate, people were like, "Oh, you don't even have 10." Like Elon was like, "You don't even have 10."

Kyle Corbitt48:20

Hmm.

Swyx48:21

Whatever. And then Satya's like, "I'm good for my 80." But like now, now we're seeing all the money start coming in, and like probably it's in the order of like 200, 300 billion like that you could probably get raised and, and committed, and they're gonna get the rest.

Like it's, it's fine. Like I think the, the plan is actually a lot bigger.

Kyle Corbitt48:38

I just-- Uh, can I just say I love this industry? It's like, yeah, they've got like 2 or 300 billion, and like what's another couple hundred billion?

Swyx48:44

The, the-- Like-

Kyle Corbitt48:44

There's no other industry in the history of the world where you could say-

Swyx48:47

Yeah, yeah

Kyle Corbitt48:47

... what's a-

Swyx48:47

It is, it is stupid, but like also like do you doubt it? Like I, I don't. I like-

Kyle Corbitt48:52

Yeah.

Alessio48:52

Yeah.

Kyle Corbitt48:53

That's fair. Yeah.

Swyx48:53

No, like the... I, I, I literally like after last week, I think maybe two weeks ago with the whole Oracle, Nvidia, and then even AMD deal, I'm like, oh, like these guys not only... They've locked down Stargate 1.

They're working on Stargate 2, whatever that, uh, that is. And, and like the sheer ambition is like freaking crazy. There is still one more shoe to drop, which is- The non-sovereign wealth funding that OpenAI needs to get, which they've promised to drop by the end of this year, and my money is on they have to do a coin.

Like it's... I'm not a crypto guy at all, but like, you know, this is-

Kyle Corbitt49:28

You think it's gonna be like an OpenAI coin

Swyx49:30

... th- this is the one AI founder that has his own coin already.

Kyle Corbitt49:32

Yeah. Hmm.

Swyx49:33

And like he needs more money, and he said that they will come up with new innovative financing meth- methods.

Kyle Corbitt49:38

I-

Swyx49:38

What else is there?

Kyle Corbitt49:39

Yeah, I mean-

Swyx49:40

They, they are already in the token-selling business, like-

Kyle Corbitt49:42

But you gotta That's a great line. Uh, but anyway-

Swyx49:45

Like buy an OpenAI token, it translates to a, a GPT-5 token. Like, you sure?

Kyle Corbitt49:50

Hmm.

Alessio49:50

It's a stable coin.

Kyle Corbitt49:54

Hmm. You'd, you'd have to- you'd have to get- you'd have to get a lot of political buy-in, I think, to, to take that level of risk.

Swyx50:00

What? Th- the White House that is most crypto friendly since- ... the dawn of time?

Kyle Corbitt50:03

Yeah. Well, and I guess like Elon's out of there now, so maybe they can get the... make a- make, make, make the friends.

Swyx50:08

Right.

Kyle Corbitt50:08

Yeah.

Swyx50:08

I, I think it's doable. We'll see.

Kyle Corbitt50:09

Yeah.

Swyx50:09

You know, like, uh, who knows? Uh, yeah, I, I... for what it's worth, I've, uh... nobody's, like this is a- this is a me theory. I, I don't have- ... any inside infor- information.

Kyle Corbitt50:18

Um-

Swyx50:18

Uh, yeah, should we go back to RULER?

RULER & Worlds50:19

Kyle Corbitt50:19

Yeah, sorry. Right. OpenPipe. Anyways, we were saying-

Alessio50:23

I think this story takes us to July '25, when you released RULER, which you got easy mode for RL rewards, and then, I mean, shortly after you get acquired in September, so maybe you just wanna talk through the summer.

You know, what was the vision, then maybe-

Kyle Corbitt50:36

Yeah

Alessio50:36

... how the acquisition came together.

Kyle Corbitt50:38

Yeah, absolutely. So, you know, I mentioned my, my initial like opinion of like how likely this, this direction was to work was maybe 25%. We're up to, you know, 55% or so, and, and RULER is actually a big update on that got me from the 25 to the 50.

So, so let me, you know, I guess just for context there. So basically, there are several problems you have to solve if you wanna use RL successfully. The problems you have to solve, I mean, some of them are just like really dumb, basic, like, hey, you gotta get the infra and like the libraries have all really sucked and been built by, you know, PhD students who don't know like how to build reliable software.

So, so like there's, there's like all these like practical issues that, that we're working through. So that's one thing, and that's, that's kinda what we're trying to solve with ART. But even after you've got that solved, you've got like major issues, which is like you, you gotta know if your- if your agent is actually...

or, you know, whatever system you're using on RL is doing a good job, right? That's, that's fundamental. You have to have a reward. You have to know it's doing well or, or poorly. Sometimes that's easy to do. If you're solving like a math problem or something, you can come up with a data set of math problems and the known solution and check if it's the same.

The- on the coding side, there's been a lot of like innovative work around... I mean, there's first of all like a lot of open data and, and a lot of, you know, the- there's like, I think the approach a lot of companies take is you, you find existing test cases and then you break them, but there's sort of like a, a way to figure out if, you know, you, you can run the test case, right?

And see if, if your code fixes it or not. In a lot of other domains, it's like much more murky. It's like what is a good job versus a bad job? How do I know if I did a good job?

And you really need that information. So we've tried a bunch of different things. RULER is a library that we released, um-

Swyx52:06

Which let me, let me... Relative universal LLM elicited rewards.

Kyle Corbitt52:10

Thank you. Yes. And the way it works is basically this depends on the sort of GRPO insight, which I was mentioning earlier, that you actually don't inor- with GRPO it has this nice property where you don't have to have like an absolute judge of the truth.

You just have to judge relatively. And so simplifying it a lot, it's basically just LMS judge on a whole group. So you say, "Okay, this is the task I'm trying to a- achieve. Here's four different runs of an agent trying to achieve it.

Which of these did best?" And it, it stack ranks them. And it turns out that works phenomenally well with GRPO. Like way better than I expected, way better than, you know, anyone who kind of like I talked to before we actually tried this expected because it sort of in, in, in the, the LLM using as judge, it, it can sort of like self ground because it's, it's just getting these relative ranks, right?

So it doesn't have to like have like a, a omniscient view of like what good or bad looks like. So that has worked at basically everything we threw it at. Um, we've done it with a bunch of client projects.

We've done it with a bunch of our own customers. It basically just works. Like it's basic- Like I, I honestly kind of feel like the reward assignment problem is like fairly solved.

Swyx53:12

Mm-hmm.

Kyle Corbitt53:13

Yeah. Which it is, it's fantastic.

Swyx53:14

Just any LMS judge off, off the hook. Like you-

Kyle Corbitt53:17

We've tried it with so many things. Like, so one of, one of the results we published was we used Qwen 2.5 14B as the model we're training, and as the judge we used Qwen 2.5 32B, which is like not...

I mean, it's fine, but it's like not a... It's, it's much worse than any frontier model, right? And even with that combination, we were able to get our, our agent doing like state-of-the-art better than any frontier model on, on the tasks we tried it on, even with like an, an extremely weak judge model.

So it, it really doesn't depend on having like a really great judge model, um, in practice. So yeah, it's, it's just like, it's just not something we've had to worry about since then at all. So that's kind of like checked off.

So that's sort of like got me like a significant increase in like, okay, this is actually something people can apply. This is now something that's packaged up. People can just use our... It's a- we open sourced everything. You can use it off the shelf.

If you stick it in your training run, it will probably just work. So that leaves the remaining problem, which we were, I guess we were talking about them out of order, but like that remain- leaves the, the environment problem, right?

And that's like the one big remaining piece that like we don't know yet how to automate or remove and requires a lot of manual work for, for every single task.

Swyx54:18

For listeners, you know, this is why I kind of refer to it as self-supervised because it, it is like removes, uh, more and more of the human judgment and like the history of machine learning all the way from like, I guess the, for the, the, the, the start of like, uh, uh, ImageNet and everything, uh, it is, is really like that, that insight of like you should just take humans increasingly out of it and scale up the data.

You can just throw in there with no supervision.

Kyle Corbitt54:42

Yeah. Yeah. Totally.

Swyx54:43

Yeah. It's, it's really awesome. Are you bullish on, um, dedicated LLM as judge models? Have you looked at those?

Kyle Corbitt54:50

I don't know.

Swyx54:50

Uh, Bespoke Labs, we did an episode with them, and they're, they're really trying to carve a niche in there.

Kyle Corbitt54:54

We've looked into it. We, we've trained some ourselves. We've also like used some off the shelf. There's, there's, uh, there's an evaluation benchmark that the AI2 people, uh, put together and a RewardBench. Um, and so RewardBench is kind of like trying to benchmark models on, on serving as LLM as judge-

Swyx55:08

And reward models or LLM as a judge is in your mind is same, same thing?

Kyle Corbitt55:11

Yeah, yeah, yeah. Um, they have-

Swyx55:13

Yeah. Mildly different.

Kyle Corbitt55:15

Okay.

Swyx55:15

Depends on the task. Like LLM as judged is, is, is usually more sort of product facing and reward is, reward modeling is much more specific within like a chat task.

Kyle Corbitt55:24

Um-

Swyx55:25

Which is that, that used to be the old meaning of reward model

Kyle Corbitt55:28

I don't know. Maybe terminology has changed. Like I, I think, I think they're, they're pretty equivalent. Um-

Swyx55:32

I, I understand that, yeah.

Kyle Corbitt55:33

Mm-hmm.

Swyx55:33

I can, I can see your side.

Kyle Corbitt55:34

Anyway, so, so yeah, RewardBench is, is kind of like... And so we've tried a bunch of off that. Um, the thing is, like I guess my, my maybe meta take on this is any task that is extremely common is gonna end up in like as a specific like part of the training data for the frontier labs, and LMS Judge is just something everybody's doing in so many different contexts that you have to assume that all the frontier labs have a bunch of like LMS Judge style tasks that they're training their models on.

And I do believe that if something does kind of like make it in, in a like more than minor way into their training data, that like they're gonna do at least as good a job as, as a dedicated model.

So I don't think there's probably a lot of alpha in dedicated LMS Judges just because it's something that like the... Uh, let me caveat that and say like if you've got like a very, very specific task that's like weird and has weird requirements and you have a lot of data on what's good or bad, then like training a reward model for your specific task I think could still work.

Um, or you know, fine-tuning an LMS Judge on your specific task could work. I'm pretty bearish on like a, "Hey, this is a model that is trained as an LMS Judge, but it's a generic LMS Judge, it can be used to judge anything."

I, I just don't think you're gonna beat the frontier labs on that.

Swyx56:45

Yeah. One other version of this that is not quite an LLM, but some people are thinking about it, is something that we're working on for a future episode, which is world models.

Kyle Corbitt56:53

Hmm.

Swyx56:53

Um, and uh-

Kyle Corbitt56:55

Sexy

Swyx56:55

... ve- yeah, very sexy. First applied in video as far as I can tell for Genie 3, uh, Genie 1, 2, 3, and then, and now with code-

Kyle Corbitt57:03

Mm-hmm

Swyx57:03

... and potentially with virtual cells for, for sy- uh, for AI bio. Any exploration there that, that's interesting to you?

Kyle Corbitt57:10

Yeah. Um, so we've been playing around with it a little bit. It's one of the directions that I'm like fairly optimistic on for solving the environment problem specifically. Because if you think about it, like, like a world model, it's, it's a simulated environment.

That's like what it, its whole purpose, right? So if you get one that's-

Swyx57:26

Yeah, but in an, in an LLM like thing, not like a Docker.

Kyle Corbitt57:29

Uh, yes. Yeah, yeah, yeah. So, so it's, it's like, you know, whatever, hallucinating, generating, imagining-

Swyx57:34

Mm-hmm

Kyle Corbitt57:34

... the responses you'll get from the world. So you can imagine, right, if you had like a really, really great world model that you were training on, yeah, it's like your, your agent that you're using, it would go out and make some tool call, and then this world model model would generate, "Hey, this is like probably what the tool call..."

And if, if you have a smart enough, strong enough one, then it could keep its own, you know, effective internal state of like the changes you made so far and how that affects. So we've played around with it some.

You know, I think if we can get it to work really well, then that could be a solution for the environment problem, where you just take a bunch of production traces and use those to condition your world model so it understands your specific system and what its failure modes are, and then train against that world model.

Um, and uh, and it works, and you know, and, and the resultant, the, you know, agent that you train with that would, would then be able to perform in your real environment. So I, I do think it's like a, like a really interesting area of research.

Swyx58:23

Yeah. And did you see the meta cold world, cold world model, um, work?

Kyle Corbitt58:27

I don't think I saw that one.

Swyx58:28

Okay. Yeah, it was like two weeks ago. Uh, we, we've just confirmed that the guy for, uh, AIE code in, in, uh, in November, and it's, it's really interesting. Like the world model is, uh-

Kyle Corbitt58:37

Oh, sorry. The... You're talking about the meta one?

Swyx58:39

Yeah.

Kyle Corbitt58:39

Oh, okay. I missed it. Yes, I did, I did. I saw that one.

Swyx58:41

I, I said a lot of syllables, so it may-

Kyle Corbitt58:43

Mm-hmm

Swyx58:43

... may not have parsed. But like, yeah, it's literally like having a debugger as the environment, as the world model and, and let, uh, opening up the execution trace to the model to see what's going on and see the state and track the state as the code executes.

Seems to be smart and, you know, exploits the unique situation of code environments where we can actually do these things.

Kyle Corbitt59:02

Mm-hmm. Yeah, I think the way they envision that model being used is a little different. Like I think they're, they're, they're try- It's-- Actually, I'm curious. I'll have to see the talk. Um, but my understanding from that paper is like the goal they're imagining is this is almost sort of like a pre-training step, and then now that this model understands code really, really well, we can then use it as basically like a code generation, um, or a coding agent of some kind.

Swyx59:24

Yeah.

Kyle Corbitt59:24

Okay. Yeah. Which, which I think makes sense. That's almost more like a different kind of pre-training, I would say. Um, the way I'm interested in applying world models is as not... It is basically as its own end, right?

Where it's like actually the goal is to come out of this with something that simulates the world. Which is not something you really need in code at all 'cause it's so easy to like run code and you don't need to model what will happen if you execute this code typically 'cause you can just execute the code and, and see what happens-

Swyx59:46

Right

Kyle Corbitt59:46

... for, for training purposes.

Swyx59:48

But it closely models how we think about code when we code-

Kyle Corbitt59:51

Yeah

Swyx59:51

... is we kind of mentally execute-

Kyle Corbitt59:53

Mm-hmm

Swyx59:53

... the model as we type, and then we go like, "Is that what we really want?" Yeah, I don't know. Anyway, it's the first model that Meta's released since the MSL reorganization. Uh, we know, you know, just based on our context that they're very, very, very interested in code models as a path to AGI, which I'm, I'm also of course, very interested in.

Kyle Corbitt1:00:11

Yeah. Totally.

Alessio1:00:12

Um, I know we kept in here for a while. Let's wrap up on the acquisition. So a lot of people say, you know, companies are not sold, they're bought. What was that process like for you? Did it just happen?

Acquired & Vision1:00:12

Alessio1:00:22

Like what was the behind the scenes?

Kyle Corbitt1:00:24

Yeah. So that was driven by actually mostly the Weights & Biases founding team.

Alessio1:00:28

Lucas?

Kyle Corbitt1:00:29

Yeah. So, so yeah, Lucas and Sean, um, uh, particularly. So they, uh, you know, had recently been acquired by CoreWeave, and CoreWeave was looking to, you know, continue, um, growing up the stack. And so yeah, they, they, they approached me and were like, "Hey, you know, like no pressure, but like this is like an area that we think is really promising and we, you know, would you like to work here?"

And so that's how the conversation started. It was, uh, like long, it was pretty painful. Um, there, there were, there were points, uh, as, as late as, you know, like the week before we actually signed where it was like unclear if it was actually going to happen.

So that part was super painful. However, we've been there a month now. We, we just shipped a product yesterday, which I'm super excited about. It's been fantastic working there so far. Like I was like very concerned. I was like, okay, yes, this is great.

We make, make a lot of money by selling our company, but like is the work environment gonna like really, really suck?

Alessio1:01:15

Right.

Kyle Corbitt1:01:15

And I was like, well, I guess that's just a risk we'll have to take. It's been fantastic. Like it's, it's honestly been way, way better than I could have imagined.

Alessio1:01:21

Do, do you go down to the office? The, the one down here?

Kyle Corbitt1:01:24

I was there today.

Alessio1:01:25

Yeah.

Kyle Corbitt1:01:25

We work for-

Alessio1:01:25

But you're Seattle

Kyle Corbitt1:01:25

... so I'm based in Seattle, so the, the-- and they have a ... small office up there that we work from

Swyx1:01:29

The Weights & Biases office in San Francisco is fantastic. If you have the chance, go visit. They do, do a lot of hackathons and co-working things.

Kyle Corbitt1:01:35

Yeah, there's a hackathon going on in a month or so. I'm sure you can come.

Swyx1:01:37

Every week there's a hackathon. Uh, but yeah. I mean, so, so do you consider yourself working for Weights & Biases or CoreWeave, or both at this point?

Kyle Corbitt1:01:47

And OpenPipe too, no. No, yeah, it's, uh- We, so we, so we- I, I report to the Weights & Biases, like-

Swyx1:01:54

Yeah

Kyle Corbitt1:01:54

... founders, so we're within that organiza- uh, in, in the org chart we're there. I don't know, like f- branding wise, they're trying to say everything kind of fa- that, that's not being sel- sold to, like, big labs is kind of Weights & Biases.

Swyx1:02:07

Oh.

Kyle Corbitt1:02:07

So, like, our stuff we're launching is Weights & Biases branded.

Swyx1:02:10

Yeah.

Kyle Corbitt1:02:10

It's not, um, yeah, not, not CoreWeave branded as much. I don't know. It's still, like, s- they're still figuring it out.

Swyx1:02:16

And what's the product you launched?

Kyle Corbitt1:02:17

We launched serverless reinforcement learning. Basically, it lets you offload all of the GPU management. You don't, like, you don't have to worry about crashes and out of memories and, like, you know, uh, uh, scaling up and down. Um, we handle all that for you, and you just, like, define your environment, you define your reward function, and then you just, like every time you run a step, you kind of, like, ship back to our back end, "Hey, these are the trajectories, these are the rewards.

Now update my model." And we just, like, make it work for you. It makes it way easier.

Swyx1:02:43

Yeah. Okay, very Thinky-like.

Kyle Corbitt1:02:46

It is very Thinky-like. I, I love the Thinking Machines launch. I, I think they have a really good idea. It's also very validating for what we're doing.

Swyx1:02:52

How did this take so long to appear? Like, it seems like-

Kyle Corbitt1:02:55

I don't know. Yeah, we were- ... it's... But that's... I felt this way about everything. Like, there's so many things that should exist, like, clearly. I just think there's, like, still not enough people, like smart people working in this space.

Like, honestly, we need... Like, I realize that there's, like, you know, like, a, a lot of people, it just feels like there's still a lot of, a lot of low-hanging fruit nobody's doing.

Swyx1:03:12

Okay. One thing I saw from your, your post was, uh, your North Star as the RL team at CoreWeave is to build an oval where every agent learns continually from its real-world experience. So you're touching on the hot topic of the moment, continual learning.

What else do we need to get there?

Kyle Corbitt1:03:27

I super believe that, and, like, that's basically the vision where I'm like, you know, I, I keep talking about these percentages, 25, 50. Like, if we get to the world where we build that, um, then I think it's just like the advantages are huge, they're clear.

Everyone should just deploy their, their agents that way. Um, we wanna be, like, the team that builds the, the, the software, um, that makes that easy to do. So I talk to a lot of engineers at our customers, and they're trying to deploy agents, and it's so easy to get the initial prototype and, like, something that, like, kinda works well.

It is so hard to get from that to something that, like, you are confident is reliable enough to actually deploy in production. And when you actually look at what those failure modes look like, it's like, "Oh, yeah, like, we know if it gets in this situation or if it gets, like, these kind of, like inputs, like, it behaves funnily," but then it's like, yeah, you can update your prompt to, to address that, but, like, that's not scalable 'cause at a certain point, it's like gonna start breaking other things.

You know, you don't know what it's breaking. You really want some way to just, like, say, "Okay, look, this thing you did there, that was the wrong thing. Just, like, adjust this behavior when you get in this, and then, you know, otherwise carry on," right?

And that's what we can do with RL, and that's what we can do with, with continual learning is, like, we don't have to, like, have this concept of, like, oh, upfront I'm, like, trying to make the perfect model that solves everything.

It's like, I'm trying to make a model that's good enough, I can deploy it in production, and then when these errors come in, I'm going to say, oh, you know exactly the... I mean, very analogous to how you train a human employee.

Like, be like, "Oh, no, actually, that's not what you should do in that situation. All right, fix that and carry on." And that's just gonna make this whole process so much easier, and I think that, you know, like, I think that there is today, like, 10 times as much AI inference that could exist than is existing right now, just purely with projects that are, like, sitting in the proof of concept stage and have not been deployed.

Because there's, like, huge bucket of those, and it's, it's all about this kinda, like, reliability issue where it's like, okay, like, it, it works in controlled circumstances. There's other areas where it doesn't work. And so if we can solve this problem, there's the, that, like, 90% of the, like, inference market, like addressable market today that's just gonna, like, come online because we've solved that problem.

So, um, that's what we wanna do. Um, I'm super excited about it, and, like, I think we have very concrete ideas on, like, the specific pieces we need to make that work, uh, and we just have to execute against them.

Swyx1:05:30

Do you feel like the online RL is more susceptible to, like, the reward hacking, especially as you're, like, shortening this loop and, like, you don't spend as much time, like, looking at the different checkpoints?

Kyle Corbitt1:05:40

I'm not that worried about it. Uh, and the reason why is because it is... Reward hacking is quite easy to detect once it starts happening because once the model's found some hack, it just starts, like, doing it all the time.

It's like, "Oh, yes, this worked great. I'm just gonna keep doing it." And so you, you, you, like, notice very quickly, whoa, it's doing this thing. And assuming you're using, at least in part, an LLM as judge to, like, determine which ones are good and bad, it's so easy to just throw in an extra term and be like, "Hey, that, like, weird thing that you keep doing, like, if it, if it does that, like, that's bad.

Give it a low reward." So we've, we've done this with a bunch of customers, and, like, reward hacking does happen, but, like, you, you just see it, and you, like, adjust your, you know, reward prompt, and it just goes away.

Swyx1:06:18

What's a thing from YC that guided you through your entrepreneurship journey? And what, what's one thing that maybe you, like, find that you disagree with YC on?

YC Advice1:06:18

Kyle Corbitt1:06:27

Oh, that's a good question. One thing that I, that I, I really identify with, and I've tried to do a good job, is kind of like, you know, sort of, uh, I, I think they say, like, "Hold your, um, problem tight and your solution loosely," right?

Swyx1:06:39

Mm.

Kyle Corbitt1:06:39

Where it's like-

Swyx1:06:39

That's what you did.

Kyle Corbitt1:06:40

Yeah. Spend a lot of time thinking about what is the problem people are trying to solve, and then it's like, don't be too bought into, like, the way you're solving it today. Um, I think that's super important. Everyone, you know, uh, it's, it's very easy to, to, to get that balance wrong if you're not thinking about it very consciously.

But something I disagree with... That's a good question. I think there's, there's lo- lots of things I disagree with, but I don't have it, like, cached in that direction in my brain. Um, I don't know. Like, I, I, like, I, I definitely have disagreed with lots of specific pieces of advice, but, um, uh, yeah, I don't, I don't have, like, a great answer right now.

Swyx1:07:11

I'll bridge it for you in case, in case, uh, something comes up. Uh, Sam Altman's like, you know, "Everything I said as president of YC was wrong for OpenAI," right? Like, do B2B, ended up doing B2C. You know, you should ship, like, products often, like, ended up being in stealth for three years.

Kyle Corbitt1:07:27

Mm-hmm.

Swyx1:07:27

Like

Kyle Corbitt1:07:30

Yeah. Yeah. Actually, I think that, that second one does resonate with me a lot. Like, we have tried to ship really quickly and just kind of like, sort of like follow the gradient of the market. I think if I do another startup, like, and I don't know, maybe this is just me, like, being beat up by the market too much.

If I do another startup, like, I would, like... I think at least some points I probably would have done better to be, like, heads down and execute on my vision for longer and, like, kind of like go for the more ambitious thing, but that would take longer to sort of, like, prove value, which is definitely not the YC way.

But I think if you have, like, I don't know, a good vision and good taste, then, like, that, that can, like, work quite well.

Swyx1:08:05

Yeah. We'll see what that is whenever that comes out. But, uh, thanks for your time. This is a great overview of everything.

Kyle Corbitt1:08:10

Yeah. Thank you, guys. This has been a super fun conversation. Thanks to both of you.

Swyx1:08:13

Awesome.