# ⚡️Every product of the future will be a living system  — Ronak Malde, Trajectory.ai

Latent Space · 2026-06-21

<https://addtry.com/41092015-8880-48a5-96c7-fe0951269393>

Ronak Malde, CEO of Trajectory.ai, recounts his journey from building AI coding agents at Windsurf (acquired by Google DeepMind in a deal involving Demis Hassabis and Sergey Brin) to launching a platform for continual learning in enterprise AI. He argues that every future product must be a 'living system' that learns from real-world user interactions, not static models. The episode details Trajectory's technical innovations, including Self-Distillation Policy Optimization (SDPO) for learning from corrections, continuous LoRA for parallel training, and open-sourcing a training stack with SkyRL. It covers partnerships with Harvey and NVIDIA to train NeMoTron 3 Super for legal workflows, improving metrics like issue spotting and citation accuracy while cutting costs. Malde explains data curation strategies that capture nuanced user edits beyond binary signals, and outlines Trajectory's roadmap from AI-native companies like Clay and Decagon to Fortune 500 enterprises. He also reveals that the idea for Trajectory emerged after giving up his acquisition equity to pursue this vision.

## Questions this episode answers

### How does Trajectory.ai help companies continuously improve their AI models?

Ronak Malde describes Trajectory as a platform that captures all user interactions and agent behaviors—called trajectories—and distills them into training signals. It creates automated evaluation and training loops so companies can continually optimize their models for specific workflows. Trajectory already helped Harvey, a legal AI company, train the NeMoTron 3 Super model to outperform frontier models on legal tasks while being cheaper and faster.

[10:34](https://addtry.com/41092015-8880-48a5-96c7-fe0951269393?t=634000)

### What happened when Google DeepMind acquired Windsurf?

Ronak Malde recounts that after news leaked of a potential OpenAI acquisition, the Windsurf CTO secretly summoned the team to a hotel, where they unexpectedly joined a call with Demis Hassabis and Google DeepMind executives. The entire process from meeting to becoming full employees at DeepMind happened within 24 hours, with the news breaking the next day.

[3:37](https://addtry.com/41092015-8880-48a5-96c7-fe0951269393?t=217000)

### What is SDPO and how does it enable continual learning with real-world corrections?

SDPO, or Self-Distillation Policy Optimization, is an algorithm developed by Trajectory that improves upon standard reinforcement learning by using "privileged information" from real-world corrections. Instead of reducing all feedback to a single reward number, SDPO gives the teacher model hints about correct behavior, then distills the guided outputs into the student model, enabling more efficient learning from textual corrections during continual learning.

[18:23](https://addtry.com/41092015-8880-48a5-96c7-fe0951269393?t=1103000)

## Key moments

- **[0:00] Windsurf Magic**
  - [0:35] Windsurf's AI agent built a 2048 game on first try, recalls Ronak Malde.
- **[1:27] Coding Thesis**
  - [1:38] Ronak Malde chose Windsurf over Google Gemini post-training for the team and fast pace.
  - [2:22] Ronak Malde bet that coding was more economically viable than math for AI's next leap.
  - [2:47] Windsurf trained SWE-1 on user signal beyond code to beat frontier models.
- **[3:28] Acquisition**
  - [3:42] Ronak Malde recounts the secret hotel meeting where Google DeepMind, not OpenAI, acquired Windsurf.
  - [5:22] Sergey Brin was the main exec excited about Windsurf and appeared in the Anti-Gravity launch video.
  - [6:11] "You can't just train on offline benchmarks; you need to understand real-world usage," says Ronak Malde.
- **[6:36] Trajectory's Start**
  - [6:53] Ronak Malde gave up $2 billion in acquisition money to start Trajectory for continual learning.
  - [7:15] Every domain—legal, healthcare, finance—will be supercharged like coding, predicts Ronak Malde.
  - [7:59] Ronak Malde aims to bring the continual learning flywheel of Cursor and Windsurf to every company.
- **[9:05] Harvey & Legal AI**
  - [9:42] Trajectory's platform distills expert traces into 'trajectories' to optimize agents and models end-to-end.
  - [11:05] Trajectory trained NeMoTron 3 Super for Harvey, achieving faster and cheaper legal models at the frontier.
- **[13:39] Model & Data**
- **[18:16] SDPO**
  - [18:23] SDPO uses privileged hints to guide model training with text instead of binary rewards, says Ronak Malde.
- **[21:43] Team & Open Source**
  - [23:22] Trajectory open-sourced Continuous LoRA for concurrent training jobs, cutting wall clock time in half.
- **[25:33] Concurrent Training**
- **[28:07] Enterprise Roadmap**
  - [29:34] Trajectory targets Fortune 500 with continual learning that observes and automates manual processes.
- **[31:44] Hiring & Close**

## Speakers

- **Ronak Malde** (guest)

## Topics

Enterprise, Language Models, Open Source Tools

## Mentioned

AnyScale (company), Clay (company), Codium (company), Cognition (company), Decagon (company), DeepMind (company), Google (company), Harvey (company), Mercor (company), Rogo (company), SkyRL (company), Trajectory (company), Cloud Code (product), Composer (product), Continuous LoRA (product), Cursor (product), Gemini (product), NeMoTron 3 Super (product), SDPO (product), Windsurf (product)

## Transcript

### Windsurf Magic

**Host** [0:04]
Okay. Uh, we are in the remote studio with Ronak from Trajectory. Congrats on your launch and welcome to Latent Space.

**Ronak Malde** [0:10]
Thanks so much. Yeah, great to be chatting with you. Thanks for having me on.

**Host** [0:14]
Yeah. Um, I don't think I, I, like, we, like, quite overlapped when it, when you were at Windsurf, but obviously I was close to the Windsurf, uh, team, Varun and the other guys. And, uh, it's really surprising to see that, uh, you know, you have sort of gone through that journey at DeepMind and come out, uh, with a such a strong thesis on what's next in continual learning.

**Ronak Malde** [0:35]
Yeah, no, it's, it was a great time honestly at, at Windsurf, and it's crazy to think that even two years ago the world just looked so different. Uh, I, I still remember it when we first launched Windsurf, it was, uh, like, Sonnet 3.5 had just come out and we were playing around with the capabilities and everything.

Um, and I distinctly remember it, there's one time where we had started the IDE, we had built the agent out and everything, and then we all started, like, circling around Varun and he was like... He told the agent to, like, make 2048.

And it was just, like, silly game, whatever. We, like, we saw the agent and it just, like, first shot made the game and that was, like, that mind-blowing experience. I mean, obviously now everyone's experienced it, but getting to, to build the first systems, like, see that in the room, it was just so magical at the time.

And then obviously getting to ride that wave to, to now DeepMind and everything is, is super cool.

### Coding Thesis

**Host** [1:27]
I, I think Codium and then Windsurf had a bit of a mix of, like, product engineering and model training. What was your background, like, coming into this?

**Ronak Malde** [1:38]
Yeah. So I had just finished up my master's at Stanford, been doing research in, in AI for, for quite a bit of time. That's actually where my, I met my co-founders as well. Uh, we, we met, like, literally first week of freshman year at Stanford, which is super awesome.

We, we can get to that story after. So then I, I joined Codium. I actually had the option to go to Google for, to work on Gemini post-training. That was kind of the status quo, but I met the team there and I was-- I, I realized just it's, it's all about the team, it's all about moving fast and getting to see that at Windsurf, it was incredible.

So took the bet on that team, joined when, back when it was called Codium. And I had this thesis that it was either gonna be math or coding, some sort of structured information that gets us to the next leap in AI.

And it turns out that coding is a lot more economically viable as well. And so that's basically how I joined the team. I started on, on the research side there. So initially we were training autocomplete models at Windsurf, like back in the Codium extension.

Afterwards we started to realize, okay, we're getting so much user signal, not just like accept, reject or what people are typing in the code, but actually from the agent, right? And like how are you building these entire websites, entire apps?

And, and so we sought out to build our own foundation model to capture all of that kind of user signal and, and really get the flywheel unlocked for Windsurf. And so I started working on this model called SWE-1.

Uh, you might have seen now there's like later iterations like SWE-1.5 and SWE-2 the, uh, cognition's working on. And this was the kind of major unlock for the company as well, is we had all this massive data, um, we were able to post-train on all of that user signal, uh, and now beat the frontier.

Um, so that all happened and it was, it was super fast 'cause then right after that, all the Google and cognition acquisition happened, I think like thirty-five days after or something. So that was, that was super fun.

### Acquisition

**Host** [3:28]
Yeah, we've done episodes around that. Just because it's a matter of just general interest, uh, what, what is your take on that day that it happens?

**Ronak Malde** [3:37]
Yeah, it was, it's crazy. I, I actually don't know if anyone's told this story, but-

**Host** [3:41]
Go ahead.

**Ronak Malde** [3:42]
So I, I guess, yeah, I was on the research side and then if you remember at the time, like, we-- everyone thought that OpenAI would acquire Windsurf, like the, the news had come out and everything and, and so, like, we, we get a, a text before from our CTO, Douglas, and he's just like, "Meet me in this hotel conference room, like, tomorrow morning and don't tell anyone else."

I'm like, "This is so suspicious." But we, we go over there and, like, I fully expected it to be, uh, like Sam Altman or something, right? And it was gonna be OpenAI.

**Host** [4:10]
Oh, you thought you were gonna, like... It was like, it's closed, like party time, you know.

**Ronak Malde** [4:14]
Yeah, yeah. I mean, I thought, like, the new-- like everything had been leaked that it was gonna be OpenAI acquiring us, and so I... Yeah, we go there, we thought it would be OpenAI, and then we go there and it's, it's like Demis on the call and then like all of the, like, DeepMind, like, Google people and so we're like, "Okay, I guess we're going to Google then."

And it all, it all just happened super fast. Like that was on a Thursday morning. Uh, that same day we just, we gave up like all of the, like, badges, computers to Windsurf, and then the following day we were at Google, like full employees, uh, or like at DeepMind.

**Host** [4:45]
Holy crap.

**Ronak Malde** [4:45]
And so, yeah, it just, it was just like insanely fast. And, and like the news came out I think Friday morning. So, uh, yeah, it, it all just happened that week.

**Host** [4:54]
Yeah. And then, uh, Cognition bought the rest on Monday.

**Ronak Malde** [4:57]
Yeah, exactly.

**Host** [4:58]
Yeah. Uh, yeah, very, very crazy. Uh, and uh, it's very cool to hear that Demis was involved in the deal. I, I know that, like, I think there was like several people, like senior people like Sergey, um, pushing for it, and obviously, like, it, it's kind of like a acqui-hire, but also I think they had to just pitch the team.

It was like thirty people or something on like, "Okay, what are you doing at DeepMind?"

**Ronak Malde** [5:22]
Yeah, it was, um... There were a, a lot of execs involved, so Sergey as well. He was, he was like the main guy that was super excited about Windsurf. Um, he was actually on-- I don't know if you saw the Anti-Gravity launch video, uh, that came out, but, uh, he was a, a main character in that as well.

And, and I think for Google and DeepMind, they're also starting to realize that you need to not just own really stellar models, but you also need to know what users are actually doing with products. And like seeing all of that signal, all the stuff that we were capturing at Windsurf, uh, and being able to like really close that loop.

And so then, you know, we, we went over to DeepMind and, uh, launched Anti-Gravity. Got to contribute to Gemini 3 as well. And it was really awesome, like to see- What a powerhouse of a company with a ton of compute can actually do and, and really, really leverage just, like, cutting-edge technology.

And I think now everyone's starting to realize that you can't just train on, like, offline benchmarks or, like, scaling up human-labeled data. You really need to understand, like, what people are doing in the real world, es-especially as AI is becoming more and more prominent for real-world use cases.

**Host** [6:27]
Yeah, totally. I mean, uh, and that thesis clearly is panning out today even with Cursor and xAI doing a, a similar deal.

**Ronak Malde** [6:36]
Exactly, yeah.

### Trajectory's Start

**Host** [6:36]
So yeah. I, I'm happy to fast-forward to when was the moment you decided... I mean, you know, you have a cushy job at DeepMind. I'm sure they are paying you, like, millions of dollars to just do what you're doing.

Um, when's the moment you're like, "Okay, nope," like, uh, you know, "It's, uh, it's time to go start this thing?"

**Ronak Malde** [6:53]
I, yeah, obviously the acquisition for s- was for $2 billion and went over to DeepMind, and then, uh, I decided to give up all the acquisition money to start Trajectory. And it, it all came down to... So I-- obviously we've been having all of these innovations in coding and, and seeing that real-world usage, but it's very clear now that we are getting that explosion for every single other domain, right?

Like legal, healthcare, finance, they're all going to be supercharged in the way that coding was for the last few years. And the other part though is as we're getting this kind of transformation, it's still that AI, like, is-- it's obviously very powerful, but it kind of acts like any other software, uh, in that it's very static.

Uh, like the model that you used yesterday is gonna be the same model and making the same mistakes tomorrow. And all of those corrections you gave it, the edits, like in any product, it's all just being put to waste.

And, and I think I realized that, okay, there are a few technologies like we did at SWE-1, like Cursor, Composer, even Cloud Code, uh, that are starting to build the model around the product and what people are actually doing in the real world, and this is that compounding loop that allows them to pull ahead.

That is what I want to bring to every single company, that power of continual learning, as people now refer to it. And, uh, and, and so, you know, I, I got together with, uh, my co-founders. I mentioned they're, uh, from Stanford.

We, we met in a math class in, in first week of freshman year, and, uh, they went on to do awesome stuff in the research world as well. So Michael was at DeepMind, uh, working on-- or he was in DeepMind proper for a long time working on, uh, robotics there, so he was on the Gemini 1.5 robotics release and, uh, leading a thinking team there.

And then Arjun, uh, was on the Vision Pro for a long time. He, he was working on the core interaction models and training those, uh, shipped those day one to, to all the Vision Pro users. Um, and so in, in some way or another, all of us were kind of working on how AI interacts with the real world, right?

Through either coding, robotics with, uh, AR/VR. And we realized continual learning is kind of the ultimate, like, paradigm to do that. It's like, how do you have humans in the loop? How do you build this intelligence around them that is constantly learning and growing on its own?

And I think that's gonna be the next major unlock of AI is, is capturing all this real-world usage and, and really going with it.

### Harvey & Legal AI

**Host** [9:06]
Yeah. I, I think that's very exciting to see, and I think broadly very true. It's weird because all three of you went to such, like, notable companies. It's almost like you planned this. Like, you were like, "Okay, we'll split up for a while, and then we'll come back together."

Yeah, so, so okay, y-you decide to work on this problem, and I guess what's the approach? Like, you know, like, uh, I think continual learning is a hot topic among a lot of folks, but I think people approach it from different angles, from, like, the train it into the weights versus just have a bunch of skills and marked down files or whatever.

You released SeeLaura recently. Just, I, I guess, like, just, just tell people more about, like, maybe what the schools of thought in continual learning and, and the one that you're choosing.

**Ronak Malde** [9:42]
Totally. So I think it's helpful to actually-- Like we've been talking about just, like, the research and, and stuff that, like, enables this. I think it's useful to think about this from the customer's perspective. So take a, a company like Harvey, for example.

They, they have a ton of complex legal workflows where agents are trying to research documents and then redline a contract. And unlike coding where we can-- obviously a technical audience can just, like, guide the agent in the right direction, we're able to be tolerant to mistakes.

For a field like legal, like getting eighty percent of the way there is the same thing as zero. And, um, and this is for a lot of fields where they have very expert data of what people are doing.

Maybe an agent attempts something, and then the user needs to go and correct it and get the rest of the twenty percent of the way there. That's a useful signal that we're able to learn off of. So actually we, uh, we-- I, I can share my screen and show some of the work we've been doing with Harvey.

**Host** [10:34]
Yeah.

**Ronak Malde** [10:34]
But, uh, I, I guess before I get there, the, the high level of what we've built for all of these companies, uh, is a platform for continual learning. We, we essentially take all of the data, all of the expert traces, the agents, and distill it into one format, which is what we call the trajectory, and it's basically all of the data that is needed to then create the evals, the judges, the environments, everything, all the components that you would need for training, and then we're able to create that end-to-end loop of a self-serve platform to optimize the agent, the model, the harness, everything.

Right now, we're starting with the model, and that's what the, the major kind of alpha is for every company is being able to own models that are faster, cheaper, better at the frontier. And so I can share some of what we've done with Harvey, which is really exciting.

One of the, the main things that Harvey really cares about is obviously they're a regulated industry. They want to have sovereign intelligence as well. Um, and so we partnered with Harvey and NVIDIA in order to train NeMoTron 3 Super, uh, in order to get to the Pareto frontier.

A lot of the legal workflows that they care about, I can actually fast-forward over here, where we were actually able to capture with-- along with Harvey's lab, uh, dataset a lot of their expert legal workflows. And so, um, the things that they care about, for example, are issue spotting, analysis explanation, citation reference, completeness coverage, um, et cetera.

And all of these things the-they have sort of like paralegals and lawyers, uh, going in-house and working on mission-critical tasks. And so with our model, we were actually able to, uh, improve these sort of metrics across the board.

And then also the, the really cool part is NeMoTron is a drastically cheaper and faster model than the frontier. And this matters at scale when you're deploying to not just a, a few users, but hundreds and thous- hundreds of thousands of users across the board.

And so this is something that's very exciting for the legal domain and something that we're able to do on our platform. The other exciting part is because we're building this, I think essentially like every company has their data sources in different formats.

We're able to build this as a scalable platform that gets all of that data without any kind of forward deployed work and really improve the model. And, uh, and the cool kind of proof point is we started out with working with a couple of companies.

Our, our current partners are Clay, uh, Harvey, Rogo, Decagon, Mercor. The first engagement, it, it took like three months to like set up the entire thing. We were kind of building the airplane as it was flying, and now we're able to train a model with Harvey in under a month and then onboarded a new customer, and within a week we were able to train a model and really get the flywheel going.

Um, so it's, it's really cool to see just like how powerful continual learning is in practice. We're obviously just getting started, and I want to be powering kind of every company that, that is pushing the frontier in knowledge work.

**Host** [13:36]
Yeah. Yeah, congrats on this. I mean, a couple observations. First, I think you're the first team doing this I've encountered using NeMoTron 3 Super. Any comments on them as a base model versus us- usually, you know, people choose like one of the Chinese models.

### Model & Data

**Ronak Malde** [13:54]
Yeah, yeah. I, I mean, so I think it's really awesome to see the open source community just in general. Like, obviously China has been like killing it with models. I remember back when we were at Windsurf.

**Host** [14:08]
Yeah, Suite One is now known to be a post-training of Chinese model.

**Ronak Malde** [14:11]
Yeah, yeah. Uh, well, and, and I think back then, like before i- in like twenty twenty-four or so, the open source was so far behind and you had like some Qwen models. You had DeepSeek, too, and it just like didn't seem promising.

And then all of a sudden, DeepSeek V3 comes out and everyone is like, "Wow," They're using frontier paradigms. They obviously open source or like showed the world GRPO. And I think since then, the gap has been closing and closing over time.

I think America or the, the Western world has some work to do still. Like obviously having one trillion parameter models like Kimi, uh, or like, and, and amazing models like GLM and DeepSeek. Uh, I don't think we're quite there yet for that size of model.

Uh, but you can already see like, um, I mean, GPT-2 SS was kind of one of the best base models a year ago that OpenAI put out, uh, in the Western world. Like since then, NeMoTron has gotten significantly better.

And keep in mind, this is a 120 B model, so same size as GPT-2 SS. And, and so I think as time goes on, like Nvidia is investing a ton of money. The, the world is, is gearing up for, I think, a lot more, uh, open source models in the Western world as well.

Um, which does matter like for legal, for finance, for healthcare. And, and so I think this will be a leading paradigm, and so we wanted to be ahead of the curve and, and really make that a powerful part of our platform.

**Host** [15:38]
Got it. I'm sure Nvidia will be happy to, to hear that.

**Ronak Malde** [15:42]
Yeah, yeah.

**Host** [15:43]
And then the other-- Yeah. Uh, I, I think the other thing also is I, I-- the... So for me, the general question of like why people haven't pursued, you know, just let's call it online models or con- continual learning is like just the curation of the data.

Like what, what is the sort of data curation solution here?

**Ronak Malde** [16:01]
Yeah, yeah, totally. So I think the, the interesting thing here is like every company is storing their usage, their, their data in slightly a different way, right? And you have this entire system o- of basically, uh, two things.

One, what are agents doing in a product? Um, so you can imagine for Cursor, they're obviously like writing code, they're running commands. Uh, and then for like, I don't know, a GTM workflow, it could be that the agent's creating tables or enriching columns.

And then the-- So that's the easy part. That's like in a OpenAI like tool called Tool Response Format. The interesting part though is what are all of the other things that users are doing in a product? And so at Windsurf, for example, this was, uh, users are modifying the code after an agent.

They're saying like, "Hey, I actually wanted this button on the left after the agent wrote a website," or like, "I wanted this class structured this way." That's all the very useful information that's happening in a product. I, I think when people think about like online learning, continual learning, they'll first think of like, like accept, reject or thumbs up, thumbs down.

Like some of those like kind of binary signals. Uh, it turns out that that's actually like very noisy. You can imagine like when is the last time you put a thumbs u- a thumbs down in Cursor or Cloud Code?

Like, uh, never, right? And everyone's kind of vibe accepting everything. So it turns out those signals are actually quite noisy. But it-- what is useful is all the directions of like the user is saying, "Hey, you should have gone this way," or like modifying some of the attempts of the agent.

And the cool part about this is this works in any product. So, uh, for a legal workflow, it's the same thing. The agent attempts to do some research and then write some red lines or create an Excel spreadsheet.

And then afterwards the lawyer goes back and maybe has to do some more research because the agent missed something, and that sort of taste. And then a- again, also like modifying some of the contracts. And so now the tricky part is how do you take all this data and then condense it into the right format to train on?

And so there's a couple of really exciting things here that we've been working on. One is just purely an infrastructure question. When we figured out a way to take all this information and, uh, build an auto research pipeline that is now able to take, uh, like whatever these vectors are of what the user is doing and make those into judges, into evals environments.

Um, but the other exciting part is the actual algorithms behind continual learning. Um, so one of the other really cool things, I'll, I'll just go ahead and make-- share my screen again, uh, that we released very recently, is scaling up, uh-

### SDPO

**Host** [18:23]
SDPO

**Ronak Malde** [18:23]
... SDPO or, uh, self distillation policy optimization. So if you're not familiar with this, I, I can give the high level real quick. So for RL, basically the way RL works, uh, and that's kind of the leading post-training paradigm, is you have an agent, and it goes and explores the world a bunch of times, and then at the end of the day, you get a reward signal.

It could be a judge, it could be like something from production. And then, uh, you basically take that number and then use that to r- update and say like, "Hey, this entire tra- trajectory was really good. This entire trajectory is really bad," and then it learns from that.

So it works pretty well. Obviously, all of our frontier models that we see today are trained on RL, but it doesn't really work for this new world of, uh, continual learning. The reason is RL, it's still taking all of this kind of useful information from the real world, like I mentioned, all the corrections and everything, and putting it into just one number, which is really broken.

And so what self distillation does, which is really cool, is takes... Or, uh, I guess the high level here is we start out with a student rollout. So i- if you're familiar with distillation in general, it's you have a student, uh, that is usually a less smart model, and then a teacher, which is a smarter model.

And you're basically saying, "Okay, let me distill all the smart stuff from the teacher into the student." The cool unlock with self distillation is you don't actually need a smarter teacher. Your student could be the smartest model. But to make the teacher smarter, I'm actually going to take some privileged information or a hint, uh, and put that into context of the teacher, and now I have a suddenly slightly smaller, smarter model that I can fit into.

So in practice, what this looks like, let's, let's take this example here, where we have an agent that's asking... Or the user asks, like, "How much is my flight ticket to New York?" We have the agent, like, look up some stuff, uh, and then we, we actually get, like, the wrong ticket information.

Now, we can actually go back, because from production we have some hidden information, and we can say, "Actually, uh, let me give the teacher a smart hint." Then we match the student log probs, uh, to that, uh, teacher information, and then suddenly we're able to take not just like a binary reward, but truly like actual text and guide the model in that direction.

**Host** [20:48]
Hmm.

**Ronak Malde** [20:48]
So this is a huge unlock of SDPO, and we've done some very exciting kind of modifications and scaling it up. It's been done in a lot of academic cases, but no one's actually been able to scale it up to real world use cases.

So we started training on Apex Agents, which is a, an awesome benchmark that Mercor put out. We're just modeling a lot of real world behavior, and we're able to see very amazing gains on some of these workflows. So for example, also the convergence rate just like goes up way faster.

This is normal, uh, GRPO style training, and then we're able to obviously be way faster. So this, uh, this on our blog post as well is very awesome to see. And we do some modifications as well to training on real off policy data, which is like what happens in production, not necessarily from your current model, and making it a lot more robust to real world use cases.

### Team & Open Source

**Host** [21:44]
Yeah. Ma- That's-- This is incredible work. I can't believe you, you basically shipped like research alongside of your products and as a, as a team of three.

**Ronak Malde** [21:53]
Well, so I, I... Actually we, we have 11 of us now. So we've been growing really fast. We, uh-

**Host** [21:57]
Yeah, you have a nice office

**Ronak Malde** [21:59]
... we have a... Yeah, yeah, yeah. Um, I, I think in order to build something really awesome in this space, we, we think of ourselves as a product company and a research lab, uh, because I think it requires cutting-edge research.

Uh, the current algorithms aren't quite there yet to get to true continual learning, and so that's a lot of what we're exploring. But then it also requires deeply understanding customers and, like, building a good interface. No one's really made the product for when intelligence is just getting smarter every single day.

And I think there's a lot of very interesting interface questions here. So, but the team is really awesome. We've, we've hired an amazing research team from OpenAI, from, uh, Meta Superintelligence, uh, from, uh, Amazon AGI, and then obviously the co-founders from, from DeepMind and Apple.

And then also really growing out the product team as well. So, uh, we, we have someone on the team from Stripe who is, uh, building a lot of backend SDKs and then, uh, someone from Figma as well, uh, really thinking about these interfaces and designing them.

So I, I think it's a concerted effort across both product and research to get here.

**Host** [22:54]
Yeah. Let's go through more of the core IP because I don't want to cut you short. You, you were just getting into some good stuff with SDPO. You have like this like five days of Trajectory, which is also like very, very strong launch.

I think one of the strongest I've ever seen, uh, you know, in, in this space. Uh, yeah, keep going. Like, uh, don't let me stop.

**Ronak Malde** [23:10]
Yeah, yeah. Um, I, I guess I can share one more as well. And, uh, yeah, it, it was super fun. Like we've just been doing so much across research and product, and so it was really awesome to, to be able to showcase that to the world in, in our five days at Trajectory.

So I, I can share one more as well is, uh... I mean our, our goal as Trajectory is to really empower every single company to just like get to the world of continual learning. It's such an exciting unlock when you can just put a product out there to a couple hundred users, and now the intelligence just grows over time.

And so one of the other exciting things we did is, is open sourcing a training stack for continual learning. And so this is, uh, in conjunction with, uh, SkyRL, uh, Berkeley's SkyRL lab, um, in any scale as well.

And so, so basically the, the high level here is up until now with training, it's very much like a linear-- Like if you ever kicked off a run with like Slurm or, or one of these like normal training clusters, it's like you kick off a job, it like starts to spin up a bunch of resources, spin up the GPUs in order to do the sampler and do the training, and then you run the entire life cycle of the training job and then spin it down.

This is like normally how training works. The problem is with continual learning, you don't really have like training explicitly starting and stopping. You're probably running concurrent jobs at the same time, and so the normal training paradigms start to break apart.

This is also very similar to the thesis like- Uh, like thinking machines, like a few of these other, like, distributed LoRA, uh, sort of companies infrastructure has had. Our contribution is to make this open source, uh, for every company.

Um, one of the things here is like... Okay, so this is like normal life cycle of a job is like set up, you do a bunch of sampling training, and then clean up. What continuous LoRA does is say, "Hey, instead of a bunch of linear jobs like this, I can actually stack things together, have one dedicated pool for training and one dedicated pool for sampling, and then now I can start to put all of these pieces together and run jobs in parallel."

And so the results are really cool. We ran this on several different scale-up of experiments, and you can see that, uh, across the board, even as you scale up experiments, so like two concurrent jobs, we're able to cut the wall clock time in half.

As this goes up to four jobs, eight jobs, uh, we're able to run concurrency really well. And, uh-

### Concurrent Training

**Host** [25:34]
Yeah, how come, how come eight goes up?

**Ronak Malde** [25:36]
I can also share it here. Let me, let's walk through here. So-

**Host** [25:39]
Oh, okay. Oh.

**Ronak Malde** [25:40]
Yeah. Okay. So the way this works is... Okay, so up here is like serial training run, right?

**Host** [25:45]
Yeah.

**Ronak Malde** [25:46]
Uh, like you're doing a bunch of sampling rollouts, whatever. Um, but if you have a fixed set of GPUs, then you have to run the job sequentially after another. With con- continual LoRA, so it, it is actually slower per training job.

Uh, like if you just care about, like, I just want run one to finish as fast as possible-

**Host** [26:02]
Yeah, makes sense

**Ronak Malde** [26:02]
... then you should be doing like completely serially, 'cause all the resources are going towards that. But with continual learning, what you care about is just like, let me have all of these runs, like learning on the job.

It also might be that like data is coming in not quite sequentially like this, but maybe in batches, right? You get new data from production, now you need to have like run four suddenly take up more resources. So it turns into a lot more of a dynamic workload, and you can start to see here like, like eight concurrent runs, for example, it takes about the same time as like sequentially, uh, like three jobs would take.

And so this also scales out, like we didn't fully push the bounds of this in this blog post. Uh, we're, we're currently scaling this up to like sixteen concurrent runs and scaling up to... Th- this was also on like smaller models as well, like scaling this up to like 235Bs.

And, uh, and the really cool part is obviously like this doesn't degrade performance whatsoever. It's still the same algorithms. We're able to see the same gains with any, any sort of normal model. And, uh, and we wanted to also make this very plug and play.

So we, we have some good instructions here of just how to get started using it in SkyRL, which is really exciting. So-

**Host** [27:10]
Yeah. It's really funny as a operating systems fan, like that you see like very similar like scheduler type problems, starvation, preemptive scheduling, all these kinds of terms. Like you just take a normal OS class, you like, then you start applying it to every, every problem domain like this.

**Ronak Malde** [27:26]
Yeah. Yeah. Exactly. I, I think it's almost like training up until now has been a very research problem, and I think people have attacked it from a research-y perspective, where you don't care about all of these like kind of distributed systems, uh, innovations.

Uh, and suddenly now, like training is becoming more production, right? Uh, like the end goal of Trajectory is that every company can be kicking off their own training runs, like continuously learning, updating, and that also requires a training stack that is built with a systems level or engineering kind of mindset.

And so, so this is just kind of the beginning, but I think there's so many different insights from, uh, distributed computing that can be applied to here.

**Host** [28:07]
Got it. By the time this comes out, you will have, uh, sort of finished your, your five days. Any other sort of overall tech bets or directions that you want to hint the audience towards?

### Enterprise Roadmap

**Ronak Malde** [28:19]
Yeah, totally. So this, uh, where we're at with the product right now is we're able to take all the kind of noisy information, uh, from a company, like all the valuable signal of what people are doing in a product, be able to train the model and then re-upload that, and then have that like observability that evals all of those things.

But right now it's still kind of like these single components, uh, of the stack and a lot of these things like we are, we are kind of like managing or, or looking at like kicking off the training runs.

The next step after this is, one, I would like kind of what we think of as act two of the company, is I wanna make this legible to the customer. So just like the right observability tools, the right abstraction levels for evals.

And then second, give them control. So like I, I want to have the customer say, like a PM should, uh, that's like managing an agent product, should be able to say, "Hey, okay, here's where my agent is really good.

Here are the areas where it's still failing." And now all I have to do is just like modify the model, modify like all of the pieces, and suddenly tomorrow I wake up and I have a smarter model that I can see in production.

All of those things like should be in control of the customer. So that's the next phase. Uh, and then after that, the most exciting part, uh, is right now we're working with these AI native companies. I mentioned Clay, Harvey, Decagon, Rogo.

Um, soon we'll be working with the kind of tech incumbents. You can imagine like the Airtable, Notion, those sort of companies. The stage after that is Fortune 500. Uh, and in order to get there, we need to get to true continual learning that I can just have just the observability layer, right?

Like I'm observing, let's say in Walmart, like what are users doing? What are all of the manual kind of processes? And then start to say, "Hey, okay, maybe a model isn't even possible today if you were to just use an off-the-shelf sort of model."

But I can now observe exactly what users are doing. I can start to see the patterns, build out an agent to harness a model that all works perfectly for that workflow. And now you can imagine any sort of company, their front using product should just dynamically be learning from their users and constantly updating.

In order to get here as well, tactically, we're starting with the models, but very soon I, I want to be also improving the harness, improving skills, doing maybe even the memory layer as well. I think all of those in conjunction are the continual learning solution.

Um, yeah, I know you mentioned that at the beginning as well, and I, I think all of those are super important and also the interplay between the two. Like no one really knows what it means to train the harn- like optimize the harness and the model together, right?

And I think there's very interesting research there in, in what that looks like. So these are all very exciting areas that we're, we're, uh, like super stoked to get into.

**Host** [30:52]
Is it worth focusing on coding alone as a verifiable domain, or are you unopi- unopinionated? 'Cause you got legal, you got the other stuff.

**Ronak Malde** [31:00]
Yeah. So for us, I mean, I, I think that just like we saw in coding for the last two years, I think that's going to happen for every single company and, and every single domain. So our thesis was basically, okay, like coding's moving super fast.

Let's accelerate the rest of these industries to first just get to where coding is, but then after that, like, truly transform all of them with continual learning. So that's why we're starting with also companies in different verticals, uh, to also build our platform so that it's agnostic to whatever domain you're in.

**Host** [31:31]
Okay, cool. Like, very ambitious. I, I, I think I finally get it more after talking to you, and hopefully people listening along can, can get a sense of the scope of the vision as well. Where should you-- should people point to for, uh, trying stuff out?

Um, you know, what are you looking for in terms of hiring? Anything like that.

### Hiring & Close

**Ronak Malde** [31:48]
Yeah, yeah, for sure. No, so it-- we're gonna be growing really fast, and it, it's exciting to see that. I would say, like, honestly, the, the main place where all of the, the research world lives on is Twitter, and so following along on there and...

Uh, no, and I, and I, like, I want people to also, like, engage and be excited about what we're building. I, I felt like when we were at, at Windsurf again, like doing, uh, just, like, seeing how developers were using coding agents and, like, being able to be really in tune with that, uh, was super important.

And, and for us, that's the same way. Like, I wanna, uh, just be in tune with, like, how customers are using our product and how, how continual learning research is, is kinda scaling up and us being able to, um, at least contribute some, uh, innovations to that field.

And, uh, and we're gonna be-- I think, like, the two areas for us are obviously, like, research and, and growing that team and product as well. We, we wanna be really robustifying our platform, scaling it up to, uh, not just five customers, but fifteen, twenty, fifty very soon.

So, uh, yeah, that- ... that's kinda the next frontiers.

**Host** [32:50]
You got five very good customers in there that, uh, are extremely good logos for a lot of people. Uh, yeah, I'm excited to hear more and, and see more, you know, and people can check you out for sure.

I think I've also invited you to my conference, uh, AIE, where w-we are doing a continual learning track. I was very cautious about doing it. I was like, "I don't know if the field is ready." But with you guys and some of the other folks, uh, also working in continual learning, I think it's time.

Even though if we don't exactly know what the end product will look like, I think there's enough test cases where, like, yeah, like people are trying to put this in production already.

**Ronak Malde** [33:22]
Yeah, I, I think it-- now is the time of continual learning, and so it's super exciting. I, I'm very excited to be at the, the fair.

**Host** [33:29]
Yeah, yeah. All right. Congrats. Amazing launch. Seriously, like, so well executed. I'm, I'm actually, like, very, very impressed. As, like, someone on the marketing side, right, like I'm like, "Man, like, I-"

**Ronak Malde** [33:38]
Actually, yeah, that, that means a lot coming from you. Yeah.

**Host** [33:42]
I'm like, "Okay, like, this is, like, the bar now for, like, how to launch a, a, a products lab," you know? So no, congrats and, uh, excited to see more.

**Ronak Malde** [33:50]
Thank you so much. Yeah, thanks for taking the time to chat. This was super fun.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
