# Windsurf: The Enterprise AI IDE

Latent Space · 2024-12-13

<https://addtry.com/6270e96a-9d07-4811-9fe3-2e86aefcdd27>

Varun Mohan and Anshul of Codeium explain why they built Windsurf, a new AI IDE, arguing that VS Code's API limitations prevented them from delivering the best agentic experience. They detail how Cascade, their agentic system, uses proprietary retrieval and planning models alongside third-party LLMs, and describe their evaluation method that masks commits and tests incomplete code states. The duo also reveals that over 800,000 developers use Codeium extensions, that they still support JetBrains and Eclipse for enterprise customers, and that they intentionally avoided a waitlist launch. They discuss the trade-offs of building first-party vs. third-party models, the importance of 'go slow to go fast' in enterprise infrastructure, and their belief that individual developer profits should come after building switching costs through superior product.

## Questions this episode answers

### Why did Codeium build their own IDE (Windsurf) instead of just improving their VS Code extension?

Varun Mohan explains that VS Code’s API limitations prevented the deep integration needed for their agentic AI. For example, to show super complete refactoring suggestions, they had to dynamically generate PNGs because no API existed. More importantly, they couldn’t observe developer trajectories—like which files were opened or edits made—to infer intent automatically. Building Windsurf gave them full control over the UI and system hooks, enabling a more intuitive, powerful AI experience.

[5:53](https://addtry.com/6270e96a-9d07-4811-9fe3-2e86aefcdd27?t=353000)

### How does Codeium evaluate the performance of their AI coding agent Cascade?

Varun describes an approach using real open-source commits: they strip the commit from a codebase, leaving it incomplete, and test whether the system can retrieve relevant context, form a plan, and execute changes to make associated tests pass—without the commit message. This mimics real developers who rarely articulate their full intent. It turns a binary success metric into a continuous one, allowing iterative improvement on retrieval, planning, and execution.

[10:25](https://addtry.com/6270e96a-9d07-4811-9fe3-2e86aefcdd27?t=625000)

### Why does Codeium keep autocomplete free for individuals while focusing on enterprise?

Varun says individual developers are too price-sensitive and churn-prone; a small price change can make them switch tools. Codeium instead targets enterprises where customers already spend millions on software, so they value deeper integration and outcomes over per-seat cost. By keeping a free tier for individuals, they gather feedback and scale, then convert that into enterprise success. Their focus is on building a durable business, not short-term consumer profit.

[28:22](https://addtry.com/6270e96a-9d07-4811-9fe3-2e86aefcdd27?t=1702000)

## Key moments

- **[0:00] Intros**
  - [0:22] Codeium's new office was previously used by Facebook/WhatsApp and Ghost Autonomy.
  - [1:17] The office was used for external shots in the TV show Silicon Valley.
  - [1:45] Codeium now has over 800,000 developers using its product.
  - [2:09] Codeium received J.P. Morgan Chase's Hall of Innovation Award within a year of deploying.
  - [2:50] GitHub has less than 10% penetration for source code management in the Fortune 500.
  - [3:28] Over 70% of developers at some enterprise customers use JetBrains IDEs.
- **[3:52] Why Windsurf**
  - [3:55] Codeium built Windsurf, its own IDE, to provide the best agentic AI experience without VS Code limitations.
  - [4:46] Windsurf aims for a magical experience where users don't need to manually tag code for context.
  - [5:08] Varun Mohan: 'As soon as you see the mountain, the AI helps you get there, and then creates it for you.'
  - [6:51] Anshul: Understanding developers' editor trajectory lets AI immediately infer intent without explicit prompts.
  - [7:27] Anshul: Codeium had been considering building an IDE for a long time, but only recently became feasible due to model capabilities.
  - [8:17] Varun Mohan: 'We ended up dynamically generating PNGs to showcase the feature because VS Code had no API.'
- **[10:12] Evals**
  - [10:24] Codeium's evaluation strips commits from open source repositories and tests if the AI can complete the code and pass tests.
  - [11:19] Codeium's eval approach converts the discrete task of making a PR work into a continuous optimization problem.
  - [11:46] Varun Mohan: 'We believe developers will never completely pose the problem statement.'
  - [12:47] Windsurf's real value is making a good first pass on large codebases, not just greenfield projects.
  - [13:23] Anshul: Most existing software development benchmarks, like SWE-bench and HumanEval, are 'kind of bogus' for professional work.
  - [13:39] Codeium builds custom retrieval evals by looking at old commits to determine which files are semantically related.
- **[16:15] Launch & Remote**
  - [16:28] When Codeium launched, the first Hacker News comment claimed 'This product is a virus,' which they found amusing.
  - [16:51] Anshul: 'We just wanna give autocomplete suggestions. That's all we wanna do.' — on early Hacker News criticism.
  - [17:15] Q: Will Cascade become a fully end-to-end agent like Devin? Anshul: Only when it can autonomously complete tasks without human intervention.
  - [18:28] Varun Mohan: The goal is for Windsurf to become fully agentic with limited human interaction, but hard problems remain.
  - [18:42] Varun Mohan: The most annoying part of Cascade is having to accept every terminal command, but fully automating risks 'rm -rf' disasters.
  - [19:11] Codeium is considering running agents on remote machines instead of locally to avoid destructive commands.
- **[25:18] Cascade Strategy**
  - [25:18] Varun Mohan: Windsurf uses Claude/OpenAI for high-level planning, but proprietary models for fast retrieval and applying plans to codebases.
  - [25:25] Shawn Wang mentions 'late interaction' (ColBERT) as similar to Codeium's retrieval approach, which Varun hadn't heard of.
  - [26:10] Anshul: Distributed compute for retrieval over raw data is not new; Codeium applies it to LLMs for code search.
  - [26:30] Varun Mohan: Codeium's same infrastructure serves both individual developers and large enterprises, avoiding custom systems.
  - [26:56] Varun Mohan: Building and owning their own serving infrastructure is a core competency that differentiates Codeium.
  - [27:15] Shawn Wang: Codeium's approach to building infrastructure deliberately for the enterprise is a 'go slow to go fast' strategy.
- **[33:12] Agents & Fixes**
  - [33:35] Q: Has Codeium explored multi-agent systems? Varun: Yes, but not deployed due to side-effect risks and latency concerns.
  - [33:51] Varun Mohan: Codeium has internally analyzed multi-agent approaches, spawning multiple trajectories to validate hypotheses.
  - [34:37] Varun Mohan: Running multiple agent rollouts in parallel is better suited for remote machines, not local ones.
  - [35:06] Anshul: Codeium's internal poll showed more excitement for the launch in a month than for the current Windsurf launch.
  - [35:38] Anshul: Upcoming Windsurf features include automatically executing terminal commands and deeper human trajectory analysis.
  - [36:03] Anshul jokes that Cascade might one day suggest actions before you type, akin to 'Clippy's coming back.'
- **[39:12] Enterprise**
  - [39:23] Anshul: Codeium remains the only AI coding assistant with an Eclipse extension, supporting enterprise developers on legacy IDEs.
  - [39:51] Anshul: Codeium's mission remains to maximize AI value for every developer, regardless of their IDE or environment.
  - [40:28] Anshul reflects on Codeium's evolution from an infrastructure company to a developer tools company.
  - [40:54] Anshul: Enterprise adoption of Codeium often hinges on whether developers love the product, as executives ask that directly.
  - [41:35] Codeium uses the same engineering team for both the individual IDE and enterprise extension products, avoiding separate silos.
  - [41:47] Varun Mohan: Codeium's ability to pivot from GPU virtualization to code AI was due to a versatile engineering team.
- **[42:01] Blog Lessons**
  - [42:03] Anshul revisits his 2022 blog post on building Copilot for X, giving himself a B-minus for accuracy over time.
  - [42:46] Anshul: Enterprise codebases can be hundreds of billions of tokens, so context windows are still insufficient.
  - [43:01] Anshul admits they were wrong in 2022 about first-party vs third-party models; now they rely on Claude/OpenAI for planning.
  - [43:17] Anshul: Windsurf's agentic features only became possible due to rapid model improvements in GPT-4 and Claude 3.5 Sonnet.
  - [44:36] Varun Mohan: Many experienced engineers at Codeium didn't get value from ChatGPT because they were already efficient with existing tools.
  - [45:10] Varun Mohan: His co-founder never used ChatGPT, highlighting the need for passive, integrated AI tools.
- **[52:49] Durable Build**
  - [52:49] Shawn Wang: Anshul's third guest post for Latent Space achieves a 'hat trick' — a rare feat.
  - [53:20] Anshul: Codeium focuses on driving real value and revenue, not just hype, unlike some San Francisco AI companies.
  - [53:21] Varun Mohan: Codeium wants to be a durable, cash-regenerative company to invest sustainably in transforming software development.
  - [54:25] Anshul's post advocates building for enterprise (security, compliance, scale) from day one — 'go slow to go fast.'
  - [55:19] Codeium's early investment in containerized deployments and security is easing FedRAMP accreditation for defense customers.
  - [56:27] Varun Mohan: Buy for undifferentiated needs, but build core competencies that you can never regain if lost, like model inference.
- **[1:01:43] Sales Scale**
  - [1:02:02] Q: What advice for hiring a VP of sales in AI dev tools? Varun: Look for intellectual curiosity, ability to build a scalable factory.
  - [1:03:16] Varun Mohan: Unlike selling databases, Codeium's sales team must understand RAG and AI to credibly partner with customers.
  - [1:04:13] Varun Mohan's hiring philosophy for sales: 'Talk to enough people, find out what good looks like, find someone humble.'
  - [1:05:24] Codeium uses 'deployed engineers' similar to Palantir, partnering with sales to deeply understand customer AI use cases.

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **Anshul** (guest)
- **Varun Mohan** (guest)

## Topics

Agent Platforms

## Mentioned

Codeium (company), Groq (company), JetBrains (company), 4o (product), Cascade (product), ChatGPT (product), Claude (product), Copilot (product), Devin (product), Eclipse (product), Llama (product), SWE-Bench (product), VS Code (product), Warp (product), Windsurf (product)

## Transcript

### Intros

**Swyx** [0:05]
Hey, everyone. Welcome to the Latent Space Podcast.

**Alessio** [0:08]
This is Alessio, partner and CTO at Decibel Partners, and I'm joined by my co-host Suwix, founder of Smol AI.

**Swyx** [0:13]
Hey, and today we are delighted to be, I think, the first podcast in the new Codeium office. So thanks for having us, and welcome Varun and Anshul.

**Anshul** [0:22]
Thanks for having us.

**Varun Mohan** [0:23]
Yeah, thanks for having us.

**Swyx** [0:24]
This is the Silicon Valley office?

**Varun Mohan** [0:26]
Yeah.

**Swyx** [0:26]
So, like, what's the story behind this?

**Varun Mohan** [0:28]
The story is that the office was previously-- So we used to be on Castro Street, so this is in Mountain View, and I think a lot of the people at the company previously, you know, were in SF or are still in SF.

And actually, one thing you-- if you notice about the company is it's actually like a two-minute walk from the Caltrain, and I think we were-- we didn't want to move the office, like, very far away from the Caltrain.

That would probably, you know, piss off a lot of the people that lived in, in, in San Francisco, this guy included.

**Anshul** [0:54]
Yep.

**Varun Mohan** [0:55]
Um, so, so we were, like, scouting a lot of spaces in the nearby area, and this area popped up. It previously was, was being leased by, I think, Facebook/WhatsApp, and then immediately after that, uh, Ghost Autonomy. And then, and now here we are.

And we also-- You know, I guess one of the things that the landlord told us was this was the place that they shot all the scenes for Silicon Valley, at least like externally and stuff like that. So that just became a meme.

Trust me, that wasn't, like, the main reason why we picked it. Um, but we've leaned into it.

**Swyx** [1:22]
It doesn't hurt.

**Varun Mohan** [1:22]
Yeah.

**Swyx** [1:23]
Yeah. Um, and obviously, that played a little bit into your launch, uh, with Windsurf as well. So, uh, let's get caught up. Uh, you were guest number four, I think.

**Alessio** [1:31]
I think it was two.

**Anshul** [1:32]
Maybe it was two.

**Swyx** [1:32]
Two.

**Anshul** [1:33]
Might have been two.

**Alessio** [1:34]
Yeah.

**Swyx** [1:34]
Um, so a lot has happened since then. Uh, you've, you've raised a huge round and also just launched your, your ID. Like, what's, what's been the progress over the, the last year or so since, since, um, the Latent Space people last saw you?

**Varun Mohan** [1:45]
Yeah. So I think the biggest things that have happened are Codeium's extensions have continued to gain a lot of sort of popularity. Uh, you know, we have over eight hundred thousand sort of developers, um, that use that product.

Lots of large enterprises also use the product. We were recently awarded J.P. Morgan Chase's Hall of Innovation Award, uh, which is usually not something a company gets, you know, within a year of deploying an enterprise product. Uh, and then large companies like Dell and stuff use the product.

So I think we've seen a lot of traction on the enterprise space. But I think one of the most exciting things we've launched recently was-- is this actually IDE called Windsurf. And I think for us, one of the things that we've always thought about is: How do we build the most powerful AI system for developers everywhere?

The reason why we started out with the extension system was we felt that there were lots of developers that were not going to be on one platform, and that still is true, by the way. Outside of Silicon Valley, a lot of people don't use GitHub, right?

This is like a very surprising finding.

**Swyx** [2:37]
What?

**Varun Mohan** [2:37]
But most people use, uh, GitLab, Bitbucket, Gerrit, Perforce, CVS, Harvest, Mercurial. I could keep going down the list, but there's probably ten of them. GitHub might have less than ten percent penetration of the Fortune five hundred, full penetration.

It's very small. And then also on top of that, GitHub has very high switching costs for source code management tools, right? Because you actually need to switch over all the dependent systems on this workflow software. It's much harder than even switching off of a database.

So because of that, we actually found ways in which we could be better partners to our customers, regardless of where they stored their source code. And then more specifically on the IDE category, a lot of developers, surprise, surprise, don't just write TypeScript and Python, right?

They write Java, they write, uh, Golang, they write a lot of different languages. And then high quality language servers and debuggers matter. Very honestly, JetBrains has the best debugger for Java. It's not even close, right? These are extremely complex pieces of software.

We have customers where over seventy percent of their developers use JetBrains. And because of that, we wanted to provide a great experience wherever the developer was. But one thing that we found was lacking was, you know, we were running into the limitations of building within the VS Code ecosystem on the VS Code platform.

And I think we felt that there was an opportunity for us to build a premier sort of experience, and that was within the reach of the team, right? The team has done all the work, all the infrastructure work to build the best possible experience, right, and plug it into every IDE.

### Why Windsurf

**Varun Mohan** [3:55]
Why don't we just build our own IDE that is by far the best experience? And as these agentic products sort of become more and more possible, and all the research we had done on retrieval and just reasoning about code bases became more and more to life, we were like, "Hey, if we launch this agentic product on top of a system that we didn't have a lot of control over, it's just gonna limit the value of the product, and we're just not gonna be able to build the best tool."

That's why we were super excited to launch Windsurf. I do think it is the most powerful IDE system out there right now, uh, in the capability, right? And this is just the beginning. I think we suspect that there's much, much more we can do, more than just the autocomplete sort of side, right?

When we originally talked, probably autocomplete was the only piece of functionality-

**Swyx** [4:34]
Yes.

**Varun Mohan** [4:34]
-the product actually had. Um, and we've come a long way since then, right? These systems can now reason about large code bases without you adding everything, right? Like, when you use Google, do you say like, @NewYorkTimes post blah, blah, blah, and like ask it a question?

No. We want it to be a magical experience where you don't need to do that. We want it to actually go out and execute code. We think code execution is a really, really important piece. And when you write software, you no long-- you not only just kind of come up with an idea, the way software kind of gets created is software is originally this amorphous blob, and as time goes on and you have an idea, the blob and the cloud sort of disappear, and you see this mountain.

And we want it to be the case that as soon as you see the mountain, the AI helps you get to the mountain, and as soon as you see the mountain, the AI just creates the mountain for you, right?

And that's why we don't believe in this sort of modality where you just write a task and it just goes out and does it, right? It's good for zero to one apps, and I think people have been seeing Windsurf as capable of doing that, and I'll let Anshul talk about that a little bit.

But we've been seeing real value in real software development, which is more to say-- this is not to say that tool-- current tools can't, but I think more in the process of, of actually evolving code from a very basic idea.

Code is not really built as you have a PRD and then you get some, some output out. It's more like you have a general vision-

**Anshul** [5:44]
Iteration.

**Varun Mohan** [5:44]
And-- Yes. And as you write the code, you get more and more clarity on approaches that don't work and do work. You're killing ideas and creating ideas constantly, and we think Windsurf is the right paradigm for that.

**Alessio** [5:53]
Can you spell out what, what you couldn't do in VS Code? Because I think when we did the, the Cursor episode, they explained, then everybody on AgriNews is like- Oh, blah, blah, blah. Why, why did you fork? Why-- You could have done it in an extension.

Like, can you maybe just explain more of those limitations?

**Anshul** [6:09]
I mean, I think a lot of the limitations around like APIs are pretty well documented. I do- I don't know if we need necess-- go-- necessarily go down that rabbit hole. I think it was when we started thinking, "Okay, what are the pieces that we actually need to give the AI to get to that kind of, you know, emergent behavior that Varun talked about, right?"

And, and yes, we were talking about all the knowledge retrieval systems that we've been building for the enterprise all this time. Like, that's obviously a component of that. You know, we were talking about all the different tools that we could give it access to so they can go, like, do that kind of terminal execution and things like that.

Then the third main category that we realized would be, like, kinda that magical thing where you're not out there writing out a PRD, you're not scoping the problem for the AI, is that if we're actually being able to understand the kind of the trajectory of what developers are doing within the editor, right?

If we actually are be able to see like, oh, the developer just went and opened up this part of the directory and tried to view it, then they made these kind of edits, and they tried to do, like, some kind of commands in the terminal.

And if we actually understand that trajectory, then our ability for the AI to just be immediately be like, "Oh, I understand your intent. This is what you wanna do," without you having to spell it all out for it, that is when, like, that kinda, like, magic would really happen.

I think that was kind of like that intuition. So you have, you have the restrictions of the APIs that are well documented. We have the kind of vision of, like, what we actually need to be able to hook into to really expose this, and I think it was that combination of those two where we're like, "I think it's about time to do the editor."

The, the editor was not, like, a, a-- necessarily like a new idea. I think we've been talking about the editor for a very long time. I think it's like, of course, we just pulled it all together in the, in the last couple of months, but it was always something in the back of the mind, and it's only when we started realizing, okay, the models are now capable of doing this.

We actually can look at this data. Like, we have a really good context awareness, and we're like, "I, I think now's the time." And, and, and we went out and executed on it.

**Alessio** [7:47]
Yeah. So it's basically not acti-- it's not like one action you couldn't do, but it's, like, how you brought it all together. It's like-

**Anshul** [7:53]
Yeah

**Alessio** [7:53]
... the VS Code's kinda like sandbox, so to speak, or-

**Varun Mohan** [7:55]
Yeah, let me, let me maybe, like, even just to go one step deeper on each of the aspects that Anshul talked about, let's go with the API aspect. So right now, I'll give you an example. Super Complete is actually a feature that I think is, like, very exciting about the product, right?

It can suggest refactors of the code. I think it can do it quickly and, and very powerfully. On VS Code, actually, the problem for us wasn't actually being able to implement the feature. We had the feature for a while.

Problem was actually even to show the feature, VS Code would not expose an API for us to do this. So what we actually ended up doing was dynamically generating PNGs to actually go out and showcase this. It was not really aligned.

We actually ended up doing it ourselves, and it took us a couple hours to actually go out and implement this, right? And that wasn't because we were bad engineers. No. Our good engineering time was being spent fighting against the system rather than being a good system.

Another example is we needed to go out and find ways to refactor the code. The VS Code API would constantly keep breaking on us. We'd constantly need to show a worse and worse experience. This actually comes down to the second point which Anshul brought up, which is like, we can come up with great work and great research.

All the work we have here is not like-- The research on Cascade is not like a couple month thing. This is like a nine months to a year thing that we've been investigating as a company. Investing in on evals, right?

Even the evals for this are, are a lot of effort, right? A lot of actually systems work to actually go out and do it. But ultimately, like, this needs to be a product that developers actually use. And I think, you know, let's even go for Cascade, for example.

Like, and looking at the trajectory, yeah, we'd like to see-

**Alessio** [9:14]
And can you define Cascade? Because that's the first time you brought it up.

**Varun Mohan** [9:16]
Yeah. So Cascade is the product that is the actual agentic part of the product, right? Um, that is capable of, of taking information from both these human trajectories and these AI trajectories, what the human ended up doing, what the AI ended up doing, to actually propose changes and actually execute code to finally get you the final work output, right?

I'll even talk about something very basic. Cascade gives you a bunch of code. We want developers to very easily be able to review this code. Okay. Then we can show developers a hideous UI that they don't wanna look at, and no one's gonna really use this product.

And we think that this is, like, a fundamental building block for us to make the product materially better. If people are not even willing to use the building block, where does this go? Right? And we just felt our ceiling was capped on what we could deliver in ex- as an experience.

Interestingly, JetBrains is a much more configurable paradigm than, than VS Code is. But we just felt so limited on both the, both the sort of directions that Anshul said, that we were just like, "Hey, if we actually remove these limitations, we can move substantially faster."

And we believe that, that this was a necessary step for us.

### Evals

**Alessio** [10:13]
I'm curious more about the evals side of it because you brought it up.

**Varun Mohan** [10:16]
Yeah.

**Alessio** [10:17]
And w- we have to ask about evals any time anyone brings up evals.

**Varun Mohan** [10:20]
Yeah.

**Anshul** [10:20]
Mm-hmm.

**Alessio** [10:20]
How do you evaluate a thing like this that is so multi-stepped and so-

**Varun Mohan** [10:24]
Yeah

**Alessio** [10:25]
... spanning, like, so much context?

**Varun Mohan** [10:26]
So what you can imagine, we can, we can sort of do, and this is, like, one of the beautiful things about code, is code can be executed. We could go take a bunch of open source code, we can find a bunch of commits, right?

And we can actually see if some of these commits have tests a-associated with them. We can start stripping the commits, and the, the approach of stripping the commits is good because it tests the fact that the code is in an incomplete state, right?

When you're writing the commit, the goal is not the commit has already been written for you. You're given it in a state that where the entire thing has not been written. And can we go out and actually retrieve the right snippets and actually come up with a cohesive plan and iterative loop that gets you to a state where the code actually passes?

So you can actually break down and, and decompose this complex pro- problem into, like, a planning, retrieval, and multi-step execution problem. And you can see on every single one of these axes is getting better. And if you do this across enough repositories, you've turned this highly discontinuous and discrete problem of make a PR work versus make it not work into a continuous problem.

And now that's a hill you can actually climb, and that's a way that you can actually apply research where it's like, hey, my retrieval got way better. This made my eval get better, right? And then notice how the way the eval works is I'm not that interested in the eval where purely it's a commit message, and you finish the entire thing.

I'm more interested in the code is in an incomplete state, and the commit message isn't even given to you. Because that's another thing about developers. They are not willing to tell you exactly what's in their head. That's the actual important piece of this problem.

We believe that developers will never completely pose the problem statement, right? Because the problem statement lives in their head, conversations that you and I have had at the coffee area, conversations that I've had over Slack, conversations I've had over Jira, right?

Maybe not Jira. Let's say Linear, right? That's, that's the cool thing nowadays.

**Anshul** [12:03]
We're talking about Jira.

**Varun Mohan** [12:03]
Yeah. So, so conversations I've had, I've had on Linear, and all of these things come together to actually finally propose sort of a, a solution there, which is why we wanna test the incomplete code. What happens if the state is in an incomplete state, and, and am I actually able to make this pass without the commit?

And can I actually guess your commit well? Now you can convert the problem into a mask prediction problem, where you want to guess both the high-level intent and as well as the remainder of changes to make the actual test pass.

And you can imagine if you build up all of these, now you can see, "Hey, my systems are getting better, retrieval quality is getting better," and you can actually start testing this on larger and larger code bases, right?

And I guess that's one thing that we, we-- honestly, to be honest, we could have done a little faster. We had the technology to go out and build these zero-to-one apps very quickly, and I think people are using Windsurf to actually do that, and it's, like, extremely impressive.

But the real value, I think, is actually much deeper than that. It's actually that you take a large code base, and it's actually a really good first pass. And I'm not saying it's perfect, but it's only gonna keep getting better and, and-- or, you know, we have, like, deep sort of infrastructure to that actually that is validating that we are getting better on this dimension.

**Anshul** [13:02]
Varun mentioned the, the end-to-end evals that we have for the system, which I think are, like, super cool. But I think you can even decompose each of those steps, right? The, the ideas of-- Just take retrieval, for example, right?

Like, how can we make eval for retrieval really good? And I think this is just a general thing that's been true about us as a company is, like, most evals and benchmarks that exist out there for software development is kind of bogus.

Th-there's not really a better way of, of putting it. Like, okay, you have SWE-bench, that's cool. No actual professional work looks like SWE-bench. Like, human eval, same thing. Like, these things are just, like, a little kind of broken, so when you're trying to optimize against a metric that's a little bit broken, you end up making kind of suboptimal decisions.

So something that we're always very keen on is like, okay, what is the actual metric that we wanna test for this part of the system? And so take retrieval, for example. A lot of the benchmarks for these embedding-based systems are, like, needle in the haystack problems.

Like, I wanna find this one particular piece of information out of all this potential context. That's not really what actually is necessary for doing software engineering, because code is a super distributed knowledge store. You actually wanna pull in snippets from a lot of different parts of the code base in order to do the work, right?

And so, you know, we built systems that instead of looking at, you know, retrieval at one, you're looking at retrieval at, like, fifty. Like, ping, what are the fifty highest things that you can actually retrieve, and are you capturing all of the necessary pieces for that?

And, and what are all the necessary pieces? Well, you can look again back at, at old commits and see what were all the different files that together were edited to make a commit. Because those are semantically similar things that might not actually show if you actually try to map out a code graph, right?

And so we can actually build these kind of golden sets. We can do this evaluation even for sub-problems within the overall task. And so now we have, like, you know, an engineering team that can iterate on all of these things and still make sure that the end goal that we're trying to build to is, like, really, really strong so that we have confidence in what we're pushing out.

**Varun Mohan** [14:44]
And by the way, just to g-talk, say one more thing about the SWE-bench thing, just to showcase these existing metrics, I think benchmarks are not a bad thing. You do want benchmarks. Actually, like, I would prefer if there are benchmarks versus, let's say, everything was just vibes, right?

But vibes are also very important, by the way, because they showcase that the-- where the benchmark is not valuable, because actually vibes sometimes show you where criminal issues are sort of exist in the benchmark. But, like, you look at some of the ways in which people have, like, optimized SWE-bench, it's like, make sure to run pytest every time X happens.

And it's like, yeah, like, sure, you can start, like, prompting it in, like, every single possible way, and, like, if you remove that, suddenly it doesn't get good at it. It's like, what really matters here? What really matters here is, like, across a broad set of tasks, you're performing, like, high quality sort of suggestions for people, and people love using the product.

And I think actually, like, the way these things work is beyond a certain point-- Because, yes, I actually think there-- it's valuable beyond a certain point, but once it starts hitting the peak of these benchmarks, getting that last ten percent actually probably is, like, counterintuitive to the actual goal of what the benchmark was.

Like, you probably should find a new hill to climb rather than sort of p-hacking or really optimizing for how you can get higher on the benchmark.

**Alessio** [15:49]
Yeah, we did an episode with Anthropic about their recent, like, SWE-agent, SWE-bench, uh, results, and we talked about the human eval versus SWE-bench and-- Or, like, human eval is kind of like a greenfield benchmark, you know? You need to be good at that.

SWE-bench is more existing. But it sounds like-- I mean, your eval creation is similar to SWE-bench as far as, like, using GitHub commits and kind of, like, the history, but then it's more, like, masking at the commit level-

**Varun Mohan** [16:12]
Yes

**Alessio** [16:12]
... versus just testing the output-

**Varun Mohan** [16:14]
That's right

**Alessio** [16:14]
... of the, of the thing.

### Launch & Remote

**Swyx** [16:16]
Cool. W-w-we have some listener questions actually about the Windsurf launch, and, uh, obviously I also wanna give you the chance to just respond to Hacker News.

**Anshul** [16:23]
Oh, man.

**Varun Mohan** [16:24]
Oh, man.

**Anshul** [16:25]
Don't make us do that.

**Varun Mohan** [16:25]
Hey, let me, let me tell you something very, very interesting.

**Swyx** [16:28]
Yeah.

**Varun Mohan** [16:28]
Uh, I love Hacker News as much as the next, next person, but the moment we launched our product, the first comment, uh, like, this was a year ago, the first comment was, "This product is a virus." And we were like-

**Anshul** [16:38]
This was the original Codeium launch, like, two years ago.

**Varun Mohan** [16:41]
This is the original.

**Anshul** [16:41]
Like, "I am analyzing the bi-binary as we speak. We'll report back."

**Varun Mohan** [16:45]
And then he's like-

**Anshul** [16:46]
We'll report back

**Varun Mohan** [16:46]
... he was like, "It's a virus." And I was like, "Dude, like, it's not a virus."

**Anshul** [16:51]
We just wanna give autocomplete suggestions. That's all we wanna do.

**Varun Mohan** [16:54]
Yeah.

**Anshul** [16:55]
Okay, Varun.

**Varun Mohan** [16:56]
O-okay.

**Swyx** [16:56]
That was fun.

**Varun Mohan** [16:56]
Wow, I didn't, I didn't expect that.

**Swyx** [16:57]
And then there was, like, Theo drama. There's enough drama-

**Varun Mohan** [17:00]
Oh, okay

**Swyx** [17:00]
... on the launch to cover, but I don't know if we wanna just-

**Anshul** [17:03]
Go ahead

**Swyx** [17:03]
... make this a Cascade piece. But we had a bunch of people in our Discord try out the product, give a lot of feedback. One question people have is, like, to them, Cascade already felt pretty agentic. Like, is that something you wanna do more of?

You know, obviously, since you just launched an IDE, you're kind of like, you're focusing on having people write the code, but maybe this is kind of like the Trojan horse to just doing more full-on end-to-end, like, code creation.

Devin style.

**Anshul** [17:27]
Yeah.

**Swyx** [17:28]
Yeah.

**Anshul** [17:28]
I, I think it's like how, how do you get there in a, in a, like, a real principled manner, right? We have obviously, like, enterprise asking us all the time like, "Oh, when's it gonna do, like, end-to-end work?"

The reality is like, okay, well, if we have something in the IDE that, again, can, like, see your entire actions and get a lot of intent that you can't actually get if you're not in the IDE, if the agent there has to always get human involvement to keep on fixing itself, it's probably not ready to become a full end-to-end automated system, 'cause then we're just gonna turn into a linter where like-

**Swyx** [17:55]
Right. Yeah, yeah

**Anshul** [17:56]
... it produces a bunch of things and no one looks at any of it. Like, that's not, that's not the great end state. But if we start seeing like, oh, yeah, there's common patterns that people do that, like, never require human involvement, it just end-to-end just totally works without, like, any, like, intent-based information, sure, that can become, like, fully agentic and, like, we will learn what those tasks are, like, pretty quickly 'cause we have a lot of data.

**Varun Mohan** [18:15]
Maybe add onto that, I think that if, if the answer is, like, full agentic is called like, is, is Devin, I think, like, yes, the answer is this product should become fully agentic and limited human interaction is the goal.

Is 100% the goal. And I think honestly, of all usable products right now, I think we're the closest right now. Of all usable products in an IDE. Now, let me caveat this by saying, I think there are lots of hard problems that have yet to be solved that we need to go out and solve to actually make this happen.

Like for instance, I think one of the most annoying parts about the product is the fact that you need to accept, um, every command that kind of gets run. It's actually fairly annoying. I would like it to go out and run it.

Unfortunately, me going out and running arbitrary binaries has some problems in that if it like rm -rf my hard disk-

**Swyx** [18:57]
It's actually a virus

**Varun Mohan** [18:58]
... I'm not gonna be-- I'm not-

**Swyx** [18:59]
It's a virus.

**Varun Mohan** [18:59]
The hacker needs to be one you. Yeah, it does, it does become a virus. Um, I think this is, this is solvable with like, with complex systems. I think w- we love working on complex systems infrastructure. I think we'll solve it.

Now, the simpler way to go about solving this is don't run it on the user's machine and run it somewhere else, because then if you bork that machine, you're kind of totally fine. Now, I think, I think though, maybe there's a little bit of trade-off of like running it locally versus remotely, and, and I think we might change our mind on this.

But I think the goal for this is not for this to be the final state. I think the goal for this is, A, it's actually able to do very complex tasks with limited human interaction, but it needs to know when to actually go back to the human, right?

Also on top of that, compress every cycle that the agent is running. Right now, actually, I even feel like the product is too slow for me sometimes right now, even with it running really fast.

**Swyx** [19:44]
Yeah, yeah.

**Varun Mohan** [19:44]
It's objectively pretty fast.

**Swyx** [19:45]
Too, too cracked.

**Varun Mohan** [19:46]
I would still want it to be faster, right? So there is like systems work and probably modeling work that needs to happen there to make the product even faster on both the retrieval side and the generation side, right?

And then finally speaking, I think another key piece here that's like really important is I actually think asking people to do things explicitly is probably going to be more of an anti-pattern if we can actually go and passively suggest the entire change for the user.

So almost imagine as the user is using the product, uh, that we're gonna suggest the remainder of the PR without the user kind of like even, even asking us for it. I think this is sort of the beginning for it, but yeah, like these are hard problems.

I can't give a particular deadline for this. I think this is like a big step up than what we had, particularly in the past. But I, I think what Anshul said is 100% true, but the goal is for us to get better at this.

**Swyx** [20:30]
I mean, the remote execution thing is interesting. You, you wrote a post about the end of localhost.

**Varun Mohan** [20:35]
Yeah. Uh-

**Swyx** [20:35]
Now it's almost like then we were kind of like, "Well, no, maybe we do need the internet," and like people wanna run things, but now it's like, "Okay, no, actually, I don't really care. Like, I want the model to do the thing."

And if you were like, you can do a task end-to-end, but it needs to run remotely, not on your computer, I'm sure most people would say, "Yeah."

**Varun Mohan** [20:50]
No, I agree with that.

**Swyx** [20:51]
I'm cool with it.

**Varun Mohan** [20:52]
I actually agree with it running remotely. That's not a security issue. I actually-- I totally agree with you that it's possible that everything could, could, uh, run remotely. Um, I do-

**Swyx** [20:59]
That's how it is at, at most com-- most like big codes, like Facebook.

**Varun Mohan** [21:02]
Yeah.

**Swyx** [21:02]
Like, nobody runs things locally.

**Varun Mohan** [21:04]
No one does. In fact, you connect to a remote machine-

**Swyx** [21:06]
SSH into the mainframe.

**Varun Mohan** [21:08]
You're right on that. Maybe the one thing that I, I do think is kind of important for these systems that is more than just running remotely is basically like, you know, when you look at these agents, there's kind of like a rollout of a trajectory, and I kinda wanna roll this trajectory back, right?

In some ways, I want like a snapshot of the system that I can like constantly checkpoint and move back and forth. And then also on top of that, I might wanna do multiple rollouts of this. So basically, I think there needs to be a way to, to almost like move forward and move backwards the system.

And whether that's locally or remotely, I think that's necessary. But every time if you move the system forward, it like destroys your machine. It's probably gonna be a hard system to kind of-- or potentially destroys your machine. That's just not a workable solution.

So I think the local versus remote, I think you still need to solve the problem of this thing is not gonna destroy your machine on every execution, if that makes sense.

**Swyx** [21:52]
Yeah.

**Varun Mohan** [21:53]
Yeah.

**Swyx** [21:53]
Uh, I, there is a category of emerging infrastructure providers that are working on time travel VMs-

**Varun Mohan** [21:58]
Yeah

**Swyx** [21:58]
... where you can kind of-

**Anshul** [21:59]
And if, and if Varun's first episode on, on this podcast was any indication, we like infrastructure problems.

**Swyx** [22:03]
Yeah. Okay. All right. Oh, so you're going there? All right. Okay.

**Varun Mohan** [22:06]
Well, that's funny, right? It's like when we first had you, you were doing so much on like actual model inference, optimization, all these things, and today it's almost like-

**Swyx** [22:14]
It's Claude, it's 4o.

**Varun Mohan** [22:16]
It, it's like, you know, people are like forgetting about the model, you know-

**Swyx** [22:19]
Yeah

**Varun Mohan** [22:19]
... now it's all about at a higher level-

**Swyx** [22:21]
Yes

**Varun Mohan** [22:21]
... of abstraction. Yeah. So maybe I can te- say like a little bit about how our strategy on this has like evolved because it, it objectively has, right? I think I would be, I would be lying if I said, if I said it hasn't.

The things like autocomplete and super complete that run on every keystroke are entirely like our own models. And by the way, that is still because properties like FIM, uh, fill-in-the-middle capabilities are still quite bad with the current, um, Claude models.

**Swyx** [22:42]
Nonexistent.

**Varun Mohan** [22:42]
Uh, they're all-- They're very bad. Nonexistent. They're not good, um, actually at it. Um-

**Swyx** [22:46]
'Cause FIM is an actual like how you order the tokens.

**Varun Mohan** [22:49]
Yes. It's how you order the tokens actually, uh, in some ways. And this is a-- this is sort of-- If you look at what these products have sort of become, and this is great, is a lot of the Claudes and the OpenAIs have focused on kind of the chat-like assistant API, where it's like complete pieces of work message, another complete piece of work.

So multi-turn kind of back-and-forth, uh, systems. In fact, like actually, even these systems are not that good at making point changes. When they make point changes, they kind of are like off here and there by a little bit.

Uh, because yeah, when you wr- when you like are doing multi-point kind of like conversations, it's, you know, exact gifts getting applied is not like even a perfect science still yet. So we care about that. The second piece where we've actually sort of trained our models is actually on the retrieval system.

And this is not even for embedding, but like actually being able to use high-powered LLMs to be able to do much higher quality retrieval across the code base, right? So this is actually what Anshul said. For a lot of the systems, we do believe embeddings work, but for complex questions, we don't believe embeddings can encapsulate all the granularity of a particular query.

Like imagine, imagine I have a question of a-- on a code base of, "Find me all quadratic time algorithms in this code base." Do we genuinely believe the embedding can encapsulate the fact that this function is a quadratic time function?

No, I don't think it does. So you are gonna get extremely poor precision recall at this task. So we need to apply something a little more high-powered, uh, to actually go out and do that. So we've actually built like large distributed systems to actually go out and, and run these at scale, run custom models at scale across large code bases.

So I think it's more a question of that. The planning models right now, undoubtedly, I think, I think the Claudes and the OpenAIs have the best products. I think Llama Four, depending on where it goes, it could be materially better.

It's very clear that they're willing to invest a similar amount of compute as the OpenAIs and the Anthropics. So we'll see. I would be very happy if they got, uh, really good, but unclear so far.

**Swyx** [24:32]
Don't forget Groq

**Varun Mohan** [24:33]
Hey, dude, I think Groq is also possible.

**Swyx** [24:35]
Yeah.

**Varun Mohan** [24:35]
Right? I think don't, don't doubt Elon. Yeah.

**Swyx** [24:38]
Okay, so I didn't actually know-- it, it's not obvious when I use Cascade. I, I should also mention that, you know, I was part of the preview-

**Varun Mohan** [24:44]
Yeah

**Swyx** [24:44]
... and thanks for letting me in, and I, I've been maining Windsurf, uh, for a long time. It's not actually obvious, you don't make it obvious that you are running your own models.

**Varun Mohan** [24:51]
Yeah.

**Swyx** [24:51]
And I feel like you should, so that, like, I feel like it has more differentiation. Like, I only have-- I have exclusive access to your models via your IDE than having the dropdown that says Cloud and 4o, 'cause I actually thought that was what you did.

**Varun Mohan** [25:04]
No, so actually the way it works is the high-level planning that is going on in the model-

**Swyx** [25:08]
Yeah

**Varun Mohan** [25:08]
... is actually getting done with products like the Cloud. But the extremely fast retrieval, as well as the ability to, like, take the high level plan and actually apply it to the code base is proprietary systems that are running internally.

### Cascade Strategy

**Swyx** [25:18]
Yeah. And then the, uh, the, the stuff that you said about, uh, embeddings not being enough, are you familiar with the, like, concept of late interaction?

**Varun Mohan** [25:24]
No, actually I've never heard of it.

**Swyx** [25:25]
Uh, yeah, so this is, um, ColBERT or, like, the guy, uh, Omar Catm from, I think, Stanford, has been promoting this a lot. It, it is basically what you've done.

**Varun Mohan** [25:33]
Okay.

**Swyx** [25:33]
Uh, so, um, sort of embedding on, on retrieval rather than pre-embedding-

**Varun Mohan** [25:37]
Okay. That's-

**Swyx** [25:38]
... in a very loose sense.

**Varun Mohan** [25:39]
I think, I think that sounds like a very good idea that is very similar to what we're doing.

**Swyx** [25:43]
Sounds like a very good idea.

**Varun Mohan** [25:44]
Yes.

**Swyx** [25:44]
I don't think we'd say that.

**Varun Mohan** [25:45]
That's like, that's like the meme of Obama giving himself a medal right there.

**Swyx** [25:50]
Yeah. No, I mean, there might be con- like in-- there might be something to learn from contrasting the ideas and seeing where-

**Varun Mohan** [25:54]
No, absolutely

**Swyx** [25:55]
... like the subtle opinion differences. It's also been applied very effectively to vision understanding because vision models tend to just consume the whole image. Uh, if you are able to sort of focus on images based on the query, I think that, that, uh, can get you a lot, a lot of extra performance.

**Anshul** [26:10]
The, the basic idea of using compute in a distributed manner to do, you know, operations over a whole set of like raw data rather than like a-

**Swyx** [26:17]
Yeah

**Anshul** [26:17]
... materialized view is not anything new, right? Like, I think it's just like how does that look like for LLMs?

**Swyx** [26:22]
When I hear you say like build large distributed systems, like you have a very strange product strategy of going to down to the individual developer, but also to the large enterprise.

**Varun Mohan** [26:30]
Yeah.

**Swyx** [26:30]
Is it the same infra that serves everything?

**Varun Mohan** [26:32]
I think the answer to that is yes. The answer to that is yes. And the only reason why for the yes, the answer is yes, and to be honest, our company is a lot more complex than I think if we just wanted to serve the individual.

And I'll tell you that because we don't really like pay other providers to do things for our indexing. We don't pay like other providers to do our, our serving of our own customer models, right? And I think that's a core competency within our company that we have decided to build.

But that's also enabled us to go and like make sure that when we're serving these products in an environment that works for these large enterprises, we're not going out and being like, "We need to build this custom system for you guys."

Right. This is the same system that serves our entire user base. So that is a very unique decision we've taken as a company, and we admit that there are probably faster ways that we could have done this.

**Swyx** [27:15]
I was thinking, you know, when, when I was working with you for your enterprise piece, I was thinking like this philosophy of go slow to go fast, like build deliberately for the right level of abstraction that can serve the market that you really are going after.

**Anshul** [27:26]
Yeah. I mean, I, I would say like I'm-- I was writing-- when writing that piece, you know, like looking back and reading it back, it sounds so like almost obvious in hindsight. Not all of those are really conscious decisions we made.

Like, I'll be the first to admit that, but like, it does help, right? When we go to like an enterprise that has tens of thousands of developers and they're like, "Oh wow, like you know, we have tens of thousands of developers and like does your infrastructure work for tens of thousands of developers?"

We can turn around and be like, "Well, we have hundreds of thousands of developers on an individual plan that we're serving. Like I, I think we'll be able to support you," right? So like being able to do those things, like we started off by just like, let's give it to individuals, let's see what people like and what they don't like and learn, but then those become value propositions when we go to the enterprise.

**Swyx** [28:03]
And to recap, when you first came on the pod, it was like auto-completion is free and Copilot was ten bucks a month. And you said, "Look, what we care about is building things on top of code completion." How did you decide to like just not focus on like short-term kind of like growth monetization of like the individual developer and like build some of this?

Because the alternative would've been, "Hey, all these people are using it." It's like, we're gonna make this other like five bucks a month plan-

**Varun Mohan** [28:28]
Yeah

**Swyx** [28:28]
... monetize.

**Varun Mohan** [28:30]
I think, I think this might be a little bit of like commercial instinct that the company has, and unclear if the commercial instinct is right. I think that right now optimizing for making money off of individual developers is probably the wrong actually strategy.

Largely because I think individual developers can switch off of products like very quickly, and unless we have like a very large lead trying to optimize for making a lot of profit off of individual developers, it's probably something that someone else could just vaporize very quickly and then, and then they move, they move to another product.

And I'm gonna say this very honestly, right? Like when you use a product like Codeium on the individual, on the individual side, there's not much thing-- not much that prevents you to switch onto another product. I think that will change with time as the products get better and better and deeper and deeper.

I constantly say this, like there's a book in business called like Seven Powers, and I think one of the powers that a business like ours need to have is like real switching costs. But like you first need something in the product that makes people switch on and stay on before you think about how do you make people switch off.

And I think for us, we believe that there's probably much more differentiation we can derive in the enterprise by working with these large companies in a way that is like, that is interesting and scalable for them. Like, I'll be maybe more concrete here.

Individual developers are much more sort of tuned towards small price changes. They care a lot more, right? Like if our product is ten, twenty bucks a month instead of fifty or a hundred bucks a month, uh, that matters to them a lot.

I think for a large company where they're already spending billions of dollars on software, this is much less important. So you can actually solve maybe deeper problems for them, and you can actually kind of provide more differentiation on that angle.

Whereas I think, I think individual developers could be churny as long as we don't have the best product. So focus on being the best product, not trying to like take price and make a lot of money off of people.

Um, and I, I don't think we will, uh, for the foreseeable future try to be a company that tries to make a lot of money off individual developers.

**Swyx** [30:16]
I mean, that makes sense. Uh, so why ten dollars a month for Windsurf?

**Varun Mohan** [30:19]
Why $10?

**Anshul** [30:20]
$10 a month was actually the, the pro plan. So we, we launched our individual pro plan before Windsurf existed, 'cause I think there's... Let's, let's all try. We all said to be financially responsible as a company. Right? And so-

**Varun Mohan** [30:30]
Yeah, yeah

**Anshul** [30:31]
... there, there comes a point-

**Varun Mohan** [30:31]
We can't, we can't run out of money.

**Anshul** [30:33]
There, there, there-

**Swyx** [30:33]
You raised $150 million. Like, you're good.

**Anshul** [30:34]
It's a cool... Like, I mean, like, there's a lot of things, 'cause of our infrastructure background, we can give, like, for essentially free, like unlimited auto-complete, you know, unlimited chat on like our, our, you know, faster models, like unlimited...

Like, we give a lot of things actually out for free. But yeah, when we start doing things like the super completes and really large amounts of indexing and all of these things, like there, there is real cogs here.

Like, we can't ignore that. And so we just created a $10 a, a month pro plan mostly just to cover the cost. Like, we're not really, like, operating, I think, on a much of a margin there either, but like, okay.

Like, just, just to cover us there. So for Windsurf, it, it just ended up being the same thing, and everyone who downloads Windsurf in the first, like, I, I forget, like, a couple of weeks, they get like two weeks for free.

Let's just have people try it out, let us know what they like, what they don't like, and that's how we've always operated.

**Swyx** [31:16]
I've talked to a lot of CTOs in like the Fortune 100, where most of the engineers they have, they don't really do much anyway. The problem is not that the developer costs 200K and you're saving 8K. It's like that developer should not be paid 200K.

**Varun Mohan** [31:30]
Mm.

**Swyx** [31:31]
But that's kinda like the base price, you know? But then you have developers getting paid 200K that should be paid 500K. So it's almost like you're averaging out the price because most people are actually not that productive anyway, so if you made them 20% more productive, they're still not very productive.

And I don't know, in the future, like, is it that the junior developer salary is like 50K, you know? And it's like the, the bottom of the end gets kinda like squeezed out, and then the top end-

**Varun Mohan** [31:56]
May-

**Swyx** [31:56]
... gets squeezed up.

**Varun Mohan** [31:57]
Yeah, maybe, Alessio, one thing that I think about a lot, because I do think about this, the per seat, anything. Every- all of this stuff I think about a good deal. Let's take a product like Office 365. I will say a lawyer at Codeium uses Microsoft Word way more than I do.

I'm still footing the same bill, but the amount of value that he's driving from Office 365 is probably, you know, tens of thousands of dollars. By the way, everyone, you know, Google Doc's a great product. Microsoft Word is a crazy product.

They made it so that the moment you review anything in Microsoft Word, the only way you can review it is with other people in Microsoft Word. So it's like this virus that penetrates everything, and it's not only penetrates it within the company, it penetrates it cross-company too.

The amount of value it's driving is way higher for him. So there w- for these kinds of products, there's always going to be, for these kinds of products, this variance between who gets value from these products, right? And you're right, it's, it's almost like a blended, 'cause you're actually totally right.

Probably this company should be paying that one developer maybe like four times as much. But in weird way, software is like this team activity enough that there's a bunch of blended outcomes, but hey, like 20% of the four, four times and there are four people is still gonna cover the cost across the four individuals, right?

And that's how roughly these products kind of get priced out.

**Swyx** [33:02]
I mean, more than about pricing, this is about like the future of like the software engineer. Like...

**Varun Mohan** [33:06]
We could be very wrong also.

**Anshul** [33:08]
Yeah. We're-

**Swyx** [33:09]
I, I think nobody knows.

**Anshul** [33:10]
Reserve-

**Swyx** [33:10]
So-

**Anshul** [33:10]
... the right to be incredibly off.

**Varun Mohan** [33:12]
Yeah.

### Agents & Fixes

**Swyx** [33:13]
I mean, business model does impact the product, product does impact the, you know, user experience, so it's all, it's all of a kind. I, I don't mind. We are, we do... are as concerned about the business of tech as the tech itself.

**Varun Mohan** [33:23]
That's cool.

**Swyx** [33:23]
Speaking of which, there's other listener questions. Uh, shout-out to Daniel Imfeld, who, who's pretty active in our Discord, just asking all these things. Uh, multi-agent, very, very hot and popular, especially from like the Microsoft Research point of view.

Have you made any explorations there?

**Varun Mohan** [33:37]
I think we have. I don't think we've called it a multi-agent, which is more so like once you-- this notion of having many trajectories that you can spawn off, um, that kind of like validate sort of some different hypotheses and you can kind of pick the most interesting one.

This is stuff that we've actually analyzed, uh, internally at the company. By the way, the reason why we have not put these things in, actually, is partially because we can't go out and execute some random stuff in parallel in the meantime, uh, in the meantime on other sides.

**Swyx** [34:02]
Because of the side effects.

**Varun Mohan** [34:03]
Because of the side effects, right? Um, so there are, there's some things that are a little bit dependent on us, u- us unlocking more and more functionality internally. And then the other thing is, in the short term, I think there is like also a latency component, and I think all of these things can kind of be solved.

I actually believe all of these things are solvable problem. They're not unsolvable problem. And then if you wanna run all of them in parallel, you probably don't want end machines to go out and do it. I think that's unnecessary, especially if most of them are IO bound kind of operations where all you're doing is reading a little bit of data and writing out a little bit of data.

It's not extremely compute intensive. I think that it's a, it's a good idea and probably something we will pursue and is gonna be in the product.

**Swyx** [34:37]
I'm still processing what you just said about things being IO bound, so you can-- so for a certain class of concurrency, you can actually just run it all on one machine.

**Varun Mohan** [34:44]
Yeah. Why not? Because if you look at, if you look at the changes that are made, right, in, for some of these, it's writing out like, what? A couple thousand bytes? Maybe like tens of thousands of bytes on every tr- It's not a lot.

Very small.

**Swyx** [34:55]
What's next for Cascade or, or Windsurf?

**Anshul** [34:57]
Oh, there's a lot. I don't know. I'd, I'd like... We did an internal poll and we were just like, "Are you more excited about this launch or, or the launch that's happening in a month? Or like what we're gonna come out with in a month?"

And it was like almost uniformly in a month. I think like, you know, there's, there's some like obvious ones. I don't know how much, Varun, you wanna say. I don't wanna to speak of, but I, I think you'd look at all the same axes of the, of the system, right?

Like, how can we improve the knowledge retrieval? Like, we'll, we'll always keep on figuring out how to improve knowledge retrieval. In the, in our launch video, we even showed some of like the early explorations we have about looking to other data sources.

That might not be the coolest thing to the individual developer building a zero to one app, but you can really believe that like the enterprise customers really think that that's very cool, right? I think, um, on the tool side, I think there's a whole lot more that we can do.

I mean, of course, I mean Varun's talked about not just suggesting the terminal command, but actually executing them. Like, I think that's gonna be a huge unlock. You look at the, the actions that people are taking, right? Like the human actions, the trajectories that we can build, like how can we make that even more detailed?

And if you take all of those things and you, you make some like even a cleaner U- UI, like the idea of looking at future trajectories, trying a few different things, and like suggesting potential next-

**Varun Mohan** [36:01]
Yeah

**Anshul** [36:01]
... like actions to be taken.

**Varun Mohan** [36:03]
Yeah, yeah, yeah. That's it.

**Anshul** [36:03]
That doesn't really exist yet, but like it's pretty obvious, I think, how that would look like, right? You open up Cascade and instead of like starting typing, it's just like, "Here's a bunch of things that we wanna do."

We kind of joked it's like Clippy's coming back, but like maybe now's the time for Clippy to really shine, right? So I, I think there's, there's a lot of ways that we can take this, which I think is like the very exciting part.

We're calling each of our launches Waves, I believe, because we wanna really double down on the aquatic themes.

**Swyx** [36:26]
Oh yeah. Does, does someone actually windsurf at the company? Is it-

**Anshul** [36:28]
I don't think so.

**Swyx** [36:30]
We're living out, we're living out our dream of being cool enough to windsurf through the products.

**Anshul** [36:34]
Yeah. Yeah, no, I don't, I don't think-

**Swyx** [36:34]
Okay. I'll say-

**Anshul** [36:34]
... I don't think we can.

**Swyx** [36:35]
Yeah. All right. Well

**Anshul** [36:36]
That, that was actually something we learned, uh, 'cause I don't think any of us are windsurfers. Like, we... Like, in our launch video, we have someone, like, using Windsurf on a windsurf. Uh, uh, that was like-

**Swyx** [36:43]
You saw that?

**Anshul** [36:44]
You saw that. In the beginning of the video-

**Swyx** [36:45]
Oh, yeah, yeah

**Anshul** [36:45]
... someone's at the computer. And we didn't realize, like, now apparently is, like, the time of the year where there's, like, not enough wind to windsurf, so we were trying to figure out how to do this, like, you know, launch video with a windsurf on the windsurf.

And, like, every windsurfer we talked to were like, "Yeah, it's not possible." And there was, like, one crazy guy who was like, "Yeah, I think we can do this." And, uh, we made it happen.

**Swyx** [37:02]
Oh, okay.

**Anshul** [37:03]
That's funny.

**Swyx** [37:04]
Oh. Is there anything that you want feedback on? Like, uh, maybe there's a fork in the road, you want feedback, you want people to respond to this podcast and tell you. What do you want?

**Varun Mohan** [37:12]
Yeah, I think, I think there's a lot of things that I think could be more polished about the product that we'd like to, to improve. Um, lots of different environments that we're gonna improve performance on, and I think we would love to hear, uh, from folks, uh, across the gamut, like, hey, like, if you have this environment, you use Windows and X version, it didn't work, or this language-

**Swyx** [37:30]
Oh, yeah

**Varun Mohan** [37:30]
... it was, like, very poor. I think we would like to hear it. Um, but also-

**Swyx** [37:33]
Yeah, I gave, I gave Prep and Kevin a lot of shit for my Python issues.

**Varun Mohan** [37:37]
Yeah. Yeah, yeah. And I think there's a lot to kind of improve on the environment side. I think, like, for instance, even just a, a dumb example, and I think, uh, sort of Swix, this was a common one, is like, yeah, like, the virtual environment, where is the terminal running?

What is all this stuff? These are all basic things that, like, to be honest, this is not rocket science, but we need to just fix it, right? Uh, we need to fix it. So, uh, we would love to hear, like, all the feedback from the product.

Like, was it too slow? Where was it too slow? Uh, what kind of environments could it work way more in? There's a lot of things that we don't know. We... Luckily, we're daily users of the product and, uh, internally, so we're getting a lot of feedback inside.

But I will say, like, there's a little bit of Silicon Valeyism, in that a lot of us develop on Mac. A lot of people, once again, over s- 80% of developers are on Windows. So yeah, there's a lot to learn and probably a lot of improvements down the line.

**Swyx** [38:19]
Are you personally tempted, as, uh, you're CEO of the company, to switch to Windows just to feel something?

**Anshul** [38:25]
Feel the Windows.

**Varun Mohan** [38:26]
You know, you know what?

**Swyx** [38:27]
Feel the pain.

**Varun Mohan** [38:28]
You know what? Maybe I should. Actually, actually, I think I-

**Swyx** [38:31]
Yeah, that's a good idea

**Varun Mohan** [38:31]
... I think I will.

**Swyx** [38:31]
I mean, like, your customers, you know, they... Yeah, I-

**Varun Mohan** [38:33]
No, you're right

**Swyx** [38:33]
... everyone says it's 80, 80, 90% on Windows, right? Like you said, you live in Windows, you will never, you will never not see something that-

**Varun Mohan** [38:40]
Yeah

**Swyx** [38:40]
... missed.

**Varun Mohan** [38:41]
So I think in the beginning, part of the reason why we, we were hesitant to do that was, like, a lot of our architectural decisions to work on across every IDE was because we built a platform-agnostic way of running the system on the user's local machine that was only buildable, uh, easily buildable on, like, on dev containers that la- that lived on a particular type of platform, so Mac was, like, nice for that.

But now there's, like, not really an excuse if it's like, if I can also make changes to the-

**Swyx** [39:05]
Sure

**Varun Mohan** [39:05]
... to the UI and stuff like that. And yeah, WSL also exists. That's actually something that we need to add to the product. That's how early it is, uh, that we have not actually added that. So-

**Anshul** [39:12]
We don't have, like, remote.

### Enterprise

**Swyx** [39:13]
Anything else about Codeium at large, right? Like, you still have your core business of the, uh, enterprise Codeium.

**Varun Mohan** [39:20]
Yeah.

**Swyx** [39:21]
Anything moving there, or anything that people should know about?

**Anshul** [39:23]
Anshul, you wanna take that? I th- I think a lot are still, still moving there, right? I think it would be a little bit like, you know, very kind of egotistical of us to be like, "Oh, we have Windsurf now.

All of our enterprise customers are gonna switch to Windsurf, and this is the on..." Like, no, we still support the other IDE.

**Swyx** [39:34]
I was gonna say, you just talked about your, your Jet- your Java guys loving JetBrains.

**Anshul** [39:38]
Yeah.

**Swyx** [39:38]
They're never gonna leave JetBrains.

**Anshul** [39:39]
They're, they're not. Like, I mean, forget JetBrains. There's still tons and tons of enterprise people on Eclipse. Like, we're still the only code assistant that has an extension of Eclipse. That's still true years in, right? And but, like, that's 'cause that's, that's our enterprise customers.

And the way that we always think about it is, like, how do we still maximize the value of AI for every developer? I, I don't think that part of who we are has changed since the beginning, right? And there's a lot of, like, meeting the developers where they are.

So I think on the enterprise side, we're still pretty invested in, in doing that. We have, like, a, a team of engineers dedicated just to making enterprise successful and thinking about the enterprise problems. But really, if we think about it from the really macro perspective, it's like if we can solve all the enterprise problems for an enterprise, and we l- have the...

we have products that developers themselves just truly, truly love, then, then we're solving the problem from both sides. And I think it's one of those things where I think when, you know, we started working with the enterprise and we started building, like, dev tools, right?

We started as an infrastructure company, now we're now we're building dev tools for, for developers. You really quickly understand and, and realize just how much developers loving the tool make us successful in an enterprise. There is a lot of enterprise software that developers hate.

**Swyx** [40:41]
There's a little bit of... I wanna draw this flywheel and sketch it.

**Anshul** [40:44]
But, like, we're giving a tool for people where they're doing their most important work. They have to love it. And, and it's not like we're, we're, like, trying to convince... The executives at this company also ask their developers a lot, "Do you love this?"

Like, that is, like, almost always a key aspect of whether or not Codeium is accepted or not, like, into the, into the organization. Like, I don't think we go from zero to 10 million ARR in less than a year in an enterprise product if we don't have a product that developers love.

So I think that's why we're just... You know, the IDE is more of a developer love kind of play. It will eventually make it to the enterprise. We still solve the enterprise problems, and again, we could be completely wrong about this, but, but we hope we're solving the right problems.

**Swyx** [41:18]
It's interesting, 'cause I, I asked you this before we started rolling, but, like, it's the same team that's... the same eng team. Like, I, in any normal company or, like, in, you know, my normal mental model of company construction, if you were to have, like, effectively two products like this, like, you would have two different teams serving two different needs, but it's the same team.

**Varun Mohan** [41:35]
Yeah. I think one of the things that's maybe unique about our company is, like, this has not been one company the whole time, right? Like, we were first, like, this GPU virtualization company pivoted to this, and then after that, we're making some changes.

And, like, I think there's, like, a versatility of the company and, like, this, this ability to move where we think the instinct... Uh, we have this instinct where... And by the way, the instinct could be wrong, but if we smell something, we're gonna move fast.

And I think it's more a testament to, I think, the engineering team rather than any one of us.

### Blog Lessons

**Anshul** [42:03]
Anshul, you had December 19, 2022, you had one of our guest posts, What Building Copilot for X Really Takes.

**Varun Mohan** [42:11]
Oh, boy.

**Anshul** [42:11]
Estimate inference to figure out latency quality. Build first party instead of using third party as API.

**Varun Mohan** [42:17]
Okay.

**Anshul** [42:17]
Figure out real time, because ChatGPT and DALL-E, DALL-E and RP are too slow. Optimize prompt because context window's limited-

**Varun Mohan** [42:26]
Uh-huh

**Anshul** [42:26]
... which is maybe not that true anymore. And then merge model outputs with the UX to make the product more intuitive.

**Swyx** [42:33]
Is there anything you would add-

**Anshul** [42:34]
I'd give myself like a B minus on that

**Swyx** [42:36]
Yeah, no, it's pretty good

**Anshul** [42:37]
Like some, some parts of that are, are, are accurate. Uh, even like the context, like the one that you called out. Like, yeah, models have like larger context lengths now. That's absolutely true. Like, it's grown a lot. But look at like an enterprise code base.

**Swyx** [42:46]
Yeah, yeah, totally

**Anshul** [42:46]
They have like, you know, tens of millions of lines of code that's hundreds of billions of tokens. Like, uh, we don't-

**Swyx** [42:51]
Never going to change, yeah

**Anshul** [42:52]
... still being really good at, like, being able to piece together this like distributed knowledge source is important. So I think, like, there are figures there that I think are, are still pretty accurate. There's probably some that are, you know, less so.

**Varun Mohan** [43:01]
Like first party versus third party.

**Anshul** [43:01]
Like first party versus third party, I think we're, like we're wrong there.

**Varun Mohan** [43:04]
We just got it wrong.

**Anshul** [43:04]
I think, I think I would nuance that to be like there are certain things that it's really important to first party, like, you know, auto-complete. You have a really specific application that you can't just prompt engineer your way out of or just maybe even like fine-tune afterwards, like you just can't do that.

I think there's truth there, but like let's also be realistic. Like the stuff that is coming out from the third model providers, like Cascade and Windsurf would not have been possible if it wasn't for the rapid improvements with 4.0 and 3.5 Sonnet.

Like, that just wouldn't have been possible. So I'll give myself a B minus. I'll, I'll say I, I passed, but yeah, it's two hour- two years later-

**Swyx** [43:33]
I don't think... Just, just to be clear, we're not grading. It's more of a-

**Anshul** [43:37]
I'll grade myself. I always grade myself

**Swyx** [43:37]
... what, what would you, you know-

**Varun Mohan** [43:39]
But where are they now, you know, kind of thing

**Swyx** [43:40]
... what would you have added? What would you like?

**Anshul** [43:42]
Yeah.

**Swyx** [43:42]
You know.

**Anshul** [43:42]
I mean, like that, that first post, right? Like that was when we had literally had... I think that was like a few weeks after we had launched Codeium. I think that's like, you know, so Six and I-

**Varun Mohan** [43:49]
Oh, yeah

**Anshul** [43:50]
... we're, we're talking like, "Maybe we can write this," 'cause we were like one of the first products that people can actually use with AI. That's cool.

**Swyx** [43:54]
I specifically like the Copilot for X thing.

**Anshul** [43:56]
Yeah.

**Swyx** [43:56]
'Cause everyone... It is so hard.

**Anshul** [43:58]
Yeah, at that time-

**Swyx** [43:59]
Like everyone wanted a Copilot, right

**Anshul** [43:59]
... like everyone was just like, you know, ChatGPT, Co- that's all that, that's all there was. So, but I think like, you know, that, that we didn't have an enterprise product. We- I don't even think we were like necessarily thinking of an enterprise product at that point, right?

So like all of the learnings that like, you know, we, we've had from the enterprise perspective, which is why I loved coming back for like a third time now on, on, on the blog. Some of those I think we kind of like figured.

Some of those we just honestly walked backwards into, had to get lucky a lot of the ways. Like we had many... We just did a lot. Like there's so many of like opportunities and deals that we had that we like lost for a variety of reasons that we had to like learn from.

There's just so much more to add that there's no way I would've gotten that right in 2022.

**Varun Mohan** [44:36]
Can I mention one thing that I think is... Hopefully this is not very controversial, but is like true about our engineering team as a whole. I don't think, uh, most of us got much value from ChatGPT, largely because I think the problem was...

And this is maybe a little bit of a different thing. It's like a lot of the engineers at the company who have been writing software for like over eight years, and this is not to say they know everything that ChatGPT knows.

They don't. They'd already gotten good enough at searching for Stack Overflow, invested a lot in searching code base, right? They can very quickly grep through the code incredibly fast, like every tool, and they've spent like eight years mastering that skill.

And ChatGPT being this thing on the side that you need to provide a lot of context to, we were not able to actually get... Like my co-founder just basically never used ChatGPT at all. Literally never did. And because of that, probably at the time one of our incorrect, uh, sort of assumptions was probably that, hey, like a lot of these passive systems need to get good because they're always there, and these active systems are gonna be behind.

I think actually Cascade was the big thing as a company where everyone is now using it, literally everyone. Biggest skeptics, and we have a lot of people at the company that are skeptical of AI. I think this is actually important-

**Swyx** [45:39]
Then why do you hire them?

**Varun Mohan** [45:40]
No, I think, I think here's the important thing. Those people that were skeptical about AI previously worked on autonomous vehicles. These are not crypto people. These are people that care about technology and wanna work on the future. Their bar for good is just very high.

They will not form a cult of, "This is awesome. This is gonna change the world." They were not gonna be the kind of people on Twitter that are like, "This is-"

**Swyx** [45:59]
WAGMI.

**Varun Mohan** [45:59]
Yeah. "This, this changes everything," like, "Software as we know it is dead." No, they are people that are gonna be incredibly honest, and we know if we hit the bar that it is good for them, we found something special.

And I think at that time we probably had a lot of sentiment like that. That has changed a lot now, and I think it's actually important that you have believers that are incredibly future-looking and people that kind of rein it in.

'Cause otherwise you just have... Y- you know, this is like autonomous vehicles. You have a, you have a very discreet problem. People are just working in a vacuum, and there's no signal to kind of bring you down to reality, right?

You have no good way to kill ideas, and there are a lot of ideas we're gonna come up with that are just terrible ideas, but we need to come up with terrible ideas, otherwise, like how does anything good come out?

And I don't wanna call these skeptics. Skeptics suggest that they don't know. They're realists. They're the type of people that when they see wait list on a product online, they just will not believe it. They will not think about it at all.

**Swyx** [46:47]
Kudos for launching without a wait list.

**Varun Mohan** [46:49]
Yeah.

**Swyx** [46:49]
Yeah.

**Varun Mohan** [46:49]
Yeah. We- by the way, we will never launch with a wait list.

**Anshul** [46:51]
Absolutely.

**Varun Mohan** [46:52]
We will never launch with a wait list. That's a thing at the company. We'd much rather be a company that's considered the boring company than a company that, that launches once in a while and, like hopefully it's good.

**Anshul** [46:59]
Also-

**Swyx** [46:59]
My, my joke is, uh, generative AI has, has really gotten really good at generating wait lists.

**Anshul** [47:04]
Yeah.

**Swyx** [47:04]
And this one's really good.

**Anshul** [47:05]
Also, just to clarify, both of us used to work in autonomous vehicles, so it doesn't come across as like-

**Varun Mohan** [47:09]
Oh, yeah, yeah. I know that's right. That's right

**Anshul** [47:11]
... should we, should we have done autonomous vehicles?

**Varun Mohan** [47:12]
No, we-

**Anshul** [47:12]
We love that technology

**Varun Mohan** [47:13]
... we love, we love it. Like I love hard technology problems. That's what I live for.

**Swyx** [47:17]
Amazing. Um, just push back on the first party thing. I accept that the large model labs have just like done a lot of work for you than, that, that you didn't need to duplicate, but you now are sitting on so much proprietary data that it's, uh, maybe worth, uh, training on the trajectories that you're collecting.

So maybe it's, maybe it's the pendulum back to first party.

**Anshul** [47:38]
Yeah, I mean, I think like, I mean, we've been pretty clear from like a security posture perspective. Like I think there's like both like, you know, customer trust and like, you know-

**Varun Mohan** [47:46]
I mean, I kind of want... Like let me offer-

**Anshul** [47:47]
So I, I think that there is signals that we do get from our users-

**Varun Mohan** [47:50]
Yeah

**Anshul** [47:50]
... that we can utilize. Like there's a lot of preference information that we get, for example.

**Varun Mohan** [47:53]
Which is effectively what you're saying, of like our trajectories-

**Swyx** [47:56]
Go ahead

**Varun Mohan** [47:56]
... our trajectory's good. We-- Like I will say this, the super complete product that we have has gotten materially better because of us not only using synthetic data but also getting the preference data from our users of like, "Hey, given these set of trajectories, here's actually what a good outcome is."

And in fact, one of the really beautiful parts about our product that is very different than a ChatGPT is we can not only see if the acceptance happened, but if something more than the acceptance happened, and, and it happened even more than that, right?

Like let's say you accepted it, but then after accepting it you deleted three or four items in there. We can see that. So that actually lets us get to even better than grou- than, than acceptance as a metric, because we're in the ultimate work output of the developer.

**Anshul** [48:35]
It's the preference between the acceptance and what actually happened. As-- If you can actually get ground truth of what actually happened, this is the beauty of being in the IDE, then, like, yeah, you get a lot, a lot of information there.

So that's-

**Swyx** [48:44]
Did, did you have this with the extension or is this pure Windsurf?

**Anshul** [48:47]
The-- We had this with the extension.

**Swyx** [48:48]
Yeah. Okay. Great.

**Varun Mohan** [48:49]
Yes. Yes.

**Swyx** [48:49]
The Windsurf just gives you more of the IDE.

**Varun Mohan** [48:51]
Yes. So that means you can also, you can also start getting more information. Like for instance, the basic thing that Anshul said, we can see if like a file explorer was opened. That's actually just a piece of information we just cannot see previously.

**Swyx** [49:02]
Sure.

**Varun Mohan** [49:02]
Yeah.

**Anshul** [49:03]
A lot of intent in there. A lot of intent.

**Alessio** [49:05]
Second one.

**Anshul** [49:06]
Oh, boy.

**Alessio** [49:06]
How to make AI UX remote.

**Anshul** [49:08]
Oh, man. Is, isn't that funny that we now, now created like the full UX experience in an IDE? I think that one, that one is pretty accurate.

**Alessio** [49:14]
That one's an A?

**Anshul** [49:15]
I, I think that one I'd give myself... I think like we were doing that within the extension. I still think that's true within the extensions as well, right? Like, we, we got very, very creative with things. Like Varun mentioned the idea of just like, you know, essentially rendering images to display things.

Like, we get creative to figure out what the right UX is doing there. Like, we could create a really like dumb UX that's like a side panel, like whatever. Like, but no, actually going the extra mile does make that experience as good as it possibly can there.

But yeah, now like look at some of the UX that we're able to build in like, in Windsurf, and it's just like, it's fun. The first time I saw... 'Cause now we can do command in the terminal, like you can not have to search for a Bash command.

The first time I saw that, I was like, I just started smiling and like, it's like it's not, it's not like Cascade, it's not like agentic system writing the lab, but I'm like, th- that is just a very, very cool UX.

**Varun Mohan** [49:58]
We literally couldn't do that in VS Code.

**Swyx** [50:00]
Yeah. I, I understand that. Yeah. Uh, I've, I've implemented a 60-line Bash command called Please, and you can, you know, do that inside of it.

**Varun Mohan** [50:07]
Oh, wow.

**Swyx** [50:08]
Yeah.

**Varun Mohan** [50:08]
That's cool.

**Swyx** [50:09]
Yeah. So Please English and then a-

**Varun Mohan** [50:10]
You know, that's actually really cool because one of the things I think we believe in is actually, I like products like AutoComplete more than Command, purely because I don't even want to open anything up. Uh, so that thing where I just can type and not have to press some button shortcuts to go in a different place, I actually like that too.

**Swyx** [50:25]
Yeah.

**Anshul** [50:26]
Yeah.

**Swyx** [50:26]
Uh, and I actually adopted Warp, the terminal Warp initially for that 'cause they gave that away for free.

**Varun Mohan** [50:30]
Wow.

**Swyx** [50:31]
Uh, but now it's everywhere, so I c- I, I can turn off a Warp and not give Sequoia my, uh, Bash commands. Yeah. Um.

**Alessio** [50:38]
I'm with you. Um. No, I use, I use War- No, no, look, I use Wa- Okay, I don't know. Gonna go on a rant. Hopefully somebody at Warp-

**Anshul** [50:47]
Let's hear that. Let's hear that

**Alessio** [50:47]
... this is like Warp product feedback. But they basically had this thing where you can do kinda like pound and then write in natural language.

**Anshul** [50:54]
Yeah, I use that one too.

**Alessio** [50:54]
But then they have also the auto infer if what you're typing is natural language, and those are different. When you do the pound, it's only like it gives you a predetermined command. When you like talk to it, it generates a flow.

**Swyx** [51:06]
Okay.

**Alessio** [51:06]
It's a bit confusing of a UX. But going back to your post, you had the, the three Ps-

**Anshul** [51:12]
Yes

**Alessio** [51:12]
... of AI UX.

**Anshul** [51:12]
What were they again? I'm curious.

**Alessio** [51:13]
Present, practical, powerful.

**Swyx** [51:16]
Actually, that was really good. I liked it.

**Alessio** [51:18]
Yeah.

**Swyx** [51:18]
Yeah.

**Alessio** [51:18]
Uh, and I think, I, I think like in the beginning, being present was enough. Maybe, you know, even when you launch it's like, oh, you have AI, like that's cool, other people don't have it. Do you think we're still in the practical where like the experience is actually...

Like the model doesn't even need to be that powerful, like just having better experience is enough? Or like do you think like really the being able to do the whole... Because your point was like you're powerful when you generate a lot of value for the customer.

You're like practical when like you're basically like wrapping it in a nicer way. Yeah. Where, where are we in the market today?

**Anshul** [51:49]
I think, I think there's always gonna be room for like practical UX, like getting it... Like, I mean, the command terminal thing, that's like a very practical UX, right? Like I do think with things like Cascade and these agentic systems, like we are starting to get onto powerful 'cause like there's so many pieces, like from a UX perspective that make Cascade really good.

Like it's, there's really like micro things that are like just all over the place. But as, you know, we're streaming in, we're showing like the changes, we're like allowing you to jump and open diffs and see it, we can run background terminal commands.

You can see what like the ter- like is, what's running, background processes are running. Like there's all these like small UX things that together come to a really powerful and intuitive UX. I think we're starting to get there.

It's definitely just the start, uh, and that's why we're so excited about where all this is gonna go. We're, I think we're starting to see the glimpses of it. I'm, I'm excited. It's gonna be a whole new ballgame.

**Swyx** [52:36]
Yeah. Awesome. First of all, it's just been really nice to work with you. It's, uh, you know, I, I do work with a, a number of guest posters and, you know, not everyone makes it through to the end and nobody else has done, done it three times, so kudos.

Um, yes.

### Durable Build

**Anshul** [52:50]
We went for a hat trick.

**Swyx** [52:52]
Um, this one was more like the money one, which I, you know, I... It's funny 'cause I think developers are like quite uninterested in money.

**Anshul** [53:00]
Just-

**Swyx** [53:00]
Isn't it weird?

**Anshul** [53:01]
Y- yeah. I mean, I think like, I don't know if this is just the nature of our company. Like I think there's, some people say like there's all like the San Francisco AI companies and like everyone's like hyping each other.

They're like on the tech and everything, which is like great, the tech's really important. We're here in Mountain View, beautiful office. We just really care about like actually driving value, making money which is, which is kind of like a core, a core part of the, a core to the company.

But yeah.

**Varun Mohan** [53:21]
I think that, I think maybe the, the selfish way of, of saying that or like a little more of the selfless way is like, yeah, we can be kind of like this VC-funded company forever, but ultimately speaking, you know, if we actually want to transform the way software happens, we need this part of the business that's cash regenerative, that enables us to actually invest tremendously in the software, and that needs to be durable cash, be cash that like churns the next year, and we want to set ourse- set ourselves up to be a company that is durable and can actually solve these problems.

**Swyx** [53:50]
Yeah. Yeah. Excellent. So for, for people, obviously we're gonna link in the show notes, but for people who are listening to this for the first time, I had a lot of trouble naming this piece.

**Anshul** [53:59]
Oh, yeah.

**Swyx** [53:59]
Uh, so we, we originally called it, uh, you had like how to make money something.

**Anshul** [54:03]
I, I for- It was, it was-

**Varun Mohan** [54:05]
He had three dollar signs.

**Anshul** [54:06]
I, I, I apologize. I was super baity. I was... I think I was like writing part of that during like-

**Varun Mohan** [54:11]
He had like three dollar signs

**Anshul** [54:11]
... like on a plane flight, so I apologize for that.

**Swyx** [54:13]
Yeah, yeah. No, he had like three dollar signs in the title.

**Anshul** [54:14]
Oh, I absolutely had three dollar signs.

**Swyx** [54:15]
I was like, "I, I can't do that." So it's either building AI for the enterprise, and then I also said the worst, the most dangerous thing an AI c- startup can do is build for other AI startups, which I think you'll, both of you will co-sign.

And I think basically the, the main thesis, which I really liked, was like, go slow to go fast. Like here's the... Like if you actually build for like security, compliance, personalization, usage analytics, latency budgets, and scale from the start, then you're gonna pay that cost now, but eventually it's gonna, gonna pay off in the long run, and this is the actual insight, you cannot do this later.

Like, if you, if you build the easy thing first as an MVP, it's like, yeah, like, just ship it with, like, whatever's easy, easy to do, and then you tack on the enterprise-ready .io set of, like, 12 things that, that you have, you actually end up with a different product or you end up worse off-

**Anshul** [55:00]
Yeah

**Swyx** [55:00]
... than if you had started from the beginning. So that I had never heard before.

**Anshul** [55:04]
Yeah, I mean, we, we see that, like, repeatedly. I mean, just, like, right now, you know, we're-- have a lot of customers in, like, the defense space, for example. We're going through, you know, FedRAMP accreditation right now, and people that we're working with, they saw, like, all the fact, like, "Oh, yeah, we already have, we already have a containerized system.

We can already, like, deploy in these manners. We can already do, like, ex-- We've already gone through, like, security..." They're like, "Oh, you guys are gonna have a much easier time doing this," right? "than most companies that are just like, 'Okay, we have, like, a big SaaS blob, and now we need to, like, do all these things.'"

It might sound like a really deep thing. I think, like, it's just anyone who's, like, worked, like, you know, for, like, an extended period of time at, like, a company, like, on a certain project has probably seen this happen, right?

Like, the technology just keeps on improving, and then you realize that, like, you have to now, like, re-architect your whole system to get something improving. Like, just making that kind of change when you've invested so much effort, like, people have, like, put in hours, they're emotionally invested, whatever it might be, it's really hard to make that change.

So I'm sure we're gonna hit that also. Like, yes, I think we've done things a little bit earlier than most companies. I think we're gonna hit points where we're gonna see parts of our systems where we're like, "Oh, we really need to re-architect that."

Actually, we've definitely hit that already, right? And I think that's just, like, at, like, the project level, the, you know, the product level, or is that, like, your whole company, right? I think the thesis behind here is, like, to some degree, your company needs to have this DNA from the beginning, and I think then you'll be able to go through those bumps a lot more smoother, um, and, and be able to drive the value.

I know Varun probably has a point.

**Varun Mohan** [56:27]
Yeah. Can I say, can I say two points? So first point I'd like to say is, um, this is something that me and Douglas, my co-founder, talk about a lot. It's like, you know, there's this constant thing of, like, build versus buy.

I think the answer is, like, a lot of the time the answer should be buy, right? Like, we're not gonna go build our own sales tool. We should go buy Salesforce, right? That's kinda dumb. That's undifferentiated. And the reason why you go with buy instead of build is, hey, like, look, the ROI of what exists out there is good.

From, like, an opportunity cost standpoint, it's better to actually go out and buy it than build it and do a shittier job, right? There's a company that's actually going out and focused on that. But here's the hidden thing that I think is, like, really important when you go out and buy.

You're losing a core competency inside the company, and that's a core competency you can never get. It's-- Or it's very hard. Like, startups are so limited on time. Let me just say, like, let's say as a company we did not invest in, I don't know, model inference.

Yeah, we have, like, a custom inference runtime. We give that up right now, we will never get it back. It's gonna be very hard to get it back, right?

**Swyx** [57:22]
You can't just use vLLM and TensorRT.

**Varun Mohan** [57:25]
vLLM and you-

**Anshul** [57:25]
I mean, that, that's-- that would be our only option.

**Varun Mohan** [57:27]
Just, just or just let me put it... If we use vLLM, we would not be talking with you right now. Like, yeah, we would've, we would've... Yeah. But the point is, this is more a question of, like, like, you know, I try to think about it from first principles.

Like, Google's a great company, makes a lot of money. What happens if they actually made the search index of the product something that someone else built for them? It's like they could. Maybe someone else could have done a good job.

May-maybe that's, like, a bad example, but, like, yeah, particularly because Google is a search index, but, like, tough luck getting that core competency back. You've lost it, right? And I think for us it's more a question of, like, what core competencies do we need inside the business?

And yeah, like, sometimes it's painful. Like, sometimes actually, like, some of these core competencies are annoying. Sometimes we'll be behind, uh, behind what exists out there, right? And we need-- just need to be very honest. That's where the truth-seekingness of the company matters.

Like, are we really honest about this core competency? Can we actually keep up? If the answer is we truly can't keep up, then why are we keeping up with this trade? We should just buy, right? Like, let's not build.

If the answer is we can, and we think that this will differentiatedly make our company a better company in the long term, then the answer is we sh- we need to. We need to because, like, the race is not won in the next year.

The race is won over the next, like, five, ten years, right? Maybe even longer, right? So that's, like, that's maybe one thing. And then the second thing actually from, like, the enterprise standpoint, I think one of the unique parts of the company is, now is we have, like, both this individual and enterprise, and usually companies stick to one or the other, and I think that needs to be part of the DNA, I, I think kind of early on in the company as, as Anshul said.

I mean, there's stories of companies like Dropbox and stuff that tried. And Dropbox is an amazing company, fantastic company that-- One of the fastest-growing consumer companies of all time. Consumer more on the software company of all time. But yeah, like, when you have everyone sort of product oriented on the consumer side, the enterprise is just...

it's checking off a lot of boxes that ultimately do not help the consumer at all, doesn't help your growth metrics, and effectively, if the original group of people didn't care, it's incredibly hard to get them to care down the line, right?

**Swyx** [59:14]
Yeah.

**Varun Mohan** [59:14]
It's incredibly hard. Why do it? And you need to feel like, hey, this is, like, this is an important part for the company's viability. So I think there's a little bit of, like, the build versus buy part and then also, like, the cultural DNA of the company that I think are both really important, and, and yeah, it's something we think about all the time.

**Swyx** [59:30]
I have the privilege of being friends with you guys off, off the air. I don't feel like-- Like, I think I know your work histories. Like, you, you say cultural DNA, but, like, it's not like you've built, like, giant enterprise SaaS before, right?

**Varun Mohan** [59:42]
Yeah. I, I think- Yeah.

**Swyx** [59:44]
So, like, where are you getting this from?

**Varun Mohan** [59:46]
Yeah. Um, in fact, in fact, I think the only other sort of... I, I guess, like, you know, when, when I look at my previous internships, maybe Anshul can provide some context here. It's like I worked at, like, LinkedIn and then Quora and then Databricks and to be honest, like, I was not that interested in B2B ETL software that much.

That's not what drives me when I wake up at night. So I-- because of that, because of that, I decided to go work in an autonomous vehicle company, uh, immediately after. I think part of it comes down to, uh, maybe a little bit of the unique aspect of the company and the fact that we pivoted as a company is, like, we want to, we want to be a durable company, and then the question is how do you work backwards from that?

There's a lot of things about being very honest about what we're good at and what we're not good at. Like, I think surprisingly, enterprise sales is, like, not something that, like, I came out of the womb knowing how to do.

I didn't really know. And because of that, like, obviously, like, a lot of, uh, sales happen between sort of folks like Anshul and I helping partner with, with companies. But very soon we hired actually a VP of sales, and we've actually been deeply involved in the process of scaling out, like, a large go-to-market team.

And I, I think it's more a question of, like, what matters to the company and how do you actually go out and build it? And I think one of the people that I think about a lot actually is someone like Alex Wang.

He dropped out of college. He was a year younger than us at MIT, and he's figured out how to constantly change the direction of the company. Effectively, it starts out as like, you know, human task interface, then an AV labeling company, then a cataloging company, then now a generative AI labeling company.

And every time, the revenue of the company kind of goes up by a factor of 10 even though the business is doing something wildly different.

**Swyx** [1:01:11]
I mean, now he's all about military contracts.

**Varun Mohan** [1:01:12]
Yeah, now it's probably gonna be military, and then after that it might be taking over the world. Like, he's just gonna keep increasing the stakes, and like there's no playbook on how this really works. It's just a little bit of like, you know, solve a hard problem and work backwards, uh, work backwards from that, right?

**Anshul** [1:01:26]
And, and we'll get lucky along the way.

**Varun Mohan** [1:01:27]
Yeah.

**Anshul** [1:01:27]
Like I, I don't think like... You think everything from first principles to the best of our abilities, but there's just so many variable unknowns that, yeah, like we, we don't know everything that's happening in every company out there, and everyone knows how fast the AI space has been moving.

Like, we have to be pretty, pretty good at adapting.

**Swyx** [1:01:43]
I wanna double click on one thing just, just because y- you, you brought it up and it's like a rare thing to touch on. Uh, VP of sales. We don't get to actually... We talk to pretty early stage founders mostly.

### Sales Scale

**Swyx** [1:01:51]
They, they don't usually have a pretty built out sales function. Advice, what kind of sales works in this kind of field, um, and you know, what, what didn't work? You know, anything you can share with other founders.

**Varun Mohan** [1:02:02]
I think one of the hard parts about, about hiring people in sales, and I, I really, like Graham Anchal can also attest, like we have amazing VP of sales at the company. One of the things is like if you're purely a developer, salespeople, their job is to like talk like really well from improper.

I mean, very obvious if you hear like me talk, like I'm not a very polished person compared to most.

**Swyx** [1:02:21]
You're great by the way. I don't know.

**Varun Mohan** [1:02:22]
Okay. Other compared to, compared to most, uh, pure, pure salespeople. So actually just checking based on the way they speak is not that interesting. I think like, you know, what matters in a space like ours that is very quickly, uh, very, moving very quickly, I think is like intellectual curiosity is very important, intellectual horsepower, understanding how to build a factory.

I'm not trying to minimize it, but in some ways scales of... You need to build something incredibly scalable here, right? It's almost like every year you're kind of making this factory twice, thrice, maybe as big, right? Because in some ways you have people that are quota carrying, like you need some number of people, and you need to make the math work, and you actually, the process of building a factory is not something you can just take someone who is a great rep at another company and just make them build a factory.

They've-- This is actually a very different skill. How do you actually make sure you have hundreds of people that actually deeply understand the product? Actually, Anchal works very closely also with, with sales to make sure that they're enabled properly, make sure that they understand the technology.

Our technology is also changing very quickly. Let's maybe take an example on how our company's very different than a company like MongoDB. When you sell a product like MongoDB, no one at the company is interested in how the data is being stored.

It's not that interesting, right? I love databases. I would be interested, but most people are like, "Solve the application problem I have at hand." People are curious about how our technology works. People are curious about RAG, right? People that are buying our technology.

And imagine we had a sales team that is scaling where no one understands any of this stuff. We're not gonna be great partners to our customers. So how do you create almost this, this growing factory that is able to actually distribute the software in a way that is true to our partners and also at the same time like taking on all the new parts of our product, right?

Like they're actually able to, uh, expound on new parts of our product. So sorry, that was more a question, more a statement of like building a scalable sales team. But in terms of like who you hire is you just need to have a sense.

Like in some ways, uh, this is maybe a, uh, an example of talk to enough people, find out what good looks like potentially in your category, and find someone who's good and humble and willing to work with you.

**Swyx** [1:04:13]
Yeah, that's just gen- generic hiring.

**Varun Mohan** [1:04:15]
Yeah, it's just generic hiring. Yeah.

**Swyx** [1:04:16]
I think here sales there's, there's sales for AI, uh, or sales for AI infrastructure-

**Varun Mohan** [1:04:22]
Okay

**Swyx** [1:04:22]
... and then there's also the sales feeding into products in the way that we're talking about here, right? Where like they basically tell you what you, what they need. I imagine a lot of that happened.

**Anshul** [1:04:31]
I think a lot of that happened, I mean, still ha- and Varun mentioned like Varun, myself, a number of other people who are, you know, developers by trade, engineers, like we're pretty involved in the sales process 'cause like there is a lot to learn, right?

Like we, before we went out and hired a sales leader, like yeah, if all we want is like neither of us had ever done a sale for Codeium in our lives, and we went and tried to find a sales leader, we, we probably would have not hired the right person.

But like-

**Varun Mohan** [1:04:54]
Yeah, we had sold a product to like 30 or 40 customers at that time.

**Anshul** [1:04:56]
Like we, we had done like hundreds and hundreds of deals cycles ourselves personally, right? Without... I mean, we read a lot of books, and we just did a lot of stuff, and we learned like what messaging worked, like what did we need to do, and then I think we found like the right person, right?

Uh, a second Varun, like Graham's amazing and they're who we brought on as our VP of sales. That just has to be part of, of the nature, and it doesn't stop now. Like just because we have a VP of sales and, and, and people dedicated to sales, it doesn't stop that we can't be involved or like engineering can't be involved, right?

Like we have lots of people, like we hire plenty of deployed engineers, right? These are people like, you know, I think like Palantir kind of made this really famous-

**Swyx** [1:05:31]
Forward deployed engineers

**Anshul** [1:05:32]
... but forward, like deployed engineers like work very, very closely with the sales team on very technical aspects because they can also understand like what are people trying to do with AI.

**Swyx** [1:05:39]
As in they work at Codeium as deployed engineers?

**Anshul** [1:05:41]
Yeah.

**Swyx** [1:05:42]
Okay.

**Anshul** [1:05:42]
And they, they partner with the, with our, with our account executives to like make our customers successful and like learn what is it that people are actually getting value with AI, right? And like that's information that we keep on collating, and it's like we will both jump into any deal cycle just to learn more because that's how we're gonna just keep on building like the best product.

It, it comes back to the same, like just care. I don't know. And hopefully we build the right thing.

**Swyx** [1:06:05]
Cool, guys. Thank you for the time.

**Varun Mohan** [1:06:06]
Yeah, thank you for your time.

**Swyx** [1:06:07]
And it's great to have you back on the pod.

**Varun Mohan** [1:06:09]
Yeah, thanks a lot for having us. Hopefully in a year we can do another one.

**Swyx** [1:06:11]
Yeah. You'll be 10 billion by then.

**Varun Mohan** [1:06:13]
Yeah, exactly.

**Swyx** [1:06:14]
At this rate, by next year.

**Anshul** [1:06:15]
Ooh. We try not thinking about that.

**Varun Mohan** [1:06:17]
Try to not be a zero billion company.

**Anshul** [1:06:19]
Yeah. That's... Well, well, I'm glad there's that, yeah.

**Swyx** [1:06:21]
All right, cool. That's it.

**Varun Mohan** [1:06:22]
All right. Awesome.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
