Origins0:00
Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, partner and CTO at Decibel, and I'm joined by my co-host, Swyx, founder of Small AI.
Hey, and this is a little bit of a Latent Space Discord, uh, reunion because we have Vasek in the house from E2B. Welcome.
Hey. Uh- Good to see you guys. Thank you for having me.
Help me with your last name because I realized I've never had to pronounce it.
Yeah.
Was it Mlejinsky?
Uh, no, no. I-- Well, I guess close enough. Mlejinsky.
Mlejinsky.
Yeah.
Okay.
And my, like, official legal first name is Vaclav, but everyone pronounces it as Vaclav, which I hate, so it's just Vasek.
Okay. Awesome. We're both, uh, invested in E2B in different ways. But you and I go back, uh, the furthest, I just realized, three years ago when you were working on DevBook, and you were interested in, like, sort of that developer experience angle, and somehow you pivoted to E2B.
Yeah.
Maybe you want to tell that story.
Yeah. So Tomas, my co-founder, who's our CTO, we've been interested in dev tools for quite a long time, like six, eight years. And before DevBook, there was like bunch of iterations. After-- There were, like, different iterations and pivots of DevBook.
We just, like, stopped renaming things at some point and just, like, went with DevBook. The one, uh, you are talking about was interactive documentation for developers. So basically, the idea was we wanted to-- Instead, like you as a developer when you come to a tools, uh, docs, uh, website, instead of reading about everything and then googling it and trying it in your coding editor, the idea was to give you interactive experience in the browser.
So you would, uh, have like pre-made interactive guides, uh, playgrounds. You could try things right away. And the, the company, the, the owner of the docs would prepare the experience for you because you would also be trying everything in the browser.
They would now see if you get stuck anywhere, what you are doing, you know. So you get a very valuable onboarding, like, analytics. We actually built an interactive playground for Prisma. I think it's still up. They are still using it the last time I checked a couple of months ago.
Uh, so it was like interactive guides and playground that you could try out Prisma without, you know, having to manually set up all the databases and everything. So you could just try Prisma right away. And that was like the very, very first version of actually our infrastructure that we are offering now.
Basically sandboxes.
It, it was sandboxes.
Yeah.
It literally was sandboxes, the same technology, but just completely unscalable.
Uh, so then did that somehow in 2024-ish turn to E2B?
2023.
In, in 2023.
I think March 23.
Okay. Yeah, yeah.
Um, and, uh, we were pretty burned out. Tomas and I, we were working from Prague, from the Czech Republic, from my apartment. Uh, like, nothing was really moving, uh, no growth, and GPT-3.5 came out, like really first model, kinda good-ish with code gen.
So we took like, let's, let's take ten days break, two weeks break from, from DevBook. And because, like, everyone was trying things with AI, it was very clear, like, this is something, uh, where the future might go. So we wanted to just like, from out of curiosity, try things out.
We wanted to build like a AI, like Devin kinda like thing. The first idea we had, had was, like, let's automate our work. Uh, because with every project we were starting, there was a set of tools you always want to, want to integrate in your backend, like Stripe, like for a size business, like Stripe Analytics, Slack notifications, emails, uh, sending out emails.
And so we gave the agent a tools to, uh, run code, and we needed some kind of sandbox. We were like, "Yeah, that's good coincidence."
Right.
"We have sandbox from, from DevBook." And, uh, we posted about it on Twitter. Basically, the agent actually pulled GitHub repository, wrote code, started the server, tested everything, and at that point, I think we deployed it through Railway. Railway had like the best DX, and it was easiest to just plug it, plug it in, into the agent.
I tweeted about it, and I think Greg Brockman, like, retweeted it. Like, hey, like, I don't remember exactly what he said. I, I need to find a tweet. But it was around the time where OpenAI was retweeting, or a p- OpenAI co-founders were, were retweeting all the things that people are doing.
Like, to show, like, what you can do with, um, GPT-3.5. And people-- Like, it had like half a million views after a few days. And so we were, with Tomas, we were like, "Uh, we gotta do something." Like, people are interested in, like...
So we, we, we just open sourced it, uh, the repository, and it was named-- The organization was, like, The AI Company. Like, we had no name.
Mm-hmm.
Uh, I kind of wish, like, we could have, like, a stick-- like, legally stick with it. Um, we just came up with a name. We was like, E2B, because you take English, you convert it to bits. That's how it started.
We started building community around it, and, like, two days later, we were like, "Let's focus on the sandbox part, not on the agent part." Um, and we, we actually had bunch of hypothesis, uh, behind it, like why it might be more interesting than building-
Mm-hmm
...the agent.
Early Agents5:19
Yeah.
And then you had the small developer, run small developer on E2B, which-
That, that was a little bit later.
Um, yeah. What was the-- What, what-- So-
This is in 2023 as well
...when you first launched E2B, it was kind of like a AI agents cloud. What did you call it?
'23, the hypothesis was, like, agents, code gen agents especially, will need some kind of environment to, uh, run the code. Um, and the same way a developer needs a laptop or something, you know. So, but we, we struggled a lot with how the product should actually look like and what's like a go-to-market, uh, the first version.
And so the high-level idea was, like, we will host your agent, everything from actually deploying it to then monitoring it, and, and it will also have this environment running the code. One of the first, like, test project was taking Shawn's, um, small agent project and deploying it inside our sandbox and just giving the agent tools to pull the GitHub repositories, work on that, do a PR, and then post it, uh, on GitHub, uh, post the g- the, the PR on GitHub, which was very, very popular.
Like, like, there was clear-
Yeah
... it was clear that, like, um, there was something very interesting-
It's amazing how, how much people went with that. It was literally not meant to do that. It was meant to do Chrome extensions.
Yeah, and I think, like- ... you really, y- you had good insight that you let the agent plan the work in, in markdown file, basically-
Yeah
... and, like, write down the spec, which I think e- even now when you look at deep research agents, like, that's sort of oftentimes what they are doing, like, they, they plan everything.
Yeah. These days I would say they should do more structured output than markdown, but, uh, you know, that's, that's an implementation detail.
Yeah. Yeah. So that was taking off, but it attracted a little bit different audience than we wanted. Um, so, uh, because it attr- attracted a lot people who nowadays would be using tools like Lovable, for example. You know, they just wanted to build projects for them, and of course, like, it wasn't really working at the time.
Like, like, it, it, it worked for, like, one simple website, but the moment you wanted something more complex, like, it was really hard.
I've been also reflecting on why I stopped working on it, and there was actually a, a... So when I built it, it was with, uh, Claude 3.
Yeah.
Uh, the, the new Claude 3 launch. It was just trying to utilize the 100K context of Claude 3, and I used the same project to do this to try to repeat the demo that I made for myself a month afterwards, and it wasn't anywhere as smart.
So Claude 3 got dumber. But it, it looks like I, like, made up the demo or something. But no, like, literally I just reran the same code, and it just, it just was not as capable.
Yeah.
And I, I mean, I think to some extent this is, like, RNGesus, you know, hurting me or, like, it's the time, it's like a different month, so, like, the, the model is, like, different or something. But I think also, like, basically people had this vision of what they wanted, and then they tried to do it in reality, and they couldn't do it 'cause the model isn't all ready.
So a lot of feedback we got, um- At some point people thought, like, small agents is, is from us because, like, we had a website, like, deploy small agent. But we were, "No, like, that's, that's, like, Shawn's work.
You should, you should go to his repository, uh, submit an issue."
No, they could, they could clone it. It was like, you know-
Yeah
... few hundred lines of code
Fork it, clone it-
Yeah, whatever
... um, edit it however you want. For, for, for us, it was more like a test if the environment, the sandbox is useful for the agent, and if it can sort of scale and, and it can work, which worked well.
But then it took us another, I would say, six months to actually find the right, you know, go-to-market strategy, which was code interpreting for us. So typically, AI data analysis, uh, data visualization inside, like, a headless Jupyter type of, uh, notebook en- environment.
Specifically Jupyter?
It doesn't need to be Jupyter. The important part is that, um, you don't need to explain the model, and the model doesn't need to care about how to keep the state of the program running. So it was, it was, especially with our earlier models, I think the models are now smarter, but they kept producing, like, co- code snippets and thought, like, they can reference to previous, uh, you know, variables and, and, and functions definitions.
Uh, which if you had only normal code execution, like you run the code snippet and then you finish, it wouldn't work. Like, so you need kind... some kind of REPL environment. And, um, what was, I think, especially early on, like, Python was the language that the models were probably the best or one of the first languages where the models were working very well.
Looking back at it, there was, like, a really strong pull from people just wanting to visualize data and talk with their data. Python was really good at it, and it mixes well with Jupyter type of environment because you get charts out of the box.
Uh, you can also support interactive charts. Jupyter itself isn't the, the right environment to actually do it because, like, uh, because all the technical problems that you will run into once you are start doing, once you start doing it on scale.
So it, it, it gets slower and slower, and actually, like, what we are coming into is, like, we are building our own thing, uh, internally just, like, to support LLM specifically. It's got, like, your own runtime.
Use Cases10:35
At what point did you start going from just code interpreter run code to, like, expand? Because now you have, you know, people doing RFT, you have Computer Use and all of those things. Like, when were the models ready for people to start using it?
Just, you know, you demoed at one of our first events as well, and I think there's always this lag between the infrastructure that you build and-
Yeah
... like, the capabilities of the model. When did you go from just code interpreter to start saying, "Okay, now it's time to do Computer Use. Now it's time to do RFT"?
That was probably end of '24, uh, start of '25. So when you look even at our data, like, you s- '24 we are growing, we are growing good, but '25 is, like, up to the right. And so it feels like at '24 people are, like, figuring out these agents, um, and, and building them and trying them, and '25 is everybody moving them into production and finding more and more use cases.
Around end of '24, start of '25, we started seeing things like using sandbox for, uh, reinforcement learning type of use cases or using sandbox as a, for Computer Use, uh, which was very interesting when Anthropic launched their Computer Use.
We had, like, a desktop version of a sandbox that was sitting, like, in our GitHub repository for six months. We were like, "Okay, this is probably interesting, but no model can actually use it." So, uh, when Anthropic announced it, it was like, "Oh, like we have something here, uh, we can show you."
Using, uh, with, with, like, Lovables and, and, and Blitz's type of products also, like, started using the sandbox for more than just, like, run code snippet, like data analysis. And then deep research agents that, that, that has been something really big in the last few months.
How does deep research agents... Are you referring to Manus or?
Yeah. Manus, for example, um, one of the company-
Deep research, I typically think of as like a, a search, like a web search-
Yeah
... heavy task. Doesn't really use a code interpreter in any way. I have some idea that Manus uses E2B in an interesting code re- code interpretee wave to do deep research. What's the difference?
Yeah. I, I think it's a good idea to stop thinking about a sandbox just for code interpreting and, and more about like a runtime, code runtime for the LLM-
Okay
... or the agent. The use case for the sandbox, it's, uh, very horizontal in a sense that it can cover everything from the agent needs to create a file, make a to-do list. It uses a browser not inside a sandbox, that's like separate.
You have browser use, uh, run, uh, use- being used for research on the website, but then you download the data somewhere. You need to transform the data. You want to d- do data analysis. You want to actually wr- uh, write a small app.
You want to create a, a Excel sheet. So it's, uh, the same way you are using, you know, your laptop as a human that is the same useful for the agent. So you can think about it as more like a DevBox and at the same time, the, the agent that's using it is also like a very, very, uh, good developer, good accountant, good, uh, like, slides creator, researcher, and so you are just basically giving it tools to let it do the job even better and faster.
Yeah.
Um, yeah, and when you say up and to the right, I just wanna share some numbers from the investor updates. So March '24-
This is clearly shared
... which we were talking about.
I don't know.
Yeah, yeah, yeah. You ta- you shared these.
This is, this is-
You shared these
... weeks ago.
No, you shared these, you shared these publicly.
Yeah, yeah.
Uh, but yeah, March '24, you were doing 40,000 sandboxes. You wanna say how many you've done last month, March 2025?
Uh, I- 15 million.
Yeah.
Around, around that.
So-
Yeah
... in one year you've gone from 40,000 to 15 million, and yeah, I think you can kinda see the slope, especially from, like, the Sonnet-3.7 release, and-
Yeah
... uh, I, I think this is, like, an interesting model versus infrastructure. I think there's this usual, like, VCEO we should invest in, like, the tools, uh, you know, and the picks and shovels instead of the application layer.
But I think it's, uh, uh, this is, like, the first time where, like, the infrastructure is lagging the applications, you know?
Yeah. That's, that's a good point. Like, I- I think '24 was all about, like, the agent couldn't use the whole sandbox, and now, uh, yeah, sometimes we are actually, like, catching up on with some features for the, for the LLMs, uh, that they need more than what we have, uh, at the moment.
LLMOS15:03
Yeah, I guess, like, what we're doing here, you know, y- you are another one of the LLMOS companies that we are talking to. We also did one with BrowserBase and, um, I'm not sure who else would qualify under that, that term but-
Exa.
Exa, yeah. Yeah. Uh, so, like, basically, like, you don't specifically yourself use LLMs internally in your products, but you enable others to work to augment their LLMs with your infrastructure.
Yeah.
Um, I don't... A- any reflections on just the general LLMOS landscape. You have other competitors, like, uh, you know, how, how is this evolving? How do you, how do you position in it?
A lot of people are saying, like, if you are a GPT wrapper in '23- ... you were really bad positioned, uh, because, like, all the value will be captured by the AI labs. It's, uh, good to be GPT wrapper because you get all the advantages from a new model.
You just switch it. I mean, the just is, like, not so simple. Probably have evals, um, need to change your prompt a little bit. But I would say it's increasingly easier to switch models. So, uh, we need to think about it the same way.
Like, u- our users are switching models a lot. We need to be agnostic, uh, to the, to the LLMs. Oftentimes, like, people want to deploy us in their, uh, cloud or on prem- on premise, so that's also something very important.
I think a good analogy here is sort of like technologically, it's kinda you want to be the Kubernetes of the world for the agent, but with much better DX, uh, and, and- ... easier, easier to use.
One thing I'm thinking about also is like what is valuable real estate to occupy in the LLMOS? Um, and I have this spectrum, uh, I, I think on the, in the BrowserBase episode I was talking about, like, either you can focus on browser emulation, or you can focus on the VM, or you can do, like, a custom mo- uh, Python sandbox, like, uh, Modal does.
Yeah.
You know, like, uh, would you say that you're the most general of all of them?
Yes. It's very general, but that's not really historically how you want to market it because people don't know what to do with it.
Okay.
Um, so that, like, we, we started as when we we- had our website, uh, in '24, it said something like computer, like cloud computer for AI. People didn't just understand, like, what to do with it. So we had to, like, literally show them, like, first you change it to code interpreting because that's something people knew from OpenAI, and then you just show them very, very specific use case, and you use that to get a early traction and early set of users.
And we spent a lot of time, and Tereza from our team spent a lot of time, especially her, uh, on educating the market-
Yes
... you know, and the developers. So, uh, and I think this is the space, the AI space is you sometimes have to, like, show developers what they might need, and you kinda have to, like, trust your gut, like, this might go in this dir- direction, may- like, this could make sense, and show them, like, what they can build with it because it's very hard to imagine what you can build with things that you don't even have, right?
So, like, why would you need forking sandboxes or checkpoint in sandboxes? Like, how is that useful? Well, it turns out if you are building some Monte Carlo type of a thing, uh, search, uh, for our agent, like, it, it's very useful.
Uh, but, like, to actually go for a developer who might not be, like, in the AI deep research type of thing. Like, it's not obvious. So, uh, you want to be agnostic, you want to be general, but you want to also show people, like, very clear use cases how they can use you.
And over the time, they then get educated-
Yeah
... and realize, "Okay, like, there's more use cases. I understand the platform." Uh, they start coming up with their own ideas. But onboard people with general use cases is very hard for... That's- at least that's what we learned hard way.
Yeah, and you also don't really tailor to, like, the more DevOps infrastructure-
Yeah
... person.
Yeah.
Which I think a lot of the other sandboxes are like, "Oh, we have a gVisor runtime, and we have all these, like, different terms" that you don't really know if you're, like, the AI engineer type.
Yeah, that's exactly true that, um, our user, uh, isn't, like, a infra engineer. Um, even ML engineer usually isn't our type of a user. It's like AI engineer, you know? Like- ... the, from definition you have, uh, uh, still remain valuable-ish web developers in JavaScript world, TypeScript world, even though we have, like, ton of usage from Python.
Um, but there are so many web developers, and it's easier and easier to use LLMs. I really think that you need to cater to these type of developers and make it, like, things simple for them, and if they want to dive in, they can.
Chances are, like, they d- don't even want to. They want to focus on building the app.
Yeah, yeah, product developers-
Yeah
... not infrastructure developers. It's very interesting. Yeah, I mean, like, uh, this, this whole GPT wrapper versus ModelLab thing, it's not like, you know, it's better to work on wrapper o- over, over ModelLabs 'cause, you know, the ModelLab's people are making a lot of money.
It's just that there-
Yeah
... there is room for wrappers, and the wrappers also do make money, and I think people were not seeing that in 2023.
Yeah.
And this is what we saw.
It's not, you know, binary.
It's not either/or.
Right.
Yeah. Yeah, I think, uh, and then, yeah, uh, just a quick check. Uh, do you know, like, the rough percentage between JavaScript and Python?
A slight less JavaScript. Um-
More Python still
... so I, like, from number of downloads of our SDK per month, it's, like, 250,000 JavaScript, close to ar- around half a million Python.
Yeah.
Something around that.
So two to one.
Yeah.
Yeah. Interesting. I mean, for, uh, if, if the use case is really for code interpreting, generating charts, and all these, then Python wins.
Yeah.
Like
Yeah, yeah, yeah. The Py- Python, exactly, Python wins for that, but once you go into more, like, for example, building apps, AI-generated apps, then it- it's JavaScript winning, right? Because you probably have- you have frameworks, all the Sveltes, Next.js's, Vue.js's frameworks, which is, like, um, JavaScript thing.
It's interesting because I would think that it doesn't really matter what code you want the LLM to produce. Like, it wouldn't dictate what kind of user is using us, right? Like, if you think about it, when a lot developers are building, like, AI data analysis, it's typically Python developer.
When developer's building, like, a type of, um, V0, like a use case, level of, like, a use case, it's a web developer. It's a product developer. So I, I don't think this is, like, super obvious, um, that... I wouldn't think that that would be the case.
There's all sorts. I mean, Volt's, um, argument is that you sh- you should want to use something like a web container where it's, like, run on your own browser and it's all, it's all, you know, free and very fast.
I guess, one, one more critical question. Like, let's say I do know my infra. Well, what is the point of an, uh, cloud for AI, like, computer for AI, right? Like, so basically, why can't I use existing tools like Railway?
Which I'm sure Railway wants to go after your customers, right? So, like, why does there need to be an AI-focused, you know, virtual machine or sandbox or execution environment, whatever you call it?
VM Specs21:50
What we offer is sort of orthogonal to cloud. So beforehand, you don't know what kind of code you will be running in our cloud, so you can't really optimize it, right? So there's no, like, build deploy step, uh, per se.
Everything happens ad hoc during runtime. So you need to solve problems like you want to install dependencies very fast. You want to be able to pull GitHub repositories ver- very fast. So then, uh, how do you... The work- workloads we are running can go from five seconds to five hours, so that also, like, changes even the pricing model a lot.
Uh, how do you make sure that everything makes sense for you and for the user, like, uh, uh, from the pricing point of view, from the, like, unit economics, and also from, like, just the infrastructure, uh, point of view where the, uh, you know, the sandboxes are getting placed inside your cluster.
The security model is, uh, is also different usually because also comes down, like, to you don't know beforehand what code you will run, so by default it's untrusted code. And you need to have complete isolation between these sandboxes to make sure first that, like, if something happens in one sandbox, it doesn't affect other sandboxes.
But also, you want to know about security inside the sandbox. So as the LLMs are getting better, you want to know what's happening inside the sandbox. There's, like, a fun story from Hugging Face, uh, when they are using, when they were using us, and, uh, are using us for their, uh, Open R1 model.
And one of the developers, uh, he shared it online, so I think it's okay to, to say that. So he lost access to their cluster because the LLM decided to change permissions. If that happens with us, we just, like, kill the sandbox and get, get a new one.
Uh, he, uh, it takes like, I don't know, 150 milliseconds. For them, like, they had to take down the whole cluster, set up everything. Didn't even know that this happened. So it's like, like the model, the compute model, uh, like TLDR is the compute model, security model is different from, like, current cloud providers, and you need to think from the first days about it differently.
And then people use, depending on the... Uh, because the difference that maybe people don't think about is, like, you can generate the code with the Python SDK, but it can still run Lua code or, like, R code.
Yeah, yeah.
Like, the, it's not always matched to the runtime-
Yeah, yeah, yeah
... of the SDK.
Yeah, you can... It's, it's a general machine, so whatever you can run on a Linux You can run inside a sandbox. And, uh, we had users like running, of course, like Python code, but C++. We had, um, Fortran, uh, someone running Fortran- ...
which is like-
Why?
They were doing some, some, like using very old API for like banking or something like that. But you can also start a server, uh, inside a sandbox, and then you want it to be accessible from the internet. It's a very general machine, and the challenge is like how do you make everything fast, for example?
How do you make everything secure? At the same time, you need to make it accessible enough and controllable by the LLM and observable for the, a human. So, you're kind of building for two personas: for a human developer, the AI engineer, and for the LLM, who's like using the sandbox.
Yeah, the composability thing is something I had not thought about before, but later, like you mentioned, it's like, you know, you might need to run Fortran to access one API, and then in the next step you'll need to take that data and run it in a Python, uh, script.
Yeah.
And then you're gonna expose that through JavaScript to something else, and like you guys can switch the runtime-
Yeah
... halfway.
Yeah. And, and in ideal world, you don't want to kill the sandbox, start a new one, uh, or maybe like create a new sandbox template. You just want to keep using the computer. And because we are in cloud, we can do it in a way that you can get more RAM, you can get more CPU.
You can get less CPU, so you are really paying only for what you need. And as the LLM is doing more and more, uh, you can have very elastic sandbox and a- keep adding features. And the goal where we, where we think this is getting, going is the LLM, like decides what it want to do and how it wants to have the sandbox configured.
So, it basically starts controlling the infrastructure itself and creating sandboxes themselves.
While we're talking about the technical details, I just wanted to let you tell people about any other technical details. Like what is the box where we get? What kind of Linux? Uh, what size of, what- whatever. What details matter here?
So, you get Ubuntu box. You can customize it, and anything Debian based is gonna work. You can even add, you know, graphical interface. So i- if you want to, you can run like legit, not headless, but, uh, like legit Ubuntu computer on it.
So like, but you can do the sort of takeo- uh, you can, like an o- operator type experience. Like you take over-
Yeah, yeah, yeah
... control of the-
Yeah, we have, we have SDK called Desktop SDK-
Yeah
... that does this for you out of the box.
Yep.
And it supports, uh, VNC. Uh, so you can also have like a human-in-the-loop type of a thing and control everything, and s- stream and see what's happening. So by default you get two CPUs and half a gig of RAM on the free tier, and you can customize it.
On our landing page we say up to eight gigs, but if you tell us what's your use case, you can go to 64 gigs of RAM. We have users, uh, using such a beefed sandboxes. You can go to 16 CPUs, I think, uh, if I'm not mistaken.
And-
Storage is free.
Storage is free. Um, I think there's like- ... a lot of things that are free that we probably need to think about a little bit more. I, uh, you know, with the dev tools, especially infra, I always see this pattern, and I had this like naive idea, idea as well, that the founder says like, "Oh, this AWS pricing, it's terrible.
Like you... Or GCP, like you have no idea what you are paying for. There's so many like small add-ons you need to pay for. We will start a new infrastructure company. You just- we're, you're just going to pay $200 per month, and, uh, then you just pay for like pure compute."
That works until you start scaling and you figure out, "Okay, like actually, like people are doing like weird stuff that I didn't expect, like having lots of traffic, producing..." We have a customer that produced petabyte of data. Uh, I mean, it's not free To host petabyte of data, and that's growing.
So then you start introducing, "Okay, probably they should pay some amount for ingress, egress, for storage." And the pricing gets increasingly complex. So, I just think it's a very interesting like phenomenon that, uh, you start with this brave idea that everything will be super simple, but-
And, and we, we only do code interpreting, so we only need to-
Yeah
... price for compute, right?
Yeah, yeah, ex- Exactly. And, and then you like wake up in the real world and it, it's messy.
Uh, yeah, so I, I know about this from like my career in cloud. And so like, the, the, the common refrain is that, uh, I, I call this the first principle of technology, everything can be broken down into some combo of compute, storage, and networking.
And if you fail to price one of them, they will... You will get abused because-
Yeah
... the, the you'll find it
Because you are essentially offering free storage or compute or, or something, right?
Yeah. Um, for those interested in this idea, there's a fourth one, which is basically the control plane or like the auth layer, the IM policies and all that. Uh, and, and that's like, maybe you can call that security as well as like the fourth layer that people kind of pay for as, as its own independent thing.
Uh, for those also interested, uh, HashiCorp has more breakdowns, uh, from this, like David McJannet, um, which I think is very interesting. If you're, if you're just in the business of running a cloud infrastructure company, you should know these things.
Many hundreds of businesses have run into the exact same problems. You should just not repeat them and just learn whatever the best practice is.
Billing29:43
Yeah. I, I think like the billing model has been figured out many times.
Yeah.
So like, I don't think it's, uh, like the challenge is not figuring it out. The challenge is I think is introducing it, like sometimes quick enough. Ma- you actually need to do changes on your infrastructure to make sure you know about all this data that's happening and moving one way or another.
But yeah, I, I completely agree with you. Like this problem, we are not the first one that- ... are having this.
Yeah. Uh-
Do, do you use one of the usage-based billing providers?
Orb.
We are talking with Orb, uh, right now. So we are, um-
Orbin, Meter
... uh, OpenMeter is another one I think.
Yeah.
Or Meter just-
Yeah, Metronome.
Metro-
Metronome
... Metro, yeah, Metro-
Yeah
... Meter is the other one.
Yeah, yeah.
We have been using Stripe, uh, Stripe's usage, which has been a little bit, uh, sometimes rougher around the edges.
Yeah, this is the thing, like, w- we shouldn't spend so, like, much engineering on it. Like, y- I want to outsource it because that's not our product. Uh, someone else should be, like, uh, focusing on this full time, and it's actually pretty non-trivial, like, to make sure you have everything right, and you really don't want to make mistakes here.
Is there anything that you're really looking for that would say, like, "Okay, th- that's really what I- what we want" that maybe Orb or Metronome haven't really adjusted for AI yet? I don't think this is AI-specific problem. Okay.
Uh, this is, this is infrastructure as, as, as you know it. For us, some things that didn't work, uh, when we look at some of the providers where, like, the cut they took from the revenue, for example, they take from you.
Uh, so some of the pricings, uh, I don't know- Mm ... what's the latest, but, like, I, I don't remember which one was it, honestly, but I knew that pricing was basically they take a small cut from the revenue each month, which is basically like what Stripe does when you are processing payments.
It's just, like, the value was really high. And then it's, it's a lot about how hard it's go- it would be to, uh, integrate it, like, how much time we are spending on it. Uh, because we know we don't want to build this in-house.
It's more like, is the switch worth it, basically. I'm curious your updated takes. I know you had the Y is in usage-based- Yeah ... billing a bigger category. I'm curious now with AI with, like, token-based pricing, if you have updated thoughts on- Uh, so for, for people who don't have context, um, at Netlify, we went from a relatively flat tier-based pricing to usage-based billing because that was, uh...
That's basically how all infrastructure companies should, should eventually go. Because you have some whales who use a lot of infrastructure and some who don't use that much, and you shouldn't, you shouldn't charge the same for both of them.
The fun insight was that you would think that at a IaaS, PaaS company, that revenue is, like, the most important problem to work on and directly impacts the company's revenue and valuation and all that. No engineers wanted to do it, and I was like, "Why?"
Like, it, it, uh... We actually looked around for a long time. We tried to hire, and then we couldn't hire, so we ended up, uh, putting one of our most senior engineers on it, and she took a year to, to ship the whole billing project.
That was presumably a board-level objective, which was like, "Hey, let's change from this pricing plan to this pricing plan. How hard can that be?" Turns out very hard, 'cause you have to instrument everything, even the things that you're like, "I, I don't know if we'll ever use this."
Yeah, um, that, that, that's, that, that's exactly what I, what I meant, uh- That's why storage is free. Yeah, but the, the first thing is, like, you need to know about everything that's happening, uh- Yeah ... inside a system- Even your logs ...
inside a cluster. How big are your logs? Yeah. It's, uh, it's crazy hard. The pricing is, like, actually the... not the figuring out the business model, but the integration of it and implementation is, is actually a lot- Yeah ...
a lot of engineering. Soft limits, like, hard limits. Like, do you cut off people when- once they bust the limits? Probably not- Yeah ... 'cause they get pissed at you. Yeah. And then they also get pissed at you if you don't cut them off, 'cause then you send them a big bill.
Yeah. So you just... There's no winning. Uh, yeah, exactly. Um, also additional problem that people are asking us, like, "I want to run the agent for five hours," and you can do that, and you might be a very early startup.
So, like, our goal isn't to, like, cash you out. Our goal is, like, you can use as much of E2b as you can and just grow. But the more longer you run the sandboxes, the larger is going to be your bill.
So I think that's like the... It, it doesn't really needs to correlate with how much product market fit you have, for example, how much users you have. Because with these agents, they can work for a long time, even if you are, like, pretty early-stage company- Yeah ...
if you have few users so. Sure. I mean, yeah, the, I think... But there's, there's still a question of soft limit, hard limit- Yeah, yeah, yeah ... that, that kind of stuff. Um, so just to answer your question on, like, billing, uh, um, for those who are interested, we actually had the CTO of Orb send in a, a talk for our remote track for the New York Summit on, like, what he thinks pricing for agents looks like.
And I think basically you are reselling, uh, tokens, right? The, the, the base layer is coming from either your open model cloud provider or your closed model lab API, and you resell them, and a lot of them, some people have a lot of markup on them, like certain unnamed AI builders, and some people have negative markup on them.
Like, they, they are basically selling you at a discount. You should buy as much as you want because they are using VC monies to subsidize your thing, and, uh, I, I think that, that seems fine. People often... Simon Willison often asks for, like, bring your own key solution where, like, you know, I will have my control over, like, my relationships and my pricing and my credits with my LLM providers.
Uh, well, but I want to use your app, so I'll give you my keys, and then you use the keys on, on, on behalf of me. Uh, that doesn't seem to be as popular as I, as I think, as I understand it.
I don't know if you've- Yeah ... had that request. It's, um, it's sort of like in crypto. Use your own hardware wallet are always gonna be, like, less popular than just going, uh, with the thing that's- Yeah, yeah.
So, so literally, uh, Alex from OpenRouter is the only person who's implemented this- Yeah ... and, like, it's fine for, you know, in- individual use cases 'cause there's a lot of free tiers for, for, uh, individuals. But y- yeah, yeah, I mean, I think pricing-wise, people are trying to move that discussion on, like, you know, are you positive margin on your tokens?
Are you negative margin on your tokens? And, um, there's some economic reality there, but they're trying to move that to the agent work, which is what is the value of the human labor you're replacing? Mm. Which is a whole different thing, right?
So, like, instead of comparing on cost of goods sold, you're comparing on value delivered- Mm-hmm ... and that is much higher. Yeah, I feel like if we could get to a good market on, like, bid ask of, like, work being done, and then you can kinda arbitrage how many tokens you need.
But I think today the, the agents are so unreliable and unpredictable that, like, it's hard to price ahead of time because you could price any software engineering task per task, right? It's like, do a new bun, it's, like, $500, and then I can arbitrage that, but today there's no certainty of that- Yeah ...
Forking36:24
and I'm curious. Yeah, you would think, like, uh... Also, this is why I was very excited about Replit when they first launched, uh, their marketplace credit thing. I'm not... If Replit can't do it, then I don't know if anyone can.
Right. Yeah, yeah.
The other technical thing we talked about before is forking. You also mentioned it before, and sandbox checkpoints and things like that. Is it something you have today, and then what are people using that for? Like, our example was the Cloud Plays Pokémon Hackathon.
Um, are there more enterprise use case where you see people request forking and checkpoints?
Yeah, we don't have this publicly now, uh, yet, uh, but that's something we are working on to release somewhat soon. Um, we have persistence, which is like the prerequisite, the core-
The amount of volume.
Yeah, yeah. And well, amount of volume, but also memory persistence-
Okay
... which is, uh, very interesting. So you basically can pause the whole sandbox while even when all the code is, you know, with all the context of the code execution and resume to it, uh, resume it later and come back to it, I don't know, like, uh, two weeks later-
Interesting
... and it's still gonna be there.
But the continuous session time is limited to 24 hours, right?
No, that's limited in the beta in, in 30 days.
Okay.
Do you mean... Are you asking about-
I don't know. I was looking at your pricing page. It said 24 hours.
Are, are you asking about, like, when the sandbox is running? The sandbox can run up to 24 hours, but when it's paused, it can be paused for like a month. I, I personally think this is, like, one of the cases where you kinda need to show the developers why it's useful, and this is so- something I think is gonna be very useful as the agents are getting better and LLMs are getting better because you will be able to paralyze problem-solving, essentially.
So you will, instead of having like a single agent doing one thing, you might have multiple agents trying different paths, and if you imagine like a tree or a graph, every node is like a sandb- uh, snapshotted sandbox, like a checkpoint sandbox, and then from that node, we fork the sandbox and go to the next state.
Eventually, you find the right path, right? It's kinda like a... It, it's a tree search. So the forking and checkpointing sol- solves the local state problem. It doesn't solve the, you know, remote state because that's something you can't really control, but it solves a, a local, uh, local state for the agent and can then come back to a state, and you don't need to, you know, replay the whole session or trying to force the LLM to do the same thing again.
You just ha- you, you have the shared history, uh, and you have even the, all the steps inside the sandbox that went to it.
Do you feel like you want to help with, like, the forking and then the remerging? Because I think people understand the forking, but then it's like, okay, how do I monitor which of the leaves-
Yeah
... is successful, and then how do I merge that back in the thing? Do you think that's something that you wanna help people do, like kinda spread out, parallelize, and then find the winner, or is that something people should do on their own?
Frameworks39:18
I think this is a sorta like a framework discussion on top of E2E, so this is some, like... We, we are looking at it. I think eventually, like, we should go more higher level, and I don't know if, like, framework is the right type of a thing.
Mm-hmm.
I, I... We, we like to, in the team, think about it as a toolkit. Like instead of like building opinionated framework, how to build agents, we give you, um, sorta like a wrapper around E2E that makes it really easy for the LLMs to, for example, merge these states or navigate this tree.
So I think, uh, it, it, it will eventually move there. I think, um, it's a good question to ask, like how it's going to look like.
Mm-hmm.
I still think, like, building a framework is very hard in the AI as, as things are moving very, uh, very fast.
Curious about frameworks. Do you see any rise in popular frameworks that we should be keeping tabs on?
Well, I have a, I have... I don't know if this unpopular opinion, but, like, people keep telling me, uh, LangChain isn't popular, but if you look at- ... the stats, like it has 20 million downloads per month.
Yeah.
How can you have not popular framework when it does, it has like 20 million download, and it's growing, I think? Um, I think there's, there's a slight bubble, uh, in, I don't know if it's like people in SF, but like developers thinking LangChain isn't popular or used.
There's one framework that's interesting. I think it's called, uh, Mastra.
Yeah. Mm-hmm.
Um, from the, one of the founders of-
Sam Bhagwat. Yeah.
Yeah.
He's in the Discord.
Yeah.
In the SP.
Yeah, yeah, yeah.
Yeah.
So that, that I, I like, which looks very promising. I also like it's like TypeScript first, uh, which I thought like is the right decision because as, as we talked about, like m- being more bullish on the product developers instead of like a Python machine learning type of a developer.
I've seen a lot of more like a toolkits or I don't know how to call them. Like, like for example, Composio-
Mm-hmm
... has been one that, like gives you all the tools advanced. There have been more tool, more tools like that, which is interesting approach. I think also browser-based is Stagehand is super exciting because in my head it's not a framework, it's like it doesn't necessarily like dictate how the agent should behave.
It just gives it the good tools-
Mm-hmm
... to navigate a website, and it's very eleg- elegant. Uh, like it has like th- I think they added like three methods, right? Act, see, and, and something-
Yeah, there's three APIs.
Yeah, three.
Is observe, I think.
Yeah, observe. Yeah, yeah, yeah.
Yeah. We talked about it in that episode. No, it's cool. Uh, actually I wasn't expecting that many names to come out, but these are good names for people to know. I would agree with most of them as, as, uh, they're in the conversation of, of tooling if, and a lot of people listen to us for like, "Oh, what's going on in SF?"
Yeah.
Right? Like, yeah.
Actually, it's a good question because like when you ask it, I would ex- expect to be like bigger framework boom out this, because like it would make more sense to build, more sense to build a framework now instead of '23 because things are a little bit more stable.
Yeah.
Big problem with frameworks is that the idea of frameworks should probably be that things aren't changing underneath your hands, which like if you launch your framework in '23, uh, it was. Uh, so that's not a fun thing to be at as a developer who's using the framework.
I think things are still changing.
I mean, still, things are 100% still changing-
Like-
... but slightly, I would say like some things are clearer than in '23.
New Protocols42:35
Like I don't think people realize, but like I'll just say it out here, like chat completions is dying.
So, like, any framework that was built in that era with, like, no conception of real time, no conception of omnimodal or multimodal native things, they will probably not age very well.
Yeah. Yeah, probably as we-- like, what's the next after chat-based interaction-
Right, right
... with LM?
Yeah, you, you need chat completions, and then you also need reasoning with, with streaming. And the streaming interactions of agents is also not mapped out very well.
Yeah.
But, like, all agent frameworks will have to adjust to that, basically.
Good question to ask when thinking all about dev tools with LMS, is my dev tool more relevant as the LMS are getting smarter and as people need less prompting? I'm, I'm, for example, like, really bad at prompting, but, like, I can get more work done over the years because the LMS are getting better.
Uh, so that means, like, there's a less need for prompt management type of a thing-
Yeah
... probably.
The, uh, talk from Ramp at our New York conference was basically about this. Like, how do you set up your inf-- or your architecture so that you benefit from-
Yeah
... 10,000X improvements in models rather than, uh, every time you, you know, model improves, you have to kinda throw out ex- your existing workflows.
Yeah, the question we ask very often when thinking about the features and, um, uh, like, how to position ourselves in the ecosystem, essentially.
We've gone 51 minutes without talking about MCPs, um, which is tragic.
Yeah. How do you do that?
Uh, yeah .
How, how do you talk about AI infrastructure and no MCP?
Right? It's kinda crazy. But since you mentioned the prompts, we just had the episode with the MCP creators, and they were, I wouldn't say not frustrated, but maybe, uh, they hope more people will use the prompts and resources in MCP servers instead of just the, the tool costs.
I'm curious if you've seen any fun use cases for, like, remote MCPs on, on E2B or any other things like that.
Yeah, MCP is watching it very closely but, like, still undecided what actual, like- ... not what it is, but kinda like- ... what to do with it. Um-
There's no MCP on your docs, bro.
Uh- I mean, like, go to-- if you go to our GitHub, we have, like, MCP server, um-
Okay
... but yeah, we, we've seen people using E2B to, uh, host MCPs. Like-
Mm
... but then I would say, like, I don't think you need E2B for that. You can just, like, probably you can run ins-- uh, it's not, like, even, like, optimized for it, probably. It's not the, like, the, the...
There are probably better ways to do that. I think people are a little bit too much focused on the protocol side of it. Even, like, calling it a protocol, I don't know. It's just like a server. Uh, and there's like a-
It's server and client.
Yeah, server and client. And there's some agreed... So I guess, like, in a sense, it is a protocol, but, like, uh, I see a lot of people, like, comparing to email protocols and such, which seems a little bit far-stretched to me, uh, at this current moment.
But maybe I'm missing something. Like, I'm... I think it's super interesting idea. I just haven't had the, like, I guess, the right insight about it yet.
Mm.
One last thing I wanted to add, like, because we have a bunch of users that added E2B, our MCP server, to their, like, registries. Uh, so I think, at least at the moment, like, it's, what's more useful is, like, higher order MCPs.
I was talking with Henry from Smithery around it, about it, uh, that, that you know, and he was saying exactly this, uh, telling me exactly this concept, like, it's unclear who's using the MCP at the moment. Is it a developer?
Is it a, like, end user? Or is it another agent? If it's a developer, then it might be, like, make sense to have like a sandbox, creation of sandbox in the MCP. But if it's a end user, probably, like, it's too low level primitive, so he wants some kind of higher order MCP.
Like, instead of us offering like a sandbox, we would be offering like a way to build a code gen agent with MCP or something like that.
Mm.
Or run code gen agent that's using E2B MCP, like, uh, in the background. So I have more like a un- un- like these type of unanswered questions in my head about MCPs.
I mean, they are all running locally right now mostly. I don't think there's many remote-
Well-
... MCP servers yet. I know that people are pushing for it, but-
I, I asked and, and it looks like they're more remotely, actually. I don't know. This is what I heard from, uh, people, like, uh, managing these r- registries of MCPs.
Well, they're incentivized to tell you that.
Yeah, yeah.
Um, I fully agree with that confusion about what to do. I do think that every dev tools company needs some kind of MCP strategy, for better or worse. It's annoying as it is. Um, I think it does start with having an API, though, rather than like a SDK-first experience, because then you...
people can just wrap in whatever language they, that they want. And then you also, for you particularly, you might want to have a distinction in your strategy versus, for MCP clients versus MCP servers, because those might be different things.
A- a- particularly, I- I like what you said about the higher order MCP for the, for the, uh, end user who doesn't really care about implementation detail. I think that is fantastic for E2B, where, like, the agents can just, um, can just spin up an E2B instance in the, in the background, and they don't even know about it.
Yeah. My... I- in ideal world, we have, like, mcp.e2b.dev. First, it needs to have figured out authentication.
Yep.
So the agent should just ask to have, I guess, account created for it, uh, whatever it is-
Yep
... it's gonna be. And then it can just, like, launch a sandbox, do the, uh, execution there, and it can come back to it month later, and the state is still there. So that, I think, makes ton of sense.
Then you can start building higher order things on top of that, uh, which could be very interesting. I like what you said about, uh, having API first versus SDK first approach. I think this is very, very in, uh, important for the LLMs, and that's how we are building our whole new dashboard, the infrastructure.
So everything needs to be API controllable from, like, public, um-
Exactly
... APIs for users, because eventually the LLMs will want to control, like, get all this data.
Okay, so that, that's a big shift for E2B, 'cause you're SDK first.
It has been, like- Underneath, uh, API based-
Of course
... all the time. But, uh, now we want- Yeah, we will go more into it and it's, it was more like, uh... In terms of prioritization, so we needed to start with humans to get to LLMs, and first build for human developer and then now building for, like, a LLM developer first.
Yeah. I'll just call out that since we did the MCP episode, they announced, uh, their update to the spec that they added an auth component to the, to the spec itself. It seems to be just based on OAuth 2.1, and, uh, I assume, like, you know, uh, that's the first easiest thing to, to do, but, like, there's no effective distinction between an agent and an end user-
Agent Web49:22
Yeah
... because we never had that.
I think this touches more, like, a broader question of you have, like, all the websites that has optimized everything for humans, but now you will have, like, agents visiting those websites. Like, what are the incentives there, you know?
Like, probably you as a website owner, you want to know it's a agent because you spend so much time and money optimizing everything for humans. So I think, like, the, the dynamics in, on the internet get, might get really weird if you don't have a distinction, this is a human, this is agent.
Yeah. We, we tried to record an episode with Matthew Prince, CEO of Cloudflare, yesterday, and then we had technical difficulties, but they're building a lot of that.
But you can see the stats. The, the stats are-
Yeah, yeah
... pretty interesting.
They mentioned, for example, used to be, you know, Google would be a 2:1 crawl to referral ratio. So for every two pages they will read, they will send you one visitor.
Yeah.
He said OpenAI is 250:1, so they'll read 250 of your pages and send you one person, and Anthropic was, like, 6,000:1.
Wow.
So they'll read 6,000 pages before they send one person back to your website, so. Obviously we have Jeremy Howard, who's been working on LLMs.txt to kinda have a separate interface for that. I, I think today people don't really curate the LLM experience.
Uh, I, I think a lot of the LLM.txt that people are making is, like, automated, take a website and turn it into LLM.txt, but, like, that's not really what you wanna do.
Yeah.
It's like, how do you separate completely the two things, you know? Like, if you're a e-commerce store, the LLM.txt should have your own inventory in, like, one thing.
Yeah.
You know, shouldn't have a search button.
I feel like obviously LLMs.txt is a good movement that improve the legibility of these doc sites for LLMs, but I feel like it's kind of like a halfway measure. Like, I think we've, we've failed with agents if we are reshaping the human environment for agents.
Mm-hmm.
Like, there must be two of everything, the, what's your human side and what's your agent side? Like, no, like, agents should just be the human side. Like-
Yeah
... why, why are we making-
Yeah
... any special dispensation for these things?
Yeah. I think the monetization is the only thing. Like, too many website are, like, ads-driven, kinda like-
Yeah, yeah
... full of-
You need to look at it, right?
Yeah. I think you need to change. Once we figure that out, you know, I think you can use the same interface.
Yeah.
But I think today it's just, like, all these pop-ups, it's like-
Okay, so-
... "Fuck me. Stop popping up"
... Bitcoin solves this.
Yeah.
I'm just kidding. Um, but anyway, yeah.
Yeah.
I, I don't know if you have a, a take on all this.
I, I have, like, a general rule that there's, uh... Usually when something new comes out, people have this tendency to recreate the thing that already exists for the new thing, like, that's, like, the internet for agents.
Yeah.
Versus so you already have all the infrastructure for the old thing and probably might be easier to teach the new thing to use the old thing. I think it, it's very common in dev- from, coming from developers, that it will be a perfect world when you have nice clear distinction between these two.
But actually, the world is super messy, so nothing comes to my mind that they will, like, end up that you have clear distinction between, like, two type of sort of entities or we created, like, a new internet for, for just mobile phones.
You actually have both website, like, desktop website version and mobile phone version, right? It's much more always, like, complicated and messy in the real world than having, like, nice clear distinction, which is something I think developers strive for, like, just from working with code because you want these clear, nice clear distinction.
But, uh, humans are more complicated than that. So, so I think, like-
Yeah
... uh, everything ends up, uh, being sort of mixed, and you will need to adapt to it.
For sure. Yeah. I think we'll do this for 20 years, and then we'll figure out the-
Yeah, yeah
... merging.
And probably it, it will be somewhere in between that you have, like, a world-
Yeah
... where agents are using human internet, but the internet change because you have agents or something like that. Uh, this reminds me of, um, some conversation that I think it's... I don't, I don't know who, who was re- making this analogy.
Advanced Uses53:00
Like, in the mobile era we had m.yourdomain.com, and then we had www.yourdomain.com, and now we c- we might have, like, just llm.yourdomain.com.
Yeah.
And that's just the LLM experience.
MCP, the-
Oh, yeah. Way better, way better. Yeah, yeah. Consumes everything. Cool. We're, we're just gonna go through all the rest of the use cases. We can refer people to your website, but, um, I'm very just-- I, I always want people to have a good mental map of how, when they should go to E2B and, like, or what other people are using E2B so that they don't miss out, right?
So it's AI data analysis, data visualization, coding agents, generative UI, codegen evals, and computer use. Do you think that's, like, the sequence of most popular is data analysis, least popular is computer use, right?
Most experimental is computer use, I would say.
Uh, yeah.
Uh, it still gets a lot of traction when you share, uh, you know, the demos of it.
Is it just Manus? Like, who else is doing this? I d- I basically don't see anyone else.
So, uh, when you say computer use, like, I imagine, like, a graphical interface.
Uh-huh.
Manus, uh, as far as I know-
Manus doesn't, yeah
... yeah, it doesn't, isn't, isn't running that. Um, so would also call it kinda computer use but without graphical interface because you are using the full computer but more from, like, code point of view.
Sure.
But that's a different discussion. I, I would say, like, computer use is very experiment- exciting but experimental from what I've seen.
Yeah. Okay.
I, I think for real computer use, you really want to support more platforms than just Linux-
Yeah
... over time.
Windows.
Yeah.
Right? And Mac.
And, and then it gets... It might be more a licensing battle than technical battle.
Yeah. Lawyers always win.
Yeah.
Um, we'll probably have Eric from, from Pig at some point, uh-
Yeah
... talk about his, his, uh, movements on Windows. So I was gonna go into evals, right? And also how that links to RFT. Yeah, just, like, t- can you tell more stories about the OpenR1 projects, how you work with them, and any other, like, academics that are working with E2B that can-
You think could be possible for like the, the research or train- like model training use case or finetuning use case?
Yeah. The way Hugging Face will build, um, the Open R1 project is using us is during like the reinforcement, uh, learn- code gen reinforcement learning step, where the R1 model, uh, the Open R1 model has a training s- step where, uh, they give it a code problem and the model needs to generate and run code somewhere.
Then you have reward function basically giving you zero or one, telling you if that was a good solution or bad solution, then you improve the model. So kinda like the feedback loop. They are using the E2B sandboxes to, you know, run many hundreds of these sandboxes, thousands of these sandboxes per, uh, training step, so they can achieve like big parallelization.
We started very fast. You don't need to use your GPU cluster for that, uh, which is like very expensive to, to use for these type of workloads. And also, and that, that goes with the, with the story that I mentioned in the beginning, is that you don't need to worry about like the LLM actual like changing permissions in your cluster, and then you can't access the cluster.
Because everything is isolated and is secure from, from i- from each other. So I think that's very, very interesting use case because we've had a few other companies like, uh, reaching out and, and, um, us- started using us this way, building models, foundation models.
It's like when we started E2B, that wasn't the use case that we had in mind, so but it's like makes total sense if you, if you... Also, if you think about like life cycle of an AI agent, like if you-- it makes a lot of sense for us to be like from the earliest stage possible, and the earliest stage is probably model training.
So this fits very, very, very nicely in that, uh, in the use case. We actually released a case study with Hugging Face that's on our website, uh, people can, can read.
Have you seen people also use that to evaluate agents they wanna use, or is it mostly people doing training?
We've seen people using E2B for evals.
Yeah, that's what I'm thinking, like it should be very easy to run like Sweet Bench and all these on, on E2B.
Yeah. This is more foreshadowing, but we will be launching, um, in the next couple of months, uh, like a startup and research program for people and universities and researchers like doing exactly these type of things. A different use case, but we, for example, work with LMArena folks from, from Berkeley that are using us to compare models, uh, in AI app generation, and we run the AI-generated app.
He wrote that integration, right?
Yes. Yes.
I did. I did.
That's, that's, uh, I, I think like that's only possible thanks to Alessio, and I'm not joking, like because like first he connected us, and then he actually wrote it.
Man, the things you have to do to win deals these days.
Full stack value add.
Well, like I think, uh, the-- he was already our investor-
Okay.
... so I think that also shows it even in a better light because you clearly see the person is not interested just to win the deal, but actually, you know, um-
To increase the value of the investment.
Value add.
Exactly.
Talk about value add. Like how many, how many VCs are actually fixing your bug in your P- uh-
Yeah
... in your code base?
Let's make a YouTube Short of this part- ... so I can share it.
You should put it on our, on our landing page.
Yeah, exactly. Um, also, man, so let's talk about since you mentioned VCs, um, a lot of VCs that passed on you before because you only do code execution. It's kinda like a small market. I, I would love for you to maybe also paint the picture, you know.
Roadmap58:12
So you just mentioned it's cheaper than the GPU cluster than Hugging Face has. Do you wanna do GPUs in the future? You mentioned Railway, uh, that was easy to do. Do you wanna compete with Railway down the line?
Like where do you see, um, E2B going?
So GPU question is an interesting one. Um, GPU market is hard. Like you are competing on the compute internal I think is hard. Because, uh, eventually you will have competitors, and everyone will be like pricing it a little lower, and no one is making any margin.
So that's like, uh, but it just opens new use cases for you. And for example, even with the data analysis, if you just run, I think it's Pandas code from-- there's a recent up- like recent update. If you just run it on GPU, it's like twice as much faster.
If we want to like do really big AI data analysis, you need that. So I think like GPUs for us and also if you want to like have LLM train small machine learning models, uh, maybe you want to like, uh, build full games, you really want to offer the full cloud, like computer but in cloud, uh, that's very elastic.
So I think GPUs are there on the roadmap. Uh, I wouldn't say it's like the thing that we are immediately, uh, working on, but eventually it makes like a lot of sense for us to, to offer this. And, uh, sorry, what was the other question?
Uh, on the-
Yeah
... do you wanna host the apps-
Yeah, yeah
... that people are building to-
That's, uh-
You know, with-
We, we are very well-positioned that the LLM is doing all the development work with us. Then you need to deploy it somewhere. So it's like very next, natural next step. Eventually, like we want the LLMs to deploy these services, apps that they are building and, uh, have them manage it, and developer is more like in the back seat, like looking at things if everything is working correctly.
If your swarm of agents is working correctly. It also requires a slightly different infrastructure for deploying, but I think there's a big advantage in knowing what developers are building on our platform, and then because then you can see like what's the ideal, even from technical point of view, you have all the insights that you need to actually effectively deploy it and kind of then cover the full life cycle of building, building the app.
But now it's not built by the human developer, it's built by the AI. So I-- like TLDR, yes, uh, probably down the road somewhere, like you-- we want to build essentially like the new AWS but for LLMs.
So we're just gonna move out and, you know, just to, to wrap up. Uh, one of the interesting things that I saw you do was you were originally Czech-based And, uh, I'll, I'll be very blunt. When I first, uh, invested in, in the early round that you did, I was like, "These guys know how to recruit in Czech Republic.
It's like a, it's like a competitive advantage." And then the next thing I know, you're moving to SF with your whole team.
You show up to our office. That's hilarious.
Um-
I actually, yeah, actually, I think you, um, I, uh, you offered even your place, like, to live for, uh-
Yeah
... for a few days. So-
So, so why move to SF? Do you think that everybody in your sim- si- kind of similar situation should? Any pros and cons that you're experiencing?
I think it's, uh, you can definitely build a dev tool company in, from Europe. I think it's a lot about question of how easy you want it to be in the earlier days. Especially if you are building like a, you know, you know, kinda like red ocean versus blue ocean waters.
If you are building in, uh, a field that already exists, and you have, like, large competitors, you are building something that's 10 times, 100 times better, it's probably, you probably don't need to be in SF. Uh, you eventually probably will need some kind of US base because for sales and customers, but you can very well build this, uh, from Europe.
I know great companies doing that, uh, because all, it's the knowledge is already among all the developers. But I would maybe argue that it might be harder to find people that are comfortable with a fast iteration loop and changing things early on.
You know, like, almost you are pivoting every week, every month. But the main motivation for us was we just wanted to be very close to our users, and it was clear after few weeks that SF is becoming this AI hub.
And what we, uh, used to do, and we, we still, we still do it sometimes, but slightly less because of, like, not, uh, having that much time and resources. But we just met with the customers that had problems, our users, and we just, like, implemented E2B for them, like next to them.
Like, we made-
Yeah
... a PR.
I call this the collision, collision-
Yeah
... installation. Yeah.
Um, by, by, by the way, I don't know if this is pub- like a known thing, but do you know how many times they did that?
200? I don't know.
Twice, three times.
Oh, long.
I ask, uh, about it like three years ago when I had a chance to ask a question from Patrick Collins, like, "How many times you did the Pa- collision installation?" And he was like, "Three times. But then after that, you probably don't want to do that because you want to automate things and focus on other stuff."
Nice.
Uh, but, um, it's exactly like, do the things that don't scale.
Three times. Did he scale it?
Um, exactly three.
But my whole point was that how many such users I can meet in Prague in a week versus in San Francisco in a week. In Prague, it's gonna be probably one, and then I can't meet anyone else for the next half of a year, uh, because I just don't have users in Prague.
But all of my users are users who are here, so we could just keep repeating doing that again and again. I would even argue we did it too much. We, we could have automated, uh, slightly faster. But it's, it's very useful feedback that you can get, and you can, uh, then start moving much, much, much faster.
Yeah.
Uh, so that was important. And I think, like, also, eventually you are in, like, a, you are in the B2B business even from the early days, and it's good to have good relationship with people and just, like, meeting in person is just better than meeting in, uh, over Zoom.
Well, I mean, that's why we do this in person.
Yeah.
But yeah, I mean, I, it's, it's, uh, you know, obviously I run a conference. I'm very sympathetic to people meeting in person, right? But I also want there to be some hope for people who are, who are never going to come to SF that, like, they can still get involved.
That's partially why we do this podcast is to-
Yeah
... get them involved in the community.
And I mean, we started a new office in Prague, so-
Yes
... um-
Yes. You're hiring again in Prague.
Yeah.
Yes.
And I think that there's really good talent in Europe, in Czech Republic, for example. I don't, I can't speak for other countries, but I can imagine it's very similar. Once, um, uh, once you have, like, a clear idea of what your product looks like, then, um, you can find really good expert on certain part of your infrastructure, on, on database and things like that and just, like, have, like, top talent, uh, get top talent for that.
The reason we didn't want to do it early on b- was because we kinda didn't know ourselves what we were building. Uh-
Mm-hmm
... and we had to figure it out in person with users here. But now we feel much more strong about, like, knowing the roadmap for the company and for our product. And so it's, it's much easier to, you know, hire people that have 8 Hour difference from us and, uh, explain them, like, what they are building even remotely sometimes if, if you are not there and you're just communicating through Slack.
Just to wrap, what are the roles that you're hiring for?
So we are hiring distributor systems engineers. We are hiring platform engineers, um, AI engineers. We are also, um, also hiring account manager and customer success engineer. Uh, so kinda all over the place. Uh, we see a lot of market pull and momentum being built up, so we want to double down and move even faster because we see all the potential where what we can do and just want to kinda like you want to pour the gas on the fire on the spark, and that's, uh, how I feel w- where we are now at E2B.
Awesome, man.
Very cool.
Thank you so much for coming on.
Yeah. Thank you for having me. It's great.






