Intros0:00
Hey everyone, welcome to the "Late in Space" podcast. This is Alessio, partner and CTO and resident at Decibel Partners, and I'm joined by my co-host, Zwickx, founder of SmallAI.
Hey, and today we have in the studio Erik Bernhardsson from Modal. Welcome.
Hi, it's awesome being here.
Yeah, awesome seeing you in person. I've seen you online for a number of years as you were building out Modal, and I think you were just making a San Francisco trip just to see people here,right? Like, I've been to like two Modal events in San Francisco here.
Yeah, that'sright. We're based in New York, so I figured sometimes I have to come out to, you know, Capitol of AI and make a presence.
What do you think is the pros and cons of building in New York?
Uh, I mean, I never built anything elsewhere. Like, I lived in New York the last 12 years. I love the city. Obviously, there's a lot more stuff going on here, and there's a lot more customers, and that's why I'm out here.
I do feel like for me, where I am in life, like, I'm a very boring person. Like, I kind of work hard and then I go home and hang out with my kids. Like, I don't have time to go to like events and meetups and stuff anyway.
So, in that sense, like New York is kind of nice. Like, I walk to work every morning. It's like five minutes away from my apartment. It's like very time efficient in that sense.
Yeah, yeah. Sounds like a good life. So I'll do a brief bio, and then we'll talk about anything else that people should know about you. So you actually, I was surprised to find out, find this out, you're from Sweden.
You went to college in KTH.
Yep, yep, Stockholm.
And your master's was in implementing a scalable music recommender system.
Yeah.
I had no idea.
Yeah, yeah, yeah, yeah. So I actually started physics, but I grew up coding, and I did a lot of programming competition. And then, and then like as I was like thinking about, you know, graduating, I got in touch with an obscure music streaming startup called Spotify, which was then like 30 people.
And for some reason, I convinced them like, why don't I just come and like write a master's thesis with you, and like I'll do some cool collaborative filtering. Despite not knowing anything about collaborative filtering really, I sort of, you know, but no one knew anything back then.
So I spent six months at Spotify basically building a prototype of a music recommendation system and then turned that into a master's thesis.
Yeah.
And then later, when I graduated, I joined Spotify full-time.
Yeah, yeah. And then, so that was the start of your data career. You also wrote a couple of popular sort of open source tooling while you were there. And then you, is that correct or?
Spotify Roots2:08
No, that'sright. I mean, I was at Spotify for seven years. This is a long stint. And Spotify was a wild place early on. And I mean, the data space is also a wild place. I mean, it was like Hadoop cluster in the like foosball room on the floor.
And, you know, so like it was a lot of crude, like very basic infrastructure, and I didn't know anything about it. And like I was hired to kind of figure out data stuff. And I started hacking on a recommendation system, and then, you know, got sidetracked into a bunch of other stuff.
I fixed a bunch of reporting things and set up A/B testing and started doing like business analytics, and later got back to music recommendation system. And a lot of the infrastructure didn't really exist. Like there was like Hadoop back then, which is kind of bad, and I don't miss it, but I spent a lot of time with that.
As a part of that, I ended up building a workflow engine called Luigi, which is like briefly like somewhat like widely ended up being used by a bunch of companies. Sort of like, you know, kind of like Airflow, but like before Airflow.
I think it did some things better, some things worse. I also built a vector database called Annoy, which is like, for a while, was actually quite widely used in 2012. So it was like way before like all this like vector database stuff ended up happening.
And funny enough, I was actually obsessed with like vectors back then. Like I was like, this is gonna be huge. Like just give it like a few years. I didn't know it was gonna take like nine years, and then it's gonna suddenly be like 20 startups doing vector databases in one year.
So it did happen. In that sense, I wasright. I'm glad I didn't start a startup in the vector database space. I would have started way too early. But yeah, that was, yeah, it was a fun seven years at Spotify.
It was a great culture, a great company.
Yeah. Just to take a quick tangent on this vector database thing, because we probably won't revisit it, but like, has anything architecturally changed in the last nine years? Or.
I mean, sort of. Like I'm actually not following it like super closely. I think, you know, the like, some of the best algorithms are still the same. It's like hierarchical, navigable, small world.
HNSW, yeah.
Exactly, yeah, HNSW. I think now there's like product quantization. There's like some other stuff that I haven't really followed super closely. I mean, obviously, like back then, it was like, you know, Annoy is like very simple. It's like a C++ library with Python bindings, and you could emmap big files into memory, and like they had some lookups.
And I used like this kind of recursive, like hyperspace splitting strategy, which is not that good, but it sort of was good enough at that time. But I think a lot of like HNSW is still like what people generally use.
Now, of course, like databases are much better in the sense like to support like inserts and updates and stuff like that. Annoy never supported that. Yeah, it's sort of exciting to finally see like vector databases becoming a thing.
Yeah, yeah. And then maybe one takeaway on most interesting lesson from Daniel Ek.
I mean, I think Daniel Ek, you know, he started Spotify very young. Like he was like 25, something like that. I don't know if it was like a good lesson, but like he, in a way, like I think he was a very good leader.
Like there was never anything like, no scandals or like, no, he wasn't very eccentric at all. He was just kind of like very like level-headed, like just like ran the company very well. Like never made any like obvious mistakes.
Or I think it was like a few bets that maybe like in hindsight were like a little, you know, like took us, you know, too far in one direction or another. But overall, I mean, I think he was a great CEO.
Like definitely, you know, up there, like generational CEO, at least for like Swedish startups.
Yeah, yeah, for sure. Okay, we should probably move to, make our way towards Modal. So then you spent six years as CTO of Better.
Yeah.
You were an early engineer, and then you scaled up to like 300 engineers.
I joined as a CTO when there was like no tech team. And yeah, that was a wild chapter in my life. Like the company did very well for a while, and then like during the pandemic.
Less well.
Yeah, it was kind of a weird story, but yeah, it kind of collapsed. And you know, then it actually went public too.
Laid off people poorly.
Yeah, yeah, there's like a bunch of stories. Yeah, I mean, the company like grew from like 10 people when I joined to 10,000. Now it's back to 1,000. But yeah, it actually went public a few months ago. Kind of crazy.
They're still around. Like, you know, they're still, you know, doing stuff. So, but yeah, very kind of interesting six years of my life for non-technical reasons, mostly like. But yeah, like I managed like 300, 400 people.
Management, scaling.
Yeah, like learning a lot of that, like recruiting. I spent all my time recruiting and stuff like that. And so managing at scale, it's like nice like now in a way, like when I'm building my own startup, like that's actually something I like don't feel nervous about at all.
Like I've managed at scale. Like I feel like I can do it again. It's like very different things that I'm nervous about as a startup founder. But yeah, I started Modal three years ago after sort of, after leaving Better.
I took a little bit of a time off during the pandemic. And but yeah, pretty quickly I was like, I gotta build something. I just want to, you know.
Yeah.
And then, yeah, Modal took form in my head.
And as far as I understand, and maybe we can sort of trade off questions. So the quick history is, started Modal in 2021, got your seed with Sarah from Amplify in 2022. Last year you just announced your Series A with Redpoint.
Modal Vision6:51
That'sright.
And that brings us up to mostly today.
Yeah.
And so like most people, I think, were expecting you to build for the data space.
But it is the data space.
It is the data space. When I think of data space, I come from like, you know, Snowflake, BigQuery, you know, FiveTrain, Airbubby, that kind of stuff.
Yeah.
And so, you know, what Modal became is more general purpose than that.
Yeah, yeah. I don't know, it was like funny. I actually ran into like Ido Liber, the CEO of Pinecone, like a few weeks ago. And he was like, I was so afraid you were building a vector database.
No, yeah, it's like, like I started Modal because, you know, like in a way, like I work with data like throughout most of my career. Like every different part of the stack,right? Like I've done everything from like business analytics to like deep learning, you know, like building, you know, training neural networks to scale, like everything in between,right?
And so one of the thoughts, like, and one of the observations I had when I started Modal, or like why I started, was like I just wanted to make, build better tools for data teams. And like very, like that's sort of an abstract thing, but like I find that the data stack is, you know, full of like point solutions that don't integrate well.
And still, when you look at like data teams today, you know, like every startup ends up building their own internal Kubernetes wrapper or whatever. And, you know, all the different data engineers and machine learning engineers end up kind of struggling with the same things.
So I started thinking about like how do I build a new data stack, which is kind of a megalomaniac project. Like because I kind of wanted to like throw out everything and start over.
It's almost a modern data stack.
Yeah, like a postmodern data stack. And so I started thinking about that, and a lot of it came from like more focus on like the human side of like how do I make data teams more productive, and like what are the technology tools that they need.
And like, you know, drew out a lot of charts of like how the data stack looks, you know, what are the different components. And it was actually very interesting like workflow scheduling, because it kind of sits in like a nice sort of, you know, it's like a hub in the graph of like data products.
But it was kind of hard to like kind of do that in a vacuum, and also to monetize it to some extent. And I got very interested in like the layers below at some point. And like at the end of the day, like most people have code they have to run somewhere.
And I started thinking about like, okay, well, how do you make that nice? Like how do you make that, and in particular, like the thing I always like thought about like developer productivity is like I think the best way to measure developer productivity is like in terms of the feedback loops.
Like how quickly when you iterate, like when you write code, like how quickly can you get feedback? And at the innermost loop, it's like running some, like writing code and then running it. And like as soon as you start working with the cloud, like it's like takes minutes suddenly, because you have to build a fucking Docker container and push it to the cloud and like run it, you know.
So that was like the initial focus for me. It was like, I just want to solve that problem. Like I want to, you know, build something less, you run things in the cloud, and like retain the sort of, you know, the joy of productivity as when you're running things locally.
And in particular, I was quite focused on data teams, because I think they had a couple of unique needs that wasn't well served by the infrastructure at that time, or like still is. And like in particular, like Kubernetes, I feel like it's like kind of worked okay for backend teams, but not so well for data teams.
And very quickly, I got sucked into like a very deep like rabbit hole of like.
Not well for data teams because of burstiness.
Burstiness is one thing, yeah, for sure. So like burstiness is like one thing,right? Like when you, like, you know, like you often have this like fan out, you want to like apply some function over very large datasets. Another thing tends to be like hardware requirements.
Like you need like GPUs. And like I've seen this at many companies. Like you go, you know, data engineers go to like, or data scientists go to a platform team, and they're like, can we add GPUs to the Kubernetes?
They're like, no. Like that's, you know, complex. We're not going to, or like, so like just getting GPU access. And then like, I mean, I also like data code, like frankly, or like machine learning code like tends to be like super annoying in terms of like environments.
Like you end up having like a lot of like custom like containers and like environment conflicts. And like so it ends up having a lot of like annoying, like it's very hard to set up like a unified container that like can serve like a data scientist, because like there's always like packages that break.
And so I think there's a lot of different reasons why, you know, the technology wasn't well suited for backend. And I think the attitude at that time was often like, you know, like you had friction between the data team and the platform team.
Like, well, what works for the backend stuff, like why can't you use it? Like, you know, why don't you just like, you know, make it work? But like I actually felt like data teams at that point, you know, or at this point now, like there's so much, so many people working with data.
And like they, to some extent, like deserve their own tools and their own toolchains. And like optimizing for that is not something people have done. So that's sort of like a very abstract, philosophical reason why I started Modal.
And then I got sucked into this like rabbit hole of like container cold start and, you know, like whatever, Linux, page cache, you know, file system optimizations.
Yeah, tell people. I think the first time I met you, I think you told me some numbers, but I don't remember. Like what are the main achievements that you were unhappy with the status quo, and then you built your own container stack?
Yeah, I mean, like in particular, it was like, like how do you, like in order to have that loop,right? Like you want to be able to start, like take code on your laptop, whatever, and like run in the cloud very quickly.
Container Magic11:56
And like run in custom containers, and maybe like spin up like 100 containers or 1,000, you know, things like that. And so container cold start was the initial, like from like a developer productivity point of view, it was like really what I was focusing on is, I want to take code, I want to stick it in a container, I want to execute in the cloud, and like, you know, make it feel like fast.
And when you look at like how Docker works, for instance, like Docker, you have this like fairly convoluted, like very resource inefficient way they, you know, you build a container, you upload the whole container, and then you download it, and you run it.
And Kubernetes is also like not very fast at like starting containers. So like, so I started kind of like, you know, going a layer deeper. Like Docker is actually like, you know, there's like a couple of different primitives, but like a lower level primitive is run C, which is like a container runner.
And I realized like, what if I just take the container runner, like run C, and I point it to like my own root file system, and then I built like my own file system, like virtual file system that exposes files over network instead.
And that was like the sort of very crude version of Modal. It's like now I can actually start containers very quickly, because it turns out like when you start a Docker container, like first of all, like most Docker images are like several gigabytes, and like 99% of that is never going to be consumed.
Like there's a bunch of like, you know, like time zone information for like Uzbekistan or whatever. Like no one's going to read it. And then there's a very high overlap between the files that are going to be read.
There's going to be like lib torch or whatever, like it's going to be read. So you can also cache it very well. So that was like the first sort of stuff we started working on. It was like, let's build this like container file system, and you know, coupled with like, you know, just using run C directly.
And that actually enabled us to like get to this point of like, you write code, and then you can launch it in the cloud within like a second or two, like something like that. And you know, there's been many optimizations since then, but that was sort of a starting point.
Can we talk about the developer experience as well? I think one of the magic things about Modal is at the very basic layers, like a Python function decorator, it's just like stub and whatnot. But then you also have a way to define a full container.
What were kind of the design decisions that went into it? Where did you start? How easy did you want it to be? And then maybe how much complexity did you then add on to make sure that every use case fit?
Yeah, like, I mean, Modal's, I almost feel like it's like almost like two products kind of glued together. Like there's like the low-level like container runtime, like file system, all that stuff like in Rust. And then there's like the Python SDK,right?
Like how do you express applications? And I think, I mean, Switch, like I think your block was like the Self-Revisioning Runtime was like to me always like the sort of, like the, you know, for me like an eye-opening thing.
It's like suddenly think about like, I want to.
You wrote your post four months before me.
Yeah?
The software 2.0, infra 2.0.
Yeah, well, I don't know, like convergence of minds. Like we're thinking, I guess we're like both thinking. Maybe you put, I think, better words on like, you know, maybe it's something I was like thinking about for a long time.
Yeah, and I can tell you how I was thinking about it on my end, but I want to hear it.
Yeah, yeah, I would love to. Like, and like to me, like what I always wanted to build was like, I don't know, like I don't know if you use like Palooma. Like Palooma is like nice, like in the sense, like it's like Palooma is like you describe infrastructure in code,right?
And to me, that was like so nice. Like finally I can like, you know, put a for loop that creates S3 buckets or whatever. And I think like Modal sort of goes one step further in the sense that like, what if you also put the app code inside the infrastructure code and like glue it all together, and then like you only have one single place that defines everything.
And it's all programmable. You don't have any config files. Like Modal has like zero config. There's no config. It's all code. And so that was like the goal that I wanted, like part of that. And then the other part was like, I often find that so much of like my time was spent on like the plumbing between containers.
And so my thinking was like, well, if I just build this like Python SDK, then, and make it possible to like bridge like different containers, just like a function call. Like, and I can say, oh, this function runs in this container, and this other function runs in this container, and I can just call it just like a normal function.
Then, you know, I can build these applications that may span a lot of different environments. Maybe they fan out, start other containers. But it's all just like inside Python. You just like have this beautiful kind of nice like DSL almost for like, you know, how to control infrastructure in the cloud.
So that was sort of like how we ended up with the Python SDK as it is, which is still evolving all the time, by the way. We keep changing syntax quite a lot, because I think it's still somewhat exploratory.
But we're starting to converge on something that feels like reasonably good now.
Yeah, and along the way,
with this expressiveness, you enabled the ability to, for example, attach a GPU to a function.
Totally, yeah. It's like you just like say, you know, on the function decorator, you're like GPU equals, you know, A100. And then, or like GPU equals, you know, A10 or T4 or something like that. And then you get that GPU.
And like, you know, you just run the code and it runs. Like you don't have to, you know, go through hoops to, you know, start an EC2 instance or whatever.
Yeah.
So it's all code.
Yeah, so on my end, the reason I wrote Self-Revisioning Runtimes was I was working at AWS, and we had AWS CDK, which is kind of like, you know, the Amazon basics Palooma.
Yeah, totally.
And then, but you know, it creates, it compiles the cloud formation.
Yeah.
And then on the other side, you have to like get all the config stuff and then put it into your application code and make sure that they line up. So then you're writing code to define your infrastructure, then you're writing code to define your application.
And I was just like, this is like obvious that it's going to converge,right?
Yeah, totally. But isn't there, it might be wrong, but like, was it like Sam or Chalice or one of those, like, isn't that like an AWS thing that, where actually they kind of did that? I feel like there's like one product.
Sam? Yeah, yeah, yeah, yeah. Still very clunky.
Okay.
It's not as elegant as Modal.
I love AWS for like the stuff it's built, you know, like historically in order for me to like, you know, like what it enables me to build. But like AWS has always like struggled with developer experience. Like, and that's been.
No, I mean, they have to not break things.
Yeah, yeah, and totally. And they have to, you know, build products for a very wide range of use cases. And I think that's hard.
Yeah, yeah. So it's easier to design for. Yeah, so anyway, I was pretty convinced that this would happen. I wrote that thing. And then, you know, imagine my surprise that you guys had it on your landing page at some point.
I think Akshad was just like, let's just throw that in there.
Did you trademark it?
No, I didn't. But I definitely got sent a few pitch decks with me, with like my post on there.
Nice.
And it was like really interesting. This is my first time like kind of putting a name to a phenomenon.
Yeah.
And I think that's a useful skill for people to just communicate what they're trying to do.
Yeah, no, I think it's a beautiful concept, yeah.
Yeah, yeah. But I mean, obviously you implemented it. What became more clear in your explanation today is that actually you're not that tied to Python.
No, I mean, I think that all the like lower-level stuff is, you know, just running containers and like scheduling things and, you know, serving container data and stuff. So I mean, I think Python is a great play. Like one of the benefits of data teams is obviously like they're all like using Python,right?
And so that made it a lot easier. I think, you know, if we had focused on other workloads, like, you know, for various reasons, like we've like been kind of like half thinking about like CI or like things like that.
But like, in a way, that's like harder, because like you also, then you have to build like, you know, multiple SDKs. Whereas, you know, focus on data teams, you can only, you know, Python like covers like 95% of all teams.
So that made it a lot easier. But like, I mean, like definitely like in the future, we can add other support, like supporting other languages. JavaScript for sure is the obvious next language. But, you know, who knows, like, you know, Rust, Go, R, like whatever, PHP, Haskell, I don't know.
Yeah, and, you know, I think for me, I actually am a person who like kind of liked the idea of programming language advancements being improvements in developer experience. But all I saw out of the academic sort of PLT type people is just type-level improvements.
And I always think like, for me, like one of the core reasons for Self-Revisioning Runtimes and then why I like Modal is like, this is actually a productivity increase,right?
Totally.
Like it's a language-level thing. You know, you manage to stick it on top of an existing language, but it is your own language.
Yeah.
A DSL on top of Python.
Yeah.
And so language-level increase on the order of like automatic memory management. You know, you could sort of make that analogy that, yeah, like maybe you lose some level of control, but most of the time you're okay with whatever Modal gives you.
And like that's fine.
Yeah. Yeah, I mean, that's how I look at it about it too. Like I think, you know, you look at developer productivity over the last number of decades, like, you know, it's come in like small increments of like, you know, dynamic, like dynamic typing or like it's like one thing.
It's not suddenly like for a lot of use cases, you don't need to care about type systems or better compiler technology or like, you know, or new ways toward, you know, the cloud or like, you know, relational databases.
And, you know, I think, you know, you look at like that, you know, history, it's a steadily, you know, it's like, you know, you look at the developers have been getting like probably 10x more productive every decade for the last four decades or something.
It was kind of crazy, like on an exponential scale, we're talking about 10x or is there a 10,000x like, you know, improvement in developer productivity. What we can build today, you know, is arguably like, you know, a fraction of the cost of what it, you know, took to build it in the '80s.
Maybe it wasn't even possible in the '80s. So to me, like that's like so fascinating. I think it's going to keep going for the next few decades.
Yeah, yeah. Another big thing in the infra 2.0 wishlist was truly serverless infrastructure. The other, on your landing page, you called them native cloud functions, something like that. I think the issue I've seen with serverless has always been people really wanted it to be stateful, even though stateless was much easier to do.
And I think now with AI, most model inference is like stateless, you know, outside of the context. So that's kind of made it a lot easier to just put a model, a model, like an AI model on model to run.
How do you think about how that changes how people think about infrastructure too?
Yeah, I mean, I think Modal is definitely going in the direction of like doing more stateful things and working with data and like high IO use cases. I do think one like massive, like serendipitous thing that happened like halfway, you know, a year and a half into like the, you know, building Modal was like GenAI started exploding.
Serverless AI21:50
And like the IO pattern of GenAI, it's like fits the serverless model like so well, because like it's like, you know, you send this tiny piece of information, like a prompt,right, or something like that. And then like you have this GPU that does like trillions of flops.
And then it sends back like a tiny piece of information,right? And that turns out to be something like, you know, if you can get serverless working with GPU, that just like works really well,right? So I think from that point of view, like serverless always to me felt like a little bit of like a solution looking for a problem.
I don't know, I don't actually like don't think like backend is like the problem that needs serverless, or like not as much. But I look at data and in particular like things like GenAI or like model inference, like it's like clearly a good fit.
So I think that is, you know, to a large extent explains like why we saw, you know, the initial sort of like killer app for Modal being model inference, which actually wasn't like necessarily what we're focused on. But that's where we've seen like by far the most usage and growth.
And was that Stable Diffusion?
Stable Diffusion in particular, yeah.
Yeah, and this was before you started offering like fine-tuning of language models. It was mostly Stable Diffusion.
Yeah, yeah. I mean, like Modal, like I always built it to be a very general-purpose compute platform, like something we could run everything. And I used to call Modal like a better Kubernetes for data team for a long time.
And what we realized was like, yeah, that's like, you know, a year and a half in, like we barely had any users or any revenue. And like we were like, well, maybe we should look at like some use cases, trying to think of use cases.
And that was around the time, same time Stable Diffusion came out. And yeah, like I mean, like the beauty of Modal is like you can run almost anything on Modal,right? Like model inference turned out to be like the place where we found initially, we're like clearly this has like 10x more ergonomic, like better ergonomics than anything else.
But we're also like, you know, going back to my original vision, like we're thinking a lot about, you know, not, okay, now we do inference really well. Like what about training? What about fine-tuning? What about, you know, end-to-end lifecycle deployment?
What about data pre-processing? What about, you know, I don't know, real-time streaming? What about, you know, large data monging? Like there's just data observability. I think there's so many things, like kind of going back to what I said about like redefining the data stack, like starting with the foundation of compute.
Like one of the exciting things about Modal is like we've sort of, you know, we've been working on it now for three years and it's maturing. But like this is so many things you can do, like with just like a better compute primitive and also go up the stack and like do all this other stuff on top of it.
Yeah. How do you think about, or rather like, I would love to learn more about the underlying infrastructure and like how you make that happen. Because with fine-tuning and training, it's a static memory that you're going to, like you exactly know what you're going to load in memory when and it's kind of like a set amount of compute.
Versus inference, just like data is like very bursty. How do you make batches work with a serverless
developer experience? You know, like what are like some fun technical challenges you solve to make sure you get max utilization on these GPUs? What we hear from people is like, we have GPUs, but we can really only get like, you know, 30, 40, 50% maybe utilization.
Yeah.
What's some of the fun stuff you're working on to get a higher number there?
Yeah, I think on the inference side, like that's where like, you know, like from a cost perspective, like utilization perspective, we've seen, you know, like very, very good numbers. And in particular, like it's our ability to start containers and stop containers very quickly.
And that means that we can, you know, we can autoscale extremely fast and scale down very quickly, which means like we can always adjust the sort of capacity, the number of GPUs running to the exact, you know, the traffic volume.
And so in many cases, like that actually leads to this sort of interesting thing where like we obviously run our things on like the public cloud, like AWS, GCP, we run on Oracle. But in many cases, like users who do inference on those platforms or those clouds, even though we charge a slightly higher price per GPU hour, a lot of users like moving their large-scale inference use cases to Modal, like end up saving a lot of money.
Because we only charge for like the time the GPU is actually running. And that's a hard problem,right? Like if you go, you know, if you have to constantly adjust the number of machines, if you have to start containers, stop containers, like that's a very hard problem.
And that, you know, and starting containers quickly is a very difficult thing. I mentioned we had to build our own file system for this. We also, you know, built our own container scheduler for that. We're looking, we've implemented recently CPU memory checkpointing so we can take running containers and snapshot the entire CPU, like including registers and everything, and restore it from that point, which means we can restore it from like an initialized state.
We're looking at GPU checkpointing next as like a very interesting thing. So I think on the inference stuff, on the inference side, like that's where serverless really shines, because you can drive, you know, you can push the frontier of latency versus utilization quite substantially, you know, which either ends up being a latency advantage or a cost advantage or both,right?
On training, it's probably arguably like less of an advantage doing serverless, frankly, because, you know, you can just like spin up a bunch of machines and try to satisfy like, you know, train as much as you can on each machine.
For that area, like we've seen like, you know, arguably like less usage like for Modal. But there are like some interesting use cases. Like we do have a couple of customers, like RAM, for instance, like they do fine-tuning with Modal and they basically like one of the patterns they have is like very bursty type fine-tuning where they fine-tune 100 models in parallel.
And that's like a separate thing that Modal does really well,right? Like we can start up 100 containers very quickly, run a fine-tuning training job on each one of them for that only runs for, I don't know, 10, 20 minutes.
And then, you know, you can do hyperparameter tuning in that sense, like just pick the best model and things like that. So there are like interesting training. Like I think when you get to like training like very large foundational models, like that's a use case we don't support super well, because that's very high IO, you know, you need to have like infinite band and all these things.
And those are things we haven't supported yet and might take a while to get to that. So that's like probably like an area where like we're relatively weak in.
Yeah. Have you cared at all about lower-level model optimization? There's other cloud providers that do custom kernels to get better performance, or are you just, given that you're not just an AI compute company?
Yeah, I mean, I think like we want to support like a generic, like general workloads in a sense that like we want users to give us a container essentially or a code and then we want to run that.
So I think, you know, we benefit from those things in the sense that like we, you know, we can tell our users, you know, to use those things. But I don't know if we want to like poke into users' containers and like do those things automatically.
That's sort of, I think, a little bit tricky from the outside to do, because, you know, we want to be able to take like arbitrary code and execute it. But certainly, like, you know, we can tell our users to like use those things.
Yeah. I may have betrayed my own biases because I don't really think about Modal as for data teams anymore. I think you started that way. I think you're much more for AI engineers. And, you know, one of my favorite anecdotes, which I think you know, but I don't know if you directly experienced it.
I went through the Vercel AI accelerator, which you supported.
Yep.
And in the Vercel AI accelerator, a bunch of startups gave like free credits and like sign-ups and talks and all that stuff. The only ones that stuck are the ones that actually appealed to engineers and the top usage, the top tool used by far was Modal.
Hmm, that's awesome.
For people building with AI apps.
Yeah, I mean, it might be also like a terminology question, like the AI versus data,right? Like I've, you know, maybe I'm just like old and jaded, but like I've seen so many like different titles. Like for a while it was like, you know, I was a data scientist and I was a machine learning engineer and then, you know, there was like analytics engineers and then it was like AI engineers.
You know, so like to me it's like, I just like in my head, that's to me just like just data, like or like engineer, you know, like I don't really. So that's why I've been like, you know, just calling it data teams.
But like, of course, like, you know, AI is like, you know, like such a massive fraction of our like workloads.
It's a different Venn diagram of things you do,right? So the stuff that you're talking about where you need like infinite bands for like highly parallel training, that's not, that's more of the ML engineer, that's more of the research scientist.
Yeah.
And less of the AI engineer, which is more sort of trying to put work at the application.
Yeah. I mean, to be fair, like we have a lot of users that are like doing stuff that I don't think fits neatly into like AI. Like we have a lot of people using like Modal for web scraping.
Like it's kind of nice, like you can just like, you know, fire up like 100 or 1000 containers running Chromium and just like render a bunch of web pages and take, you know, whatever. Or like, you know, protein folding, is that, I mean, maybe that's, I don't know, like, but like, you know, we have a bunch of users doing that or like, you know, in terms of in the realm of biotech, like sequence alignment, like people using, or like a couple of people using like Modal to run like large, like mixed integer programming problems, like, you know, using Garuba, like things like that.
So video processing is another thing that keeps coming up. Like, you know, let's say you have like petabytes of video and you want to like transcode it, like, or you can fire up a lot of containers and just run FFmpeg.
So there are those things too. Like, I mean, like that being said, like AI is by far our biggest use case, but, you know, like again, like Modal is kind of general purpose in that sense.
Yeah. Well, maybe, so I'll stick to the Stable Diffusion thing and then we'll move on to the other use cases of sort of for AI that you want to highlight. The other big player in my mind is Replicate.
Competition31:10
Yeah.
In this era. They're much more, I guess, custom built for that purpose, whereas you're more general purpose. How do you position yourself with them? Are they just for like different audiences or are you just hit on competing?
I think there's like a tiny sliver of the Venn diagram where we're competitive and then like 99% of the area we're not competitive. I mean, I think for people who, in particular like front engineers, I think that's where like really the found good fit is like, you know, people who built some cool web app and they want some sort of AI capability and they just, you know, an off-the-shelf model is like perfect for them.
That's like, like use Replicate, that's great,right? Like I think where we shine is like custom models or custom workflows, you know, running things at very large scale, we need to care about utilization, care about cost, you know, we have much lower prices because we spend a lot more time optimizing our infrastructure.
And, you know, and that's where we're competitive,right? Like, you know, and you look at some of the use cases like Suno is a big user. Like they're running like large scale like AI.
Oh, we're talking with Mikey.
Oh, that's great. Cool.
In a month.
Yeah. So I mean, they're using Modal for like production infrastructure. Like they have their own like custom model, like custom code and custom weights, you know, for AI generated music, Suno.ai. You know, those are the types of use cases that we like, you know, things that are like very custom or like it's like, you know, and those are the things like it's very hard to run in Replicate,right?
And that's fine. Like I think they focus on a very different part of the stack in that sense.
And then the other
company pattern that I pattern match you to is Modular.
Is it the names?
No, well, no, but yes, the name is very similar. I think there's something that might be insightful there from a linguistics point of view. But no, they have Mojo, the sort of Python SDK.
Yeah.
And then they have the Modular Inference Engine, which is their sort of cloud stack, their sort of compute inference stack. I don't know if anyone's made the comparison to you before, but like I'm seeing you evolving a little bit in parallel there.
No, I mean, maybe. Yeah, like it's not a company I'm like super like familiar with. Like, I mean, I know the basics, but like I guess they're similar in the sense like they want to like do a lot of, you know, they have sort of big picture.
Yes, they also want to build very general purpose.
Yeah.
And they also are.
Which I admire.
Marketing themselves as like, if you want to do off-the-shelf stuff, go somewhere else. If you want to do custom stuff, we're the best place to do it.
Yeah.
Yeah. There is some overlap there. There is not overlap in the sense that
you are a closed source platform. People have to host their code on you.
That's true.
Whereas for them, they're very insistent on not running their own cloud service.
Yeah.
They're a boxed software.
Yeah.
They're licensed software.
I'm sure their VCs at some point are going to force them to reconsider.
No, no, Chris is very, very insistent and very convincing. So anyway, I'll just make that comparison, let people make the links if they want to, but it's an interesting way to see the cloud market develop from my point of view, because I came up in this field thinking cloud is one thing, and I think your vision is like something slightly different, and I see the different takes on it.
Yeah, and like one thing I've, you know, like I've written a bit about it in my blog too, is like I think of us as like a second layer of cloud provider in the sense that like I think Snowflake is like kind of a good analogy.
Like Snowflake, you know, is infrastructure as a service,right? But they actually run on the like major clouds,right? And I mean, like you can like analyze this very deeply, but like one of the things I always thought about is like why did Snowflake arguably like win over Redshift?
And I think Snowflake, you know, to me, one, because like, I mean, in the end, like AWS makes all the money anyway. Like, and like Snowflake just had the ability to like focus on like developer experience or like, you know, or user experience.
And to me, like really prove that you can build a cloud provider a layer up from, you know, the traditional like public clouds. And in that layer, that's also where I would put Modal. It's like, you know, we're building a cloud provider.
Like, you know, we're like a multi-tenant environment that runs the user code, but also building on top of the public cloud. So I think there's a lot of room in that space. I think it's a very sort of interesting direction.
Yeah. How do you think of that compared to the traditional past history? Like, you know, yeah, AWS, then you had a Heroku, then you had Render, Railway.
Yeah, I mean, I think those are all like great. Like, I think the problem that they all faced was like the graduation problem,right? Like, you know, Heroku or like, I mean, like also like Heroku, there's like a counterfactual future of like what would have happened if Salesforce didn't buy them,right?
Like, that's a sort of separate thing. But like, I think what Heroku, I think, always struggled with was like eventually companies would get big enough that you couldn't really justify running in Heroku. So they would just go and like move it to, you know, whatever AWS or, you know, in particular.
And, you know, that's something that keeps me up at night too. Like, what does that graduation risk like look like for Modal? I always think like the only way to do, to build infrastructure, to build a successful infrastructure company in the long run in the cloud today is you have to appeal to the entire spectrum,right?
Or at least like the enterprise. Like, you have to capture the enterprise market. And, but the truly good companies capture the whole spectrum,right? Like, I think of companies like, I don't know, like Datadog or Mongo or something like that, where like they both capture like the hobbyists and
acquire them, but also like, you know, have very large enterprise customers. So I think that arguably was like where I, in my opinion, like Heroku struggled was like how do you maintain the customers as they get more and more advanced?
I don't know what the solution is, but I think there's, you know, that's something I would have thought deeply if I was at Heroku at that time.
What's the AI graduation problem? Is it, I need to fine-tune the model, I need better economics, any insights from customer discussions?
Yeah, I mean, better economics certainly, but although like I would say like even for people who like, you know, needs like thousands of GPUs at very, you know, like just because we can drive utilization so much better, like there's actually like a cost advantage of staying on Modal.
But yeah, I mean, certainly like, you know, and then like the fact that VCs like love, you know, throwing money, at least used to, you know, at companies who needed to buy GPUs, I think that didn't help the problem.
Yeah, and in training, I think, you know, there's less software differentiation. So in training, I think there's certainly like better economics of like buying big clusters.
But I mean, my hope it's going to change,right? Like I think, you know, we're still pretty early in the cycle of like building AI infrastructure. And, you know, I think a lot of these companies over in the long run, like, you know, they're except for maybe super big ones like, you know, the Facebook and Google, they're always going to build their own ones.
But like everyone else, like to some extent, you know, I think they're better off like buying platforms. And, you know, someone's going to have to build those platforms.
Yeah. Cool. Let's move on to language models and just specifically that workload, just to flesh it out a little bit. You already said that RAMP is like fine-tuning 100 models at once simultaneously on Modal. Closer to home, my favorite example is ErikBot.
AI Features38:03
Maybe you want to tell that story.
Yeah, I mean, it was a prototype thing we built for fun, but it was pretty cool. Like we basically built this thing that you can, it like hooks up to Slack, it like downloads all the Slack history and, you know, fine-tunes a model based on a person and then you can chat with that.
And so you can like, you know, clone yourself and like talk to yourself. I mean, it's like a nice like demo in a sense like I think like it's like fully contained Modal. Like there's a Modal app that does everything,right?
Like it downloads Slack, you know, it integrates with the Slack API, like downloads the stuff, the data, like just runs the fine-tuning and then like creates like dynamically an inference endpoint. And it's all like self-contained and like, you know, a few hundred lines of code.
So I think it's sort of a good kind of use case for Modal or like it kind of demonstrates a lot of the capabilities of Modal.
Yeah. On a more personal side, how close did you feel ErikBot was to you?
It definitely captured like the language. Yeah, I mean, I don't know, like the content. I mean, like it's like I always feel this way about like AI. And it's gotten better, but like you look at like AI output of text, like, and it's like when you glance at it, it's like, yeah, this seems really smart, you know?
But then you actually like look a little bit deeper. It's like, what does this mean? What does this person say? It's like kind of vacuous,right? And that's like kind of what I felt like, you know, talking to like my cloned version.
Like it like says like things like the grammar is correct. Like some of the sentences make a lot of sense, but like what are you trying to say? Like there's no content here. I don't know. I mean, it's like I got that feeling also with ChatGPT in the like early versions,right?
Now it's like better, but.
That's funny. Yeah, I built this thing called Small Podcaster to automate a lot of our back office work, so to speak. And it's great at transcript, it's great at doing chapters. And then I was like, okay, how about you come up with a short summary?
And it's like, it sounds good, but it's like, it's not even the same ballpark as like what we end up writing.
Right.
And it's hard to see how it's going to get there.
Oh, I have ideas. I'm working on it.
Yeah.
I'm certain it's going to get there, but like I agree with you,right? And like I have the same thing. I don't know if you've read like AI generated books. Like they just like kind of seem funny,right? Like they're just off,right?
But like you just glance at them and it's like, oh, it's kind of cool. Like looks correct, but then it's like very weird when you actually read them.
Well, so for what it's worth, I think anyone can join the Modal Slack. Is it open to the public?
Yeah, totally. If you go to modal.com, there's a button in the footer.
Yeah, and then you can talk to ErikBot. And then sometimes I really like pigging ErikBot and then you answer afterwards, but then you're like.
Really?
Yeah, mostly correct or like whatever.
Cool.
No, so okay, any other broader lessons, you know, just broadening out from like the single use case of fine-tuning, like what are you seeing people do with fine-tuning or just language models on Modal in general?
Yeah, I mean, I think language models is interesting because so many people get started with APIs and that's just, you know, they're just dominating a space in particular, OpenAI,right? And that's not necessarily like a place where we aim to compete.
I mean, maybe at some point, but like it's just not like a core focus for us. And I think sort of separately it's sort of a question if like there's economics in that long term, but like so we tend to focus on more like the areas like around it,right?
Like fine-tuning, like another use case we have is a bunch of people, RAMP included is doing batch embeddings on Modal. So let's say, you know, you have like a, actually we're like writing a blog post like where we take all of Wikipedia and like paralyze embeddings in 15 minutes and produce vectors for each article.
So those types of use cases, I think Modal suits really well for. I think also a lot of like custom inference, like we have like, you know, structured output, guided generation or things like that. We have, if you want more control, like those are the things like we see a lot of users using Modal for.
But for a lot of people, it's like, you know, just go use like GPT-4 and like, you know, that's like a great starting point and we're not trying to compete necessarily like directly with that.
Yeah. When you say parallelize, I think you should give people an idea of the order of magnitude of parallelism because I think people don't understand how parallel. So like I think your classic hello world with Modal is like some kind of Fibonacci function,right?
Yeah, we have a bunch of different.
Like a percent function.
Yeah, yeah. I mean, like, yeah, I mean, it's like pretty easy in Modal like fan out to like, you know, at least like 100 GPUs like in a few seconds. And, you know, if you give it like a couple of minutes, like we can, you know, you can fan out to like thousands of GPUs.
Like we run at relatively large scale. And, yeah, we've run, you know, many thousands of GPUs at certain points when we needed, you know, big backfills or some customers had very large compute needs.
Yeah, yeah. And I mean, that's super useful for a number of things. One of the reasons actually I, so one of my early interactions with Modal as well was with Small Developer, which is my sort of coding agent.
The reason I chose Modal was a number of things. One, I just wanted to try it out. I just had an excuse to try it. Akshad offered to onboard me personally.
Yeah, good excuse.
But the most interesting thing was that you could have that sort of local development experience as I was running it on my laptop, but then it would seamlessly translate to a cloud service or like a cloud-hosted environment. And then it could fan out with concurrency controls.
So I could say like, because like, you know, the number of times I hit the GPT-3 API at the time was going to be subject to the rate limit from there.
Yeah.
But I wanted to fan out without worrying about that kind of stuff.
Yeah.
With Modal, I can just kind of declare that in my config and that's it.
Oh, like a concurrency limit?
Yeah.
Yeah, yeah, yeah. There's a lot of control. Yeah, yeah.
Yeah, yeah. So like I just wanted to highlight that to people as like, yeah, this is a pretty good use case for like, you know, just like writing this kind of LLM application code inside of this environment that just understands fan out and rate limiting natively.
You don't actually have an exposed queue system, but you have it under the hood, you know, that kind of stuff.
Totally. It's a self-provisioning cloud.
So the last part of Modal I wanted to touch on and obviously feel free, I know you're working on new features, was the sandbox that was introduced last year. And this is something that I think was inspired by Code Interpreter.
You can tell me the longer history behind that.
Yeah, like we originally built it for the use case. Like there was a bunch of customers who looked into code generation applications and then they wanted, they came to us and asked, is there a safe way to execute code?
And yeah, we spent a lot of time on like container security. We used JVisor, for instance, which is a Google product that provides pretty strong isolation of code. So we built a product where you can like basically like run arbitrary code inside a container and monitor its output or like get it back in a safe way.
I mean, over time it's like evolved into more of like, I think the long-term direction is actually, I think, more interesting, which is that I think Modal as a platform where like I think the core like container infrastructure we offer could actually be like, you know, unbundled from like the client SDK and offered to like other, you know, like we're talking to a couple of like
other companies that want to run, you know, through their packages, like run, execute jobs on Modal like kind of programmatically. And so that's actually the direction like sandboxes are going. It's like turning into more like a platform for platforms is kind of what I've been thinking about it.
Oh boy, platform. That's an old Kubernetes line.
Yeah, yeah, yeah. But it's like, you know, like having that ability to like programmatically, you know, create containers and execute them, I think is really cool. And I think it opens up a lot of interesting capabilities that are sort of separate from the like core Python SDK in Modal.
So I'm really excited about seeing. I mean, it's like one of those features that we kind of released and like, you know, then we kind of look at like what users actually build with it. And people are starting to build like kind of crazy things.
And then, you know, we double down on some of those things because when we see like, you know, potential new product features. And so sandboxes, I think in that sense, it's like kind of in that direction. We found a lot of like interesting use cases in the direction of like platformized container runner.
Can you be more specific about what you're double down on after seeing users in action?
Yeah, I mean, we're working with like some companies that, I mean, without getting into specifics, like that need the ability to take their users' code and then launch containers on Modal. And it's not about security necessarily. Like they just want to use Modal as a backend,right?
Like they may already provide like Kubernetes as a backend, Lambda as a backend, and now they want to add Modal as a backend,right? And so, you know, they need a way to programmatically define jobs on behalf of their users and execute them.
And so I don't know, that's kind of abstract, but does that make sense?
I totally get it. It's sort of one level of recursion to sort of be the Modal for their customers.
Exactly. Yeah, yeah, exactly.
And Cloudflare has done this. You know, Kenton Vardar from Cloudflare was like the tech lead on this thing, called it sort of functions as a service as a service.
Yeah. That was exactlyright. Facets.
Facets.
Facets.
Yeah, like, I mean, like that, I think any
base layer, second layer cloud provider like yourself, compute provider like yourself should provide. It is like very, very, you know, it's a marker of maturity and success that people just trust you to do that. They'd rather build on top of you than compete with you.
Like the more interesting thing for me is like what does it mean to serve a computer, like an LLM developer rather than a human developer,right? Like that's what a sandbox is to me.
Yeah, for sure.
That you have to redefine Modal to serve a different non-human audience.
Yeah, yeah, yeah. And I think there's some really interesting people, you know, building very cool things.
Yeah, so I don't have an answer, but, you know, I imagine things like, hey, the way you give feedback is different. You maybe have to like stream errors, log errors differently. I don't really know.
Yeah.
Obviously, there's like safety considerations. Maybe you have an API to like restrict access to the web.
Yeah.
I don't think anyone would use it, but it's there if you want it.
Yeah, yeah.
Any other sort of design considerations? I have no idea.
With sandboxes?
Yeah, yeah. Open-ended question here.
Yeah, I mean, no, I think, yeah, the network restrictions I think make a lot of sense. Yeah, I mean, I think, you know, long-term, like I think there's a lot of interesting use cases where like the LLM in itself can like decide I want to install these packages and like run this thing.
And like obviously for a lot of those use cases, like you want to have some sort of control that it doesn't like install malicious stuff and steal your secrets and things like that. But I think that's what's exciting about the sandbox primitive is like it lets you do that in a relatively safe way.
Yeah.
Cool.
Do you have any thoughts on the inference wars? So a lot of providers are just rushing to the bottom to get the lowest price per million tokens. Some of them, you know, the Sean Rendemat, they're just losing money.
Inference Wars48:56
Like the physics of it just don't work out for them to make any money on it. How do you think about your pricing and like how much premium you can get and you can kind of command versus using lower prices as kind of like a wedge into getting there, especially once you have Modal instrumented?
Yeah, what are the trade-offs and any thoughts on strategies that work?
Yeah. I mean, we focus more on like custom models and custom code. And I think in that space, there's like less competition. And I think we can, you know, have a pricing markup,right? Like, you know, people will always compare our prices to like, you know, the GPU power they can get elsewhere.
And so how big can that markup be? Like it never can be, you know, we can never charge like 10x more, but we can certainly charge a premium. And like, you know, for that reason, like we can have pretty good margins.
The LLM space is like the opposite. Like the switching cost of LLMs is zero. Like if all you're doing is like straight up, like at least like open source,right? Like if all you're doing is like, you know, using some, you know, inference endpoint that serves some open source model and, you know, some other provider comes along and like offers a lower price, you're just going to switch,right?
So I don't know. To me, that reminds me a lot of like all this like 15-minute delivery war. So like, you know, like Uber versus Lyft or DraftKings versus Fanny. Like or like maybe that's not, but like, you know, and like maybe going back even further, like I think a lot about like the sort of, you know, flip side of this is like there's actually a positive side of it.
It's like, like I think I thought a lot about like fiber optics boom of like '98, '99, like the other day or like, you know, and also like the overinvestment in GPU today. Like, yeah, like, you know, I don't know.
Like in the end, like I don't think VCs will have the return they expected, like, you know, in these things. But guess who's going to benefit? Like, you know, it's the consumers,right? Like someone's like reaping the value of this.
And that's, I think, an amazing flip side is that, you know, we should be very grateful, you know, the fact that like VCs want to subsidize these things, which is, you know, like you go back to fiber optics.
Like there was the extreme like overinvestment in fiber optics network in the '99 and like '98. And no one made money who did that. But consumers, you know, got tremendous benefits of all the fiber optics cables that were led, you know, throughout the country in the decades after.
I feel something similar about like GPUs today and also like specifically looking like more narrowly at like the LLM and inference market. Like that's great. Like, you know, I'm very happy that, you know, there's a price war. Modal is like not necessarily like participating in that price war,right?
Like I think, you know, it's going to shake out and then someone's going to win and then they're going to raise prices or whatever. Like we'll see how that works out. But it's not, for that reason, like we're not like focused, like we're not hyper-focused on like serving, you know, just like straight up, like here's an endpoint to an open source model.
We think the value in Modal comes from all these, you know, the other use cases, the more custom stuff, like fine-tuning and very complex, you know, guided output, like type stuff or like also like in other, like outside of LLMs, like we focus a lot more on like image, audio, video stuff, because that's where there's a lot more proprietary models.
There's a lot more like custom workflows. And that's where I think, you know, Modal is more, you know, there's a lot of value in software differentiation. I think focusing on developer experience and developer productivity, that's where I think, you know, you can have more of a competitive moat.
Yeah. I'm curious what the difference is going to be now that it's an enterprise. So like with DoorDash, Uber, they're going to charge you more and like as a customer, like you can decide to not take Uber. But if you're a company building AI features in your product using the subsidized prices and then, you know, the VC money dries up in a year and like prices go up, it's like you can't really take the features back without a lot of backlash, but you also cannot really heal your margins by paying the new price.
So I don't know what that's going to look like, but.
But like margins are going to go up for sure, but I don't know if prices will go up because like GPU prices have to drop eventually,right? So like, you know, like in the long run, I still think like prices may not go up that much, but certainly margins will go up.
Like I think you said SwiftStack margins are negativeright now. Like, you know.
Oh, for some people.
Obviously, that's not sustainable. So certainly margins will have to go up. Like some companies are going to have to make money in this space. Otherwise, like they're not going to provide the service. But that's the equilibrium too,right? Like at some point, like, you know, they sort of stabilize this and one or two or three providers make money.
Yeah, what else is a maybe underrated model, something that people don't talk enough about or, yeah, that we didn't cover in the discussion?
Yeah, I think what are some other things? We talked about a lot of stuff. Like we have the bursty parallelism. I think that's pretty cool. Working on a lot of like trying to figure out like, I'm kind of thinking more about the roadmap, but like one of the things I'm very excited about is building primitives for like more like I/O intensive workloads.
And so like we're building some like crude stuffright now where like you can like create like direct TCP tunnels to containers and that lets you like pipe data and like, you know, we haven't really explored this as much as we should, but like there's a lot of interesting applications.
Like you can actually do like kind of real-time video stuff in Modal now because you can like create a tunnel to, yeah, exactly. You can create a raw TCP socket to a container, feed it video, and then like, you know, get the video back.
And I think like it's still like a little bit like, you know, not fully ergonomically like figured out, but I think there's a lot of like super cool stuff like when we start enabling those more like high I/O workloads, I'm super excited about.
I think also like, you know, working with large data sets or kind of taking the ability to map and fan out and like building more like higher level like functional primitives like filters and group buyers and joins. Like I think there's a lot of like really cool stuff you can do, but this is like maybe like, you know, years out.
Yeah. We can just broaden it out from Modal a little bit, but you still have a lot of, you have a lot of great tweets. So it's very easy to just kind of go through them. Why is Oracle underrated?
Hardware & Tweets54:53
I love Oracle's GPUs. I mean, like I don't know why, you know, what the economics looks like for Oracle, but like I think they're great value for money. Like we run a bunch of stuff in Oracle and they have bare metal machines for like two terabytes of RAM.
They're like super fast SSDs. And yeah, like compared to, you know, I mean, we love AWS and AGCP too. We have great relationships with them. But I think Oracle's surprising, like, you know, if you told me like three years ago that I would be using Oracle Cloud, like I'd be like, what?
Wait, why? But now I'm, you know, I'm a happy customer.
And it's a combination of pricing and the kinds of SKUs, I guess, they offer.
Yeah. Great, great machines, good prices, you know.
That's it.
Yeah. Yeah. That's all I care about. Yeah. The sales team is pretty fun too. Like I like them.
In Europe, people often talk about Hetzner.
Yeah. I'm not, I don't like super, like we've focused on the main clouds,right? Like we've, you know, Oracle, AWS, GCP will probably add Azure at some point. I think, I mean, there's definitely a long tail of like, you know, CoreWeave, Hetzner, like
Lambda, like all these things. And like over time, I think we'll look at those too, like, you know, wherever we can get theright, you know, GPUs at theright price. Yeah, I mean, I think it's fascinating. Like it's a tough business.
Like I wouldn't want to try to build like a cloud provider, you know, it's just, you just have to be like incredibly focused on like, you know, efficiency and margins and things like that. But I mean, yeah, I'm glad people are trying.
Yeah. And you can ramp up on any of these clouds very quickly,right? Because it's.
Yeah, I mean, yeah, like I think so. Like we, like a lot of, you know, what Modal does is like programmatic, you know, launching and termination of machines. So that's like what's nice about the clouds is, you know, they're relatively like mature APIs for doing that as well as like, you know, support for Terraform for all the networking and all that stuff.
So that makes it easier to work with the big clouds. But yeah, I mean, some of those things, like I think, you know, I also expect the smaller clouds to like embrace those things in the long run, but also think, you know, in the, you know, we can also probably integrate with some of the clouds like even without that.
There's always an HTML API that you can use, just like script something that launches instances like through the web. Yeah, endpoints.
I think a lot of people are always curious about whether or not you will buy your own hardware someday. I think you're pretty firm in that it's not in your interest. But like your story and your growth does remind me a little bit of Cloudflare, which obviously, you know, invests a lot in its own physical network.
Yeah. I don't remember like early days, like did they have their own hardware or?
I think they also have.
They bush out a lot with like agreements through other.
Acronyms or one of those things.
Yeah, okay, interesting.
But now it's all their own hardware.
Yeah.
So I understand.
Yeah, I mean, my feeling is that when you're a venture-funded startup, like buying physical hardware is maybe not the best use of the money.
I really wanted to put you in a room with ISO Cat from Poolside.
Yeah.
Because he has the complete opposite view.
Yeah.
This is great.
I mean, I don't know. I just think from like a capital efficiency point of view, like do you really want to tie up that much money in like, you know, physical hardware and think about depreciation and like, like as much as possible, like I, you know, I favor a more capital efficient way of like we don't want to own the hardware because then, and ideally we want to, we want the sort of margin structure to be sort of like 100% correlated revenue and COGS in the sense that like, you know, when someone comes and pays us, you know, $1 for compute, like, you know, we immediately incur a cost of like whatever, 70 cents, 80 cents, you know, and there's like complete correlation between cost and revenue because then you can leverage up in like
a kind of a nice way. You can scale very efficiently. You know, like that's not, you know, turns out like that's hard to do. Like you can't just only use like spotting on-demand instances. Like over time, we've actually started adding a pretty significant amount of reservations too.
So I don't know. Like reservations is always like one step towards owning your own hardware. Like I don't know. Like do we really want to be, you know, thinking about switches and cooling and HVAC and like power supplies and like.
Disaster recovery.
Yeah. Like is that the thing I want to think about? Like I don't know. Like I like to make developers happy, but who knows? Like maybe one day. Like, but I don't think it's going to happen anytime soon.
Yeah. Obviously, for what it's worth, obviously I'm a believer in cloud, but it's interesting to have the devil's advocate on the other side. The main thing you have to do is be confident that you can manage your depreciation better than the typical assumption, which is two to three years.
Yeah. Yeah.
And so the moment he, the moment you have a CTO that tells you, no, I think I can make these things last seven years, then it changes the math.
Yeah. Yeah. But, you know, are you deluding yourself then? That's the question,right?
Yeah, yeah.
It's like the waste management scandal. Do you know about that? Like they had all this like accounting scandal back in the '90s. Like this garbage company, like where they like started assuming their garbage trucks had a 10-year depreciation schedule, booked like a massive profit.
You know, the stock went to like, you know, up like, you know, and then it turns out actually all those garbage trucks broke down and like you can't really depreciate them over 10 years. And so then the whole company, you know, they had to restate all the earnings and they just.
Nice. Let's go into some personal nuggets. You received the IOI Gold Medal, which is the International Olympiad in Informatics.
Personal1:00:20
20 years ago.
Yeah. How have these models and like going to change competitive programming? Like do you think people still love the craft? I feel like over time we're kind of like programming has kind of lost maybe a little bit of its luster in the eyes of a lot of people.
Yeah, I'm curious to see what you think.
I mean, maybe, but like, I don't know. Like, you know, I've been coding for almost 30 or more than 30 years. And like I feel like, you know, you look at like programming and, you know, where it is today versus where it was, you know, 30, 40, 50 years ago, there's like probably a thousand times more developers today than, you know, so like, and every year there's more and more developers.
And at the same time, developer productivity keeps going up. And so like I actually don't expect, like, and when I look at the real world, I just think there's so much software that's still waiting to be built. Like I think we can, you know, 10X the amount of developers and still, you know, have a lot of people making a lot of money, you know, building amazing software and also being, while at the same time being more productive.
Like I never understood this like, you know, AI is going to, you know, replace engineers. That's very rarely how this actually works. When AI makes engineers more productive, like the demand actually goes up because the cost of engineers goes down because you can build software more cheaply.
And that's, I think, the story of software in the world over the last few decades. So I mean, I don't know how this like relates to like competitive programming as a, like I don't know, kind of going back to your question.
Competitive programming to me was always kind of a weird kind of, you know, niche, like kind of, I don't know, I loved it. It's like puzzle solving. And like my experience is like, you know, half of competitive programmers are able to translate that to actual like building cool stuff in the world.
Half just like get really in, you know, sucked into this like puzzle stuff and, you know, it never loses its grip on them. But like for me, it was an amazing way to get started with coding and or get very deep into coding and, you know, kind of battle off with like other smart kids and traveling to different countries when I was a teenager.
Yeah. There was another. Oh, sorry. I was just going to mention like it's not just that he personally is a competitive programmer. Like I think a lot of people at Modal are competitive programmers. I think you met Akshat through.
Akshat, co-founder is also IOI Gold Medal. By the way, Gold Medal doesn't mean you win. Like, but although we actually had an intern that won IOI. Gold Medal is like the top 20, 30 people roughly.
Yeah. And so like obviously it's very hard to get hired at Modal, but like what is it like to work with like such a talent density? Like, you know, how is that contributing to the culture at Modal?
Yeah, I mean, I think humans are the root cause of like everything at a company,right? Like, you know, bad code is because it's bad human or like whatever, you know, bad culture. So like I think, you know, like talent density is very important and like keeping the bar high and like hiring smart people and, you know, it's not always like the case that like hiring competitive programmers is theright strategy,right?
If you're building something very different, like you may not, you know, but we actually end up having a lot of like hard, you know, complex challenges. Like, you know, I talked about like the cloud, you know, the resource allocation.
Like turns out like that actually, like you can phrase that as a mixed-engineered programming problem. Like we now have that running in production, like constantly optimizing how we allocate cloud resources. There's a lot of like interesting, like complex, like scheduling problems and like how do you do all the bin packing of all the containers?
Like, so I, you know, I think for, you know, for what we're building, you know, it makes a lot of sense to hire these people who like those very hard problems.
Yeah. And they don't necessarily have to know the details of the stack. They just need to be very good at algorithms.
No, but my feeling is like people who are like pretty good at competitive programming, they can also pick up like other stuff like elsewhere. Not always the case, but, you know, there's definitely a high correlation.
Yeah. Oh yeah. I'm just, I'm interested in that just because, you know, like there's competitive mental talents in other areas, like competitive speed memorization or whatever. And like they don't really see those transfer. And I always assumed in my narrow perception that competitive programming is so specialized, it's so obscure, even like so divorced from real-world scenarios that it doesn't actually transfer that much.
But obviously, I think for the problems that you work on, it does.
But it's also like, you know, frankly, it like translates to some extent, not because like the problems are the same, but just because like it sort of filters for the, you know, people who are like willing to go very deep and work hard on things,right?
Like I feel like a similar thing is like a lot of good developers are like talented musicians. Like why? Like why is there a correlation? And like my theory is like, you know, it's the same sort of skill.
Like you have to like just hyper-focus on something and practice a lot. Like, and there's something similar that I think creates like good developers.
Sweden also had a lot of very good Counter Strike players. I don't know. Why did Sweden have fiber optics before all of Europe? I feel like I grew up in Italy and our internet was terrible. And then I feel like all the Nordics had like amazing internet.
I remember getting online and people in the Nordics had like 5 ping, 10 ping.
Yeah, we had very good network back then. Yeah.
Yeah. Do you know why?
I mean, I'm sure like, you know, I think the government, you know, did certain things quite well,right? Like in the '90s, like there was like a bunch of tax rebates for like buying computers. And I think there were similar like investments in infrastructure that, you know, I mean, like and I think like I always think about, you know, it's like I still can't use my phone in the subway in New York.
And that's, you know, and that was something I could use in Sweden in '95. You know, we're talking like 40 years almost,right? Like, like why? And I don't know. Like I think certain infrastructure, you know, Sweden was just better at it.
I don't know.
Yeah. And also, you never owned a TV or a car?
Never owned a TV or a car. I never had a driver's license.
How do you do that in Sweden though? Like that's cold.
I grew up in a city. I mean, like I took the subway everywhere with a bike or whatever. Yeah. I always lived in cities, so I don't, you know, I never felt. I mean, like we have a, like me and my wife has a car, but like I.
That doesn't count.
I mean, it's in her name because I don't have a driver's license. She drives me everywhere. It's nice.
Nice.
That's fantastic.
Great. I know. Any, I was going to ask you, like the last thing I had on this list was, you know, your advice to people thinking about running some sort of run code in a cloud startup is only do it if you're genuinely excited about spending five years thinking about load balancing, page faults, Cloud Security, and DNS.
Startup Advice1:06:47
So basically, like it sounds like you're summing up a lot of pain running Modal.
Yeah.
Yeah.
I mean, like that's like, like one thing I struggle with, like I talked to a lot of people starting companies in the data space or like AI space or whatever, and they could sort of come at it at like, as you know, from like an application developer point of view, and they're like, I'm going to make this better.
But like guess how you have to make it better? It's like you have to go very deep on the infrastructure layer. And so, and so one of my frustrations has been like so many startups are like, in my opinion, like Kubernetes wrappers and like, you know, like and not very like thick wrappers, like fairly thin wrappers.
And I think, you know, every startup is a wrapper to some extent, but like you need to be like a fat wrapper. You need to like go deep and like build some stuff. And that's like, you know, if you build a tech company, you're going to want to, you're going to have to spend, you know, 5, 10, 20 years of your life like going very deep and like, you know, building the infrastructure you need in order to like make your product truly stand out and be competitive.
And so, you know, I think that goes for everything. I mean, like you're starting a whatever, you know, online retailer of, I don't know, bathroom sinks. You probably have to spend, you know, 10, you have to be willing to spend 10 years of your life thinking about, you know, whatever bathroom sinks.
Like otherwise, it's going to be hard.
Yeah. Yeah. Makes sense. I think that's good advice for everyone. And yeah, congrats on all your success. It's pretty exciting to watch and it's just the beginning.
Yeah. Yeah. Yeah. It's exciting. And everyone should sign up and try out Modal.
Yeah. Now it's GA. Yay.
Yeah.
Used to be behind a waitlist.
Yeah.
Awesome, Erik. Thank you so much for coming on.
Yeah, it was amazing. Thank you so much.
Thanks.






