Intro0:00
Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, partner and CTO at Decibel, and I'm joined by my co-host Swyx, founder of Small AI.
And we're so excited to be back in the studio with Chris Lattner from Mojo Modular. Welcome back.
Yeah, I'm super excited to be here. As you know, I'm a huge fan of both of you-
Yeah
... and also the podcast, and a lot, a lot of what's going on in the industry.
Yeah.
So thank you for having me.
Thanks for keeping us involved in your journey. And I think, like, you spend a lot of time writing when obviously your, your time is super valuable, educating the rest of the world. Like, I just saw your, like, a two and a half hour workshop with the GPU Mode guys, and that, that was super exciting.
Um, we're decked out in your swag from-
I have the socks.
Amazing. I love it.
We'll get to the part where you are a personal human machine of just productivity, and you do so much with... And, and, um, you know, I, I think, I think there's a lot to learn from people just on a personal level.
But I think a lot of people are gonna be here just for the state of Modular. We're also calling it The Shape of Compute, I think is gonna be probably the, the podcast title.
Yeah, it's super exciting. I mean, there's, there's so much going on in the industry with hardware and software and just innovation everywhere.
Most people can catch up on, like, the, the first episode that we did, and we, we introduced Modular. I think people wanna know, like, I think since then you said you open sourced it. There's been a lot of updates.
Uh, what would you highlight as, like, the past year or so of updates?
R&D Phase1:21
Yeah. So, so if you zoom out and say, what is Modular? We're a company that's just over three years old. Uh, three and a quarter, so-
Yeah
... we're quite a ways in. The first three years was a very mysterious time to a lot of people because I didn't want people to really use our stuff. Okay. And so why, why that? Why do you build something that you don't want people to use?
Well, it's because we're figuring it all out. And so the way I explain it is we were very much in a research phase. And so we were trying to solve this really hard problem of how do we unlock heterogeneous compute?
How do we en-enable GPU programming to be way easier? How do we enable innovation that's full stack across the AI stack by making it way simpler and driving out complexity? Like, there are these core questions, and I had a lot of hypotheses right?
But building an alternate AI stack that is not as good as the existing one isn't very useful, 'cause people will always evaluate it against state-of-the-art. What I did and what the team did and what we all did together is we said, "Okay, well, let's get to the point where at least Chris is happy."
And I have very high standards, and so we need to be state-of-the-art on NVIDIA GPUs, meeting NVIDI- beating NVIDIA's best on things like a Llama 3 model, which by the way, is serving end-to-end, like, very high bar. By the way, this is, like, something that's a pretty well-studied problem.
Like, it's not NVIDIA that's working on this. It's the entire industry-
You're up against, like, the best in the world
... that are-- that's doing this. Exactly. And so I-
And what, what is state-of-the-art, just as a rough?
I don't remember the number of tokens per second.
It's, it's tokens per second, and it's roughly between 800 to a thousand.
Yeah, shared GPT benchmark and, like, industry standard-
Yeah
... uh, kinds of things. And so, and so I said to the team, like, "Look, like, until we can do that, I don't believe it's real."
Yeah.
It turns out that, yeah, there's all these things. There's page attention, continuous batching. There's GPU curls. There's programming languages. There's, uh, a whole bunch of hardware stuff, and there, there's all this stuff. And I said, "Oh, by the way, we can't use CUDA."
That's the whole point.
That's the whole point, right? And so it's not a let's pick the shortest, easiest path to get to a demo. This is a let's do the hardest, most fundamental thing that everybody is telling me, as usual, that it's impossible.
Right? Th-this can't be done, right? And so for those first three years, right, the challenge is prove that we can do something that people think is impossible, and at least prove to me. And so what do you do for that?
What you do is you clear the deck. You try to get enough distractions out of the way so you can focus, you can iterate, you can move quickly. You don't have to argue with committees. And so we wanna be open to a certain extent because we want people to be aware of, like, Mojo and the things we were doing before, so we can hire people and there, there's very specific reasons.
We don't really want design by committee. We don't want a lot of these other things. And so across that research mode, it's like the primary thing I care about is prove that it's possible and make me happy. And we transitioned right at the end of December.
We had this release where we achieved that goal, and it was a super narrow release. It ran just on A100, just one model, but it had state-of-the-art performance. Uh, it was a full stack, vertically integrated thing. And my whole team is telling me, like, "Oh my God, it's fi-" Like, they pass out and could take time off for Christmas.
But then they come back and it's like, "Okay, well, everything sucks. It's not very good." Like, we got it to work, but-
They changed their minds after they went on break?
No, no, it worked.
Okay.
We achieved the goal, but it's not very good. The code is ugly. There's technical debt. There's this and that and the other thing. It only runs on A100. It's only one model. This isn't useful. This isn't, like, a valuable contribution.
And, uh, you know, I, I built some things in the past, right? And so I said, "Well, that's okay. We have one thing that works end to end, and we know 400 ways to make it better."
Yeah.
And so instant mode switch. And at that point you say, "Okay, well, let's just break it down like a normal engineering problem." It's not R&D anymore. It's an engineering problem. You say, "Okay, cool. Let's refactor these APIs. Let's deprecate this thing.
Let's, oh yeah, let's add H100 support. Let's add function calling and token sampling and, like, all the different things you need." You can project manage that.
Yeah.
Right? And so every six weeks, we've been shipping a new release. And so we added all the function calling features, and now you have agentic workflows. We have 500 models. We have H100 support. Uh, we're about to launch our AMD MI300 and 325 support.
That'll be a big deal for the industry. And as you do that, and Bl-Blackwell, like, all this stuff is, like, all now in the product. And so as this happens, now suddenly it's like, "Oh, okay, I get it."
But this is a very fundamentally different phase for us because it works, right? And once it works, you can see it end to end. Lots of people can put the pieces together in their brain, it's not just in my brain, and have few other, few other people that understand how all the different individual pieces work.
Now everybody can see it. So as we phase shift, now suddenly it's like, "Yeah, okay, let's open source it." Well, now we want more people involved. Now let's do each of these things. It's actually, we had hackathon, and so we invited 100 people to come visit and spend the day with us, and we learned GPU programming from scratch.
And so, you know, some of it... We, we built a very fancy inference framework. You know, the winning team for the hackathon took that, a four-person team, in one day. They had not used Mojo before. They hadn't programmed GPUs before, and they had built a training system.
They wrote an atom optimizer, a bunch of training kernels. They built a simple backprop system, and they actually showed that you could train a model using all the stuff we'd built for inference because it's so hackable, and also because AI coding tools are awesome.
Right.
But, uh, but th- this is the power of what you can do when, when you're ready to scale. Now, if we had done that six months ago or 12 months ago or something, it would've been a huge mess, right?
Because everything would break, and there are a lot of bugs and, you know, honestly, today it's still an early state system. There's still some bugs, but now it's useful. It can solve real-world problems, and so that's, that's the difference, and that's kind of the evolution that we've gone through as a team.
CPU to GPU6:55
I remember when we first had you, at the start, you were focused on CPU-
Yes
... actually optimization. How long did you work on CPU, and then how much of a jump was it to go from CPU to GPU?
Yeah. So, so the way I would... The way I explain Modular is if you take that first three-year R&D journey and if we round a little bit, first year was, um, prove compilation philosophy, and so this was writing very abstract compiler stuff and then prove that we could make a matrix multiplication go faster than Intel MKL's matrix multiplication on Intel silicon.
Mm-hmm.
And make it configurable and m-multiple D types and prove, like, a very narrow problem, um, and do that by writing this MLIR compiler representation directly by hand, which was really horrible.
Right.
But we proved the technology. It's a technology milestone. Year two was then say, "Okay, cool. I believe the, the fundamental approach can work, but guess what? Usability is terrible." Writing internal compiler stuff by hand sucks, and also MatMul is a long ways from an AI framework.
Mm-hmm.
And so year two embarked on two paths. One is Mojo. So programming language, syntax, member of the Python family, make it much more accessible and easy to write kernels and performance and all that kind of stuff. And then second, build an AI framework for CPUs, as you say, where you could go beat OpenVINO and things like this-
Mm
... on Intel CPUs. End of year two, we, we got and we're like, "Foof, we've achieved this amazing thing, but you know what's cool? GPUs." Right?
Mm-hmm.
And so again, two things. We said, "Okay. Well, let's prove we can do GPUs, number one. Also, let's not just build, uh, you know, air quotes, a 'CUDA replacement.' Let's actually show that we can do something useful. Let's take on LLM serving."
Yeah.
No big deal, right?
Right.
And so again, two things where you say, "Let's go prove that we can do a thing, but then validate it-
Mm-hmm
... against a really hard benchmark," and then that was... That's what, what brought us to year three.
Yep.
So... And each of these stages is really hard and lots of interesting technical problems. Um, probably the biggest problem that you face is you face people that are constantly telling you it's impossible. But again, you just have to be a little bit stubborn and believe in yourself and work hard and stay focused on milestones.
When they say impossible, do they mean impossible or very, very hard?
Well, so I mean, it's, it's common sense that CUDA is n-nearly 20 years old. NVIDIA's got hundreds of thousands of people working on it. The entire world's been writing CUDA code for many years.
Yeah, so just very, very hard.
And so it's... No, it's, I mean, many people think it's impossible for a startup to do anything in the space.
Yep.
Like, that's just common sense. All these people have thrown all this money at all these different systems. They've all failed. Why is your thing gonna succeed when all these other things built by other smart people have failed, right?
And so it's conventional wisdom that change is impossible. But hey, we're in AI. You know this. Like, change is all around us all the time, right?
Yeah.
And so what you need to do is you need to map out what are the success criteria, what causes change to actually work. And across my career, like with LLVM, all the GCC people told me it was impossible.
Right.
Like, "LLVM will fail because GCC is 20 years old, and it's had hundreds of people working on it, and blah, blah, blah, and spec benchmarks and whatever." Nobody told me it was impossible outside because it was secret. And so that was a little bit different.
But everybody inside Apple that knew about it said, "No, no, no. Objective-C is fine. Like, we should just improve Objective-C. The world doesn't need new programming languages. That new programming languages never get adopted."
Yeah.
Right? And, and it's common sense that, like, new programming languages don't go anywhere. That's conventional wisdom. You know, MLIR, super funny. MLIR is another compiler thing, and so built this thing, brought it to the LLVM community and said, "Hey, we open sourced this.
Does LLVM want it?" I know a few LLVM people, right? It was my PhD project, and all the LLVM illuminati in the community had been working on LLVM for 15 years or something, and they're like, "No, no, no.
LLVM's good enough. We don't need a new thing. Machine learning is not that important." You know? It's... And so again, uh, obviously have developed a few skills to work through this kind of challenge-
Mm-hmm
... but you get back to the reality that humans don't like change, and when you do have change, it takes time to diffuse into the ecosystem for people to process it, and this is where when you talk about hackathon, well, you kinda have to teach people about new things.
And so this is why it's, um, like, really important to take time to do that, and the blog post series and things like this are all kind of part of this, like, education campaign because none of this stuff is actually impossible.
It's just really hard work. It requires an elite team and good vision and guidance and stuff like this, but, um, but it's understandably conventional wisdom that it's impossible to do this.
MAX Framework11:14
And is the idea to build serving basically, I don't need to tell you how much better this is. I can just actually serve the models using our platform, and then you can see how much faster it is, and then obviously you'll adopt it.
Yeah. Well, so, uh, today you can download MAX. It's available for free. You can scale it to thousands of GPUs you want for free. I'll tell you some cool things about it. So it's not as good as something like VLLM because it's missing some features, and it only supports NVIDIA and AMD hardware, for example.
But by the way, the container's like a gigabyte.
Wow.
Right? Well, why is it a gigabyte? Well, it's a completely new stack. It doesn't use CUDA.
Mm-hmm.
You can run arbitrary PyTorch models, and if, if that's the case, then okay, you pull in some dependencies. But, like, if you're running the common LLMs and the GenAI models that people really care about, well, guess what? It's a completely native stack.
It's super efficient, doesn't have Python in the loop and for eager mode opt dispatch and stuff like this. Because you don't have all that dependency, guess what? Your server starts up really fast.
Yeah.
So if you care about horizontal auto-scaling, that's actually pretty cool. If you care about reliability, it's pretty cool that you don't have all these weird things that are s-stacked in there. If you're, uh, wanting to do something slightly custom, guess what?
You have full control over everything, and stuff's all open source, and so you can go hack it. This thing's more open source than VLLM. 'Cause VLLM depends on all these crazy binary CUDA kernels and stuff like this that are just opaque blobs from NVIDIA, right?
And so this is a very different world. And so I don't want everybody just like overnight drop VLLM 'cause I think it's a great project, right? But, but I think there's some interesting things here, and there's, uh, specific reasons that it's interesting to certain people.
Can you just maybe run people through the pieces of MAX?
Yeah.
Because I think last time none of this existed .
Mojo Language12:52
Yeah, exactly. Uh, last time has, uh, been a long time ago, both in AI and also in-
Yeah
... Modular time. Um, yeah, so the bottom of our stack is... We, we think about it as concentric circles. So the inside is a programming, programming language called Mojo. Chris, why did you have to build another programming language, right?
Mm-hmm.
Well, the answer is that none of the existing ones actually solve the problem. So what's the problem? Well, the problem is that all of compute is now accelerated. You have GPUs, you have TPUs, you have all these chips.
You have CPUs also, and so we need a programming language that can scale across this. And so if you go shopping and you look at that, like the best thing-ish is C++. Like, there are things like OpenCL and SYCL, and there's like a million things coming out of the HPC community that have tried to scale across different kinds of hardware.
Let me be the first to tell you, and I can say this now, I feel, I feel comfortable saying this, that C++ sucks.
Yeah .
I've earned that right- ... having written so much. And so... And let me also claim, AI people generally don't love C++. What do they love? They love Python, right? And so what we decided to do is say, "Okay, well even...
And even within the space of C++ there isn't an actually good unifying thing that can talk to tensor cores and things like this." And so we said, "Okay, I want..." Again, Chris being unreasonable . I want something that can expose the full power of the hardware, not just one chip-
Mm-hmm
... coming from one vendor, but the full power of any chip. It has to support the next level up, very fancy compilers and other stuff like this with graph compilers and stuff like that. It needs to be portable, so portable across vendors, and enable portable code.
So yeah, it turns out that, uh, an H100 and an AMD chip are actually quite different.
Mm-hmm.
That is true. But there's a lot that you can reuse across that, and so the more common code that you can have, the better. The other piece is usability. And so we want something people can actually use and actually learn, and so it's not good enough just to have Python syntax.
Mm-hmm.
We wanna have performance and control and, again, full power of the hardware. And so this is where Mojo came from. And so today, Mojo's really useful for two things. And so, um, Mojo will grow over time, but today I, I recommend you use it for places you care about performance, so something running on a GPU, for example, or really high performance.
Uh, you're doing continuous batching within a web server or something like this where you have to do fancy hashing and... Like, if, if you care about performance, Mojo is a good thing. Other cool thing that we're about to ship, so, you know, stay tuned, is it's the best way to extend Python.
And so if you have a big blob of Python code, you care about performance, you wanna get the performance part out of Python, you can make it... We make it super easy to move that to Mojo. Mojo is not just a little faster than Python, it's faster than Rust .
And so it's like tens of thousands of times faster than Python, and it's in the Python family. And so you can literally just, like, rip out some for loops, put it in Mojo, and now you get performance wins.
And then you can start improving the code in place. You can offload it to GPU, you can do this stuff, and all the packaging is super simple, and it's just a beautiful way to make Python code go fast.
Just, just to double-click, you, you said you're about to ship this?
It's technically in our nightlies, but we haven't actually announced it yet.
Is it? Okay. Um-
We, we do, we do a lot of that, by the way, 'cause we're very developer-centric.
Yeah.
And so if you join our Discord or Discourse-
Nightly's is great. It's like the best program. Uh-
Yeah, yeah. And so we have a ton of stuff that's-
Yeah
... kind of unannounced but well-known in the community.
Okay.
And so-
I thought this was already released.
Yeah. Yeah.
Um- And when you say rip out, is this... You have a Python binding to then run the Mojo, like you have C bindings or however-
This is the cool thing, is it's binding-free. So think about like... So today, again... Sorry, I, I get excited about this. I forget how tragic the world is without- ... outside of our walls. Um, the thing we're competing with is if you have a big blob of Python code, which a lot of us do, you build it, build it, build it, build it.
Performance becomes a problem. Okay, what do you do? Well, you have a couple of different things. You can say, "I'm gonna rewrite my entire application in Rust or something," right? Some people do that. The other thing you can do is you can say, "Okay, I'm gonna use PyBind or NanoBind or some Rust thingy and rewrite a part of my module, the, the performance critical part."
Mm-hmm.
"And now I have to have this binding logic and this build system goop and all this complexity around-
Yeah
... I have Rust code over here and I have Python code over here. Oh, by the way, you now have to hire people who can work on Rust and Python, and like that .
Right.
Like this is, this fragments your team. It's very difficult to hire Rust people. Like, I love them, but it's, there's just too few of them, right? And so what we're doing is we're saying, "Okay, well, let's keep the languages basically the same."
Mm-hmm.
So Mojo doesn't have all the features that Python does. Notably, Mojo doesn't have classes. But it has functions. Yeah, you can use arbitrary Python objects. You can get this binding-free experience, and it's very similar. And so you can look at this as being like a super fast, slightly more limited Python, and now it's cohesive within the ecosystem.
Yep.
And so I think this is useful way beyond just the AI ecosystem. I think this is a pretty cool thing for Python in general. But now you bring in GPUs and you say like, "Okay, well, I can take your code and you can make it real fast."
'Cause CPUs have lots of fancy features like their own tensor cores and SIMD and all this stuff. Then you can put it on a GPU, and then you can put it on eight GPUs, and then you can... A- And that, you know...
And so, like, you can keep going, and this is something that Rust can't do.
Right.
You can't put Rust on a GPU really, right? And things like this, and this is, this is part of, like, this new wave of technology that we're bringing into the world. So-
So-
And, and also, like... I'm sorry. Like I get excited about the stuff we're, we've shipped and what we're about to ship. But, but this fall is gonna be even more fun .
Okay.
So stay tuned.
All right.
There may, there may be more hardware beyond AMD and NVIDIA.
MAX AI Stack18:25
Nice. So that was Mojo.
The first concentric circle Yeah. Thank you for keeping me on track. So the inner circle is the programming language, right? And so it's a good way to make stuff go fast. Mm-hmm. The next level out is you say, "Okay, well, you know what's cool?
AI." Mm-hmm. Have I convinced you? Um And so if you get into the world of AI, you start thinking about models. And beyond models, now we have GenAI, and so you have things like pipelines, right? An entire pipeline where you have KV cache orchestration, and you have, uh, uh, stateful batching, and I mean, you guys are the experts.
Agentic everything and all this kind of stuff. And so next level out is an AI-- a very simple GenAI inference-focused framework we call MAX. And so MAX has a serving component- Is it MAX Engine or- Or MA-- yeah.
Okay. There's different sections. We just call it MAX. Okay. Um, we, we got too complicated with sub-brands. This is also part of our R&D on branding is that- HP has the same problem. Yeah, exactly. Well, and also- Like who named LLVM?
I mean, what the heck is that, right? So, so- Honestly, it's short, it's Googleable. Not the worst. Yeah, and VLLM came and decided to mess with it. Yeah, yeah. So MAX, the way to think about it is it's, it's not really a-- it's not a PyTorch.
That's, that's not what it wants to be, but it's really focused on inference. It's really focused on performance and control and latency and, and if you wanna be able to write something that then gets Python out of the loop of the model logic, like it's really good for control.
And so it dovetails and is designed to work directly with Mojo, and so within a lot of LLVM-- or LLM applications, as you all know, there's a lot of very customized GPU kernels. And so you have a lot of crazy forms of attention, like the DeepSeek things that just came out and, like, all this stuff is always changing, and so a lot of those are custom kernels.
But then you have a graph level that's outside of it, and the way that has always worked is you have, for example, CUDA or things like Triton Lang or things like this on the inside, and then you have Python on the outside.
We've embraced that model. Like, don't fix what ain't broken. So we use straight Python for the model level. And so we have an API. It's very simple. It's not, it's not, like, designed to be fancy, but it feels kind of like a very simple PyTorch.
And you can say, "Give me an attention block, give me these things. Like, configure out these ops," but it directly integrates with Mojo. And so now it's-- you get full integration in a way that you can't get because none of the other frameworks and things like this can see into the code that you're running.
And so this means that you get things like automatic kernel fusion. What's that? Well, that's a very fancy compiler technology- Mm-hmm ... that allows you to say, "Okay, you will write one version of, uh, flash attention," and then cool, we can autofuse in Silu and the other activation functions that you might wanna use, and you don't have to write all the permutations of these kernels.
Mm-hmm. Well, that just means you're more productive. That means you get better performance. I mean, it's like a lot of thing-- it just drives down complexity in the system, and so you shouldn't have to know there's a fancy compiler.
Yeah. Everybody should hate compilers. Like, the only reason people should know about compilers is if they're breaking, right? Right. And so it just feels like a, a very nice, very ergonomic and efficient way to, like, build custom models and customize existing models and things like this.
And so with MAX we have five, six hundred very common model families and, and, uh, implemented in that. You can see, uh, builds.modular.com. We have a whole bunch of models, and you can scroll through them, and you can get the source code and play with them and do that.
That's actually really great for people who care about serving and research and, and, and all this, this kind of stuff. Next layer out is you say, "Okay, well, you have a very fancy, uh, way to do serving on a single node."
That's pretty useful and pretty important- Mm-hmm ... but you know what's actually cool? Large scale deployment. And so we have a next lever out-- level out, cluster level. And so that's the, "Okay, cool, I have a Kubernetes cluster.
I've got a platform team. They've, uh, got a three-year commit on three hundred GPUs, and now I have product teams, and I wanna throw workloads against this shared pool of compute." And the folks carrying the pagers want the product teams to behave- Mm-hmm ...
and so they wanna keep track of what's actually happening, and so that's the cluster level that goes out. And each of the... And so very fancy, um, prefix caching on a per-node basis, and then you have intelligent routing, and there's a whole bunch of, uh, disaggregated prefill, and, like, a whole bunch of cool technologies at each of these lay-layers.
But the, the cool thing about it is that they're all co-designed. Yep. And because the inside is heterogeneous, you can say, "Hey, I have some in AMD, I have some NVIDIA. Hey, I throw a model on there, run it on the best architecture."
Well, this actually simplifies a lot of the management problems, and so a lot of the complexity that we've all internalized as being inherent to AI is actually really a consequence of these systems that are being designed together. And so to me, what, you know, my number one goal right now is to drive out complexity, both of our stack, 'cause we do have some tech debt that we're still fixing, but, um, but of AI in general.
And AI in general has way more tech debt than, than it deserves. Well, yeah, I mean, there's only so much that most people, um, excluding you, can do to comprehend their s- part of the stack and optimize it.
I'm curious, uh, if you have any views or insider takes on what's happening with VLLM versus SGLang and everything coming out of Berkeley. Yeah. I, I, I, I don't. I have outsider perspective. Yeah. I don't have, have any inside knowledge.
Um, SGLang seems, to me as an outsider, so I'm not directly involved with either community- Sure ... I just don't have time. Um, so I can be a fanboy without h- without participating. But, um, SGLang- They, they have a beef going on.
So, see, I, I don't know the politics- Yeah. ... and drama. But SGLang to me seems like a very focused team that has specific objectives and things they wanna prove, and they're executing really hard towards a specific set of goals, much like the Modular team, right?
And so that's-- so I think there's some kindred spirits there. VLLM seems much more like a massive community with a lot of stakeholders, a lot of stuff going on, and it's kind of a hot mess. And so- But they wanna be like- It's, it's really cool- ...
the default inference platform that everyone kind of benchmarks against. Well, uh, and, and I mean, it's, as far as I know, it's, like, crushing TrtLLM and some of the other older systems and things like that. And so I think, like, metrics are good probably for that.
And so I, I can't speak to their ambition. Mm-hmm. But, but it seems structurally like they're very different approaches. One is say, "Let's be really good at a small number of things that are really important." Yeah. That's our f- that's our approach too.
One is saying, "Let's just say yes to lots of things, and then we'll have lots of things, and some of them will work and some of them don't." One of the challenges I hear constantly about VLLM- For what it's worth is if you go to, go to their webpage, they say, "Yeah, we support all of this hardware.
We support Google TPUs and Inferentia and AMD and NVIDIA, obviously, and CPUs," and this and that and the other thing. So they have this huge, huge list of stuff that they, they support. But then you go and you try to follow any demo on some random webpage saying, "Here's how you do something with VLLM," and if you do it on a non-NVIDIA piece of hardware, it fails.
Mm.
And so what is the value of something like VLLM? Well, the value is you wanna build on top of it generally. You don't wanna be obsessing about the internals of it. If you have something that, you know, advertises that it works and then you pick it up and it doesn't work and now you have to debug it, it's kind of a betrayal of the goal.
And so to me, um, again, I know how hard it is. I know the fundamental reason why they're trying to build on top of a bunch of stuff below them that they didn't invent that doesn't work very well, honestly, right?
'Cause we know that har- hardware is hard and software for hardware is even harder these days. Um, so I understand why that is, but, but that approach of saying, "Here's all the things we can do," and then having a sparse matrix of things that actually work is very different than the conservative approach, which is saying, "Okay, well, we do a small number of things really well, and you can rely on it."
And so I can't say to which one's better. I mean, I appreciate them both, and they both have great ideas, and the fact that there's the competition is good for everybody in the industry.
It is.
And so, you know-
They're, they're in the benchmark wars state of the way that this industry evolves.
That's right. But, but, but I also hear just on the enterprise side of things that they're also good. They don't wanna follow the drama, right?
Yeah.
And so, like, this is where having less chaos and more, "I can work with somebody," is actually really valuable these days.
Yeah.
And, um, and you need state-of-the-art, you need performance and things like this, but, but actually having somebody that, uh, is executing well and works together, w- can work with you is also super important.
Have we rounded out the offerings? You, you talked about MAX.
Yeah, yeah.
I realized that it impresses me that you named the company Modular, and you've designed all these things in a very modular way.
Yeah.
I'm wondering if there are any sacrifices to that, or is modular everything the right approach? Like there's no trade-offs, it seems.
If you allow me to show my old age, um, like my, my obsession with modular design came from back when creating LLVM, okay? So at the time there was GCC. GCC is a v- again, I love it as a C and C++ compiler and stuff, right, and have tons of respect for it, but it's a monolith.
It was designed from the old school Unix, like, like compilers are Unix pipe. Everything is global variables, C code. It was written in K&R, if you even know what that is anymore. And so it was from a different epoch of software, and LLVM came and said, "You know what everybody teaches you in school is you have a front end, an optimizer and a code generator.
Let's make that clean," right? And so at the time, people told me, again, first of all, it's impossible to replace GCC because it's so established, et cetera, et cetera. But then also that you can't have clean design because look at what GCC did and it's successful, and therefore you can't have performance or good support for hardware or whatever it is without that.
They might be right if in a infinitely perfect universe, but we don't live in an infinitely perfect universe. Instead, what we live in is a universe where you have teams. Like you have people that are really good at writing parsers, people that are really good at writing optimizers, people that are really good at writing a code generator for x86 or something like that, right?
And so these are very different skill sets. Also, as like we see in AI, the requirements change all the time. And so if you write a monolith that is super hard code and hack today, well, two years from now is it gonna be relevant?
And this is why what we see a lot in AI is we see these like really cool and very promising systems that rise and grow very rapidly, and then they end up falling. It's because they're almost disposable frameworks.
This concept really comes from-
The language is a framework, yeah.
It's, uh, each of these systems end up being different, but, but what I believe in very strongly is that in AI we want progress to go faster, not slower. That's controversial because it's already going so fast, right?
Mm-hmm.
But the, uh, but if we can accelerate things, we get more product value into our lives, we get more impact, we get all the, all the good things that come with AI. I'm a, I'm a maximalist, by the way.
But with that comes the reality that everything will change and break. And so if you have modularity, if you have clean architecture, you can evolve and change the design. It's not overspecialized for a specific use case. The challenge is you have to set your baseline and metrics right.
This is why it's like compare against the best in the industry, right? And so you can't say, or at least it's unfulfilling to me to say, "Let's build something that's 80% as good as the best, but it's got this other benefit."
I wanna be best of in the categories we care about.
Open Source & Competition29:16
And when it comes to using like a prefix caching, page attention-
Mm-hmm
... how do you decide what you wanna innovate on versus like, "Hey, you are actually the team that has built the best thing. We're just gonna go, go ahead and use that"?
Yeah. I, I'm very shameless about using good ideas from wherever they come. Like everything's a remix, right? And so if somebody, if, I don't care who it is, if it's NVIDIA or if it's SGLang or VLLM, and if somebody has a good idea, let's pull it together.
But the key thing is make it composable and orthogonal and flexible and expressive. To me, what I look at is not just the things that people have done and put into VLLM, for example, but the continuous stream of archive papers, right?
And so I follow, you know, there's a very vibrant industry around inference research.
Yep.
It used to be it was just training research, right? And much of that doesn't, never gets into these standard frameworks, right? And the reason for that is you have to write massively hand-coded CUDA kernels and all this stuff for any new thing, and I would want one new D type and now I have to change everything because nothing composes.
Mm-hmm.
And so this is again, where if you get some of these software architecture things right, I mean, which admittedly requires you to invent a new programming language-
Right
... there's a, there's a few hard parts to this problem, but, but the cool thing about that is you can move way faster. And so I'll give you an example of that. This is fully public 'cause not only do we open source our thing, we open source all the version control history.
And so you can go back in time and say, uh, and like, Chris likes open source software, by the way. Let me, let me convince you of this. I don't like people meddling in my stuff too early But, but I, I like, I, I like open source software, and so you can go look at how we brought up H100, built Flash Attention from scratch in a few weeks.
Built, like, all the stuff. We're beating the tree data reference implementation that everybody uses, for example, right? Written fully in Mojo. Again, all of our GPU kernels are written in Mojo. Uh, you can go see the history of the team building this, and it was done in just a few weeks, right?
And so we brought up the H100, entirely new GPU architecture. If you're familiar, it has very fancy asynchronous features and tensor memory accelerator things, and like all this goofy stuff they added, which is really critical for performance. And again, our goal is meet and beat cuDNN and-
Mm-hmm
... TRTLM and these things, and so it's not... You can't just do a quick path to success. You have to get everything right and everything to line up, 'cause any little thing being wrong nerfs performance. And we did that in, uh, I, I think it was less than two months.
Public on GitHub, right? And so, um, like, that velocity is just incredible. I think, uh, it took nine months or 12 months to invent Flash Attention, and it took another six months to get into VLLM. And this is just in- Like, the, that latter part is just integration work.
Mm-hmm.
Right? And, and so now you're talking about, like, building all this stuff from scratch in a composable way that scales against other architectures and has these advantages. It's just a different, different level. So anyways, I mean, our stuff's still early in many ways, and so we're missing features.
But, um, if you're interested in what we're doing, you can totally follow the change logs and you can f- like follow them nightly, where we publish all the cool stuff, and all the kernels are public, and so you can see and contribute to this as well, if you're interested.
Yeah. Do you have any requests for projects that people should take on outside of Modular-
Yeah
... that you don't wanna bring in?
Absolutely. Um, so there's... So we're a small team. I mean, we're over 100 people, but we're compared to the size of the problem we're taking on-
Of the problem. The size of your ambitions. Yeah.
Yeah. We're, we're infinitesimal compared to the size of the AI industry, right? And so this is where, um, for example, we didn't really care about Turing support. Turing was an older GPU architecture, and so somebody on the community is like, "Okay, I'll enable a new GPU architecture too."
So they contributed Turing support. Now you can use MAX in Colab for free. That's pretty cool. There's a bunch of operators. So we're very focused on AI and gen AI and things like this. By the way, our stuff isn't AI specific.
Right.
So people have written ray tracers, and there's people doing flight safety and, like, all kinds of weird things- ... that... Flight simulation. At our hackathon somebody made a demo of, uh, like, looking at, I think it was a voice transcript, the, the black box type traffic, and then predicting when the pilot had made a mistake and the airplane was gonna have a big problem, and predicting that with high confidence.
So it's like your car, you know. You're, you're driving your car until it starts beeping when you're not slowing down.
It's like lane assist, yeah. Yeah, yeah, yeah.
Yeah. Okay, that kind of thing-
Right, right
... but for, for the FAA
Wow.
And stuff like this, right? And I know nothing about that. Like, this is not my domain-
Right
... trust me. This is the power of there's so many people in our industry that are almost infinitely smart, it feels like. They're way smarter than I am in most ways. And you give them tools, and you enable things, and also you have AI coding tools and things like this that help, like, bridge some of the gap.
I think we're gonna have so many more products . Like, this is, this is really what motivates me, right? And this is where, um, I think we talked about it last time, one of the things that really frustrated me years ago and inspired me to start Modular in the first place is that I saw what the trillion-dollar companies could do, right?
You look at the biggest labs with all the smart people that built all the stuff vertically top to bottom, and they could do incredible product and research and other work that, you know, nobody with a five-person team-
Mm-hmm
... or a startup or something could afford to do. And here we're not even talking about the compute, we're just talking about the talent. Well, and the reason for that is just all the complexity. It only worked, uh, for example, at Google, because you could walk, proverbially walk down the hallway-
Mm-hmm
... tap on somebody's shoulder and say, "Hey, your stuff doesn't work. How do I get it to work? I'm off of the happy path. Like, how do I get this new thing to work?" And they'll say, "I'll hack it for you."
Right.
Right? Well, that doesn't work at scale. Like, we need things that compose right, that are simple, that you can understand, that you can cross boundaries, and, you know, so much of AI is a team sport, right? And we want it so more people can participate and grow and learn.
And if you do that, then I think you get, again, more product innovation and-
Yeah
... less just, like, gigantic AI will solve all problem kind of solutions, and more fine-grained, purpose-built, product-integrated AI.
Yeah. I, I think one way of phrasing what you're doing is, you know, y- y- you have this line about it's, it's... You're, you want AI to accelerate, and that's contrarian because it's already fast and people are already uncomfortable.
But I think you're more... You're, you're accelerating distribution. Um, I have mixed feelings about the words democratizing, but that's really what you're doing.
Well, it used to be... Democratizing AI used to be the cool thing back in 2017.
I know. It's not, not cool anymore.
Right.
But-
Yeah, I know. But it used to be the cool thing, and what it meant and what it came to mean is democratizing, uh, model training, right? And it's super interesting, again, as veterans. Like, uh, back in 2017, AI was about the research-
Hmm
... because nobody knew both what the product applications were, but then we didn't know how to train models. And so things like PyTorch came on the scene, and I think PyTorch gets all credit for democratizing model training.
Mm-hmm.
Right? It's taught to pretty much every computer science student that, uh, graduates. That's a huge deal. But nobody democratized the inference. Inference always remained a black art, right? And so this is why we have things like VLLM and SGLang.
Yep, 'cause-
These are the black box that you can just hopefully build on top of and not have to know how any of that scary stuff works. It's because we haven't taught the industry how to do this stuff. And, you know, it, it'll take time, but I think that that's not an inherent problem.
I think it's just that we don't have, like, a PyTorch for inference or some- something like that. And so as we start making this stuff easier and breaking it down, we can get a lot more diffusion of good ideas.
Yeah. I wanna double-click a little bit more on some technical things.
Yeah.
But just, just to sidetrack on... You know, VLLM is an open source project led by academics and all that, and I think a lot of the other inference teams, 'cause effectively every team is, is a startup, like the Fireworks together, you know, all the, all those guys.
Your business model is very different-
Yeah
... from them, and I, I wanna spend a little bit of time on that.
Happy to talk about it.
You, you, you intentionally... Well, you believe in open source, but it, you know, it's not, it's not just that. You, you just choose not to make money from the, a lot of the normal sort of hosted cloud offerings that everyone else does.
Yeah. There's a philosophical reason and a differentiation reason. There's a whole bunch of reasons.
Yeah. So, like, uh, maybe remind people of what that is. Like, how, basically, how do they pay you money, and wh- what they get for that, and like-
Yeah
... why you, and why you picked that.
So again, we're, we're, we're doing this on hard mode, right?
Yeah.
We took a long path to product. We're, we're not just write a few CUDA kernels and have some alpha and go buy a GPU reserve thing and resell our GPUs. Like, that, that path has been picked by many companies, and they're really good at it .
So that's not a contribution that I'm very good at and, um, and I'm not gonna go build a data center for you. Like, that's-
Mm-hmm
... there's people that are way, way better than that.
Okay.
And so all the best luck. I want that-
Literally Crusoe, like, walking the Stargate grounds like...
I, I want, I want those people to use MAX and so-
Yeah
... like, you know, I'm, I'm pretty good at this-
classic Crusoe, yeah.
Yeah, I'm pretty good at this software thing, and so you can handle all the compute. Moreover, if you get out of startups, you have a lot of people that are struggling with GPUs in cloud, right? And so GPUs in cloud fundamentally are a different thing than CPUs in cloud.
And a lot of people walk up to it and say, "It's all just cloud, right?" But let me convince you that's not true . And so first of all, CPUs in cloud, why was that awesome? Well, all the workloads were stateless.
They all could horizontally autoscale. CPUs are comparatively cheap, and so you get elasticity. That's really cool. Like-
You can load them pretty quickly.
Yeah.
It's not like gigabytes of weights.
Yeah. And so it turns out what business do you know knows what they're doing in two and a half years?
Nobody.
Nobody, right? And so cloud for CPUs is incredibly valuable because you don't have to capacity plan-
Mm-hmm
... that far out, right? Now you fast-forward to GPUs. Well, now you have to get a three-year commit. A three-year commit on a piece of hardware that Jensen's gonna make obsolete in a year.
Mm-hmm.
Right? And so now you get this thing, and so you make some big commit, and what do you do with it? Well, you have to overcommit because you don't know what your needs are gonna be, and you're not ready to do this.
Also, all the tech is super complicated and scary. Also, GPU workloads are stateful. And so you talk about, like, the fancy agentic stuff and-
Catching pipelines, yeah
... y'all know, know all this stuff. Yeah. It's, it's stateful, and so now you don't get horizontal autoscaling. You don't get stateless, uh, elasticity. And so you get a huge management problem. And so what we think we can do and we can help with people is say, "Okay, well, let's give you power over your compute."
And a lot of people are, have different systems, and there's very simple systems that go into this, but you can get, like, five X performance TCO benefit by doing intelligent routing. That's actually a big deal. For a platform team, they don't like to have to deal with this.
They want hardware optionality to get to AMD. They want this kind of power and technology, and so we're very happy to work with those folks. The way I explain it in, in a, in a simple way is that a lot of the endpoint companies, and there's a lot of them out there, and so you can't make...
They're not all one thing. They obviously have pros and cons and trade-offs, but, but generally the value prop of an endpoint is to say, "Look, AI and AI software and applications and workloads, it's all a hot mess. It's too complicated.
Don't worry your little head about it. I'll take AI off your plate so that you don't have to worry about it. We'll take care of all the complexity for you, and it'll be easy. Just talk to our endpoint."
Our approach is to say, "Okay, well, guess what? It's all a hot mess. Yes, 100%. Like, it's horrible. It's e- it's worse than you probably even know, and get, get to tomorrow it's gonna be even worse 'cause everything keeps changing.
We'll make it easy. We'll give you power. We'll give you superpowers for your enterprise and your team," because every CEO that I talk to not only wants to have AI on their products, they want their team to upskill in AI.
And so we don't take AI away from the enterprise. We give power over it back to their team and allow them to both have an easy experience to get started, 'cause a lot of people do wanna run standard commodity models, and you do want stuff to just work as table stakes.
But then when they say, "Hey, I wanna actually fine-tune it," well, I don't wanna give my proprietary data to some other startup, right? Or even some big startups out there, and you, you know whose these are. That's my proprietary IP.
And then you get to people who say, "Hey, I have a fancy data model. I actually have data scientists. I actually have a few GPUs. I'm gonna train my model." Cool. That got democratized. Now how do I deploy it?
Right? Well, again, you get back into hacking the internals of VLLM, and PyTorch isn't really designed for KV cache optimizations and all the modern transformer features and things like this, and so suddenly you fall off this complexity cliff if you care about it being good.
And so we say, "Okay, well, yeah, this is another step in complexity, but you can own this, and you can scale this, and so we can help you with that." And so it's a different, it's a different trade-off in the space.
But I will admit that their time to market and revenue growth and stuff like that has been much faster because they didn't have to, like, build an entire replacement for CUDA to get there.
Business Model41:37
Nice.
And when it comes to, like, charging, people are buying this as a platform. It's not tied to, like, token, inferred, anything like that.
Yeah.
We have two, two things going on. So the MAX framework and the Mojo language, free to use on Nvidia and CPUs, any scale. Go nuts. Do whatever you want. Please send patches. Like-
Right
... it's free, right? Why is this? Well, it turns out CUDA's already free. Nvidia's already dominant in here. Uh, we want the technology to go far and wide. Use it for free. Like, it would be great if you send us patches, but you don't even have to do that, right?
Um, we do ask you to allow us to use your logo on our webpage, and so send us an email and say, "Hey, you're using it," and if you're scaling it on 10,000 GPUs, that'd be awesome. But, but that's the only requirement.
Um, if you want cluster management and you want enterprise support and you want things like this, then you can pay on a per GPU basis.
Yeah.
And you can contact our sales team, and then we can work out a deal, and we can work with you directly. And so that's, that's how we break this down. And also let, let me say, like, one thing I would love to see, and again, it's still early days, but I would love to see PyTorch.MAX- I'd love to see VLLM adopt MAX.
I'd love to see SGLang adopt a MAX. Like, like we have our own little serving thing, but go look at it. It's really simple. Like, it would be amazing. And, again, we're in the phase now where I do want people to actually use our stuff.
Yeah.
And so we have historically been in, in the mode of like, "No, our stuff's closed. Stay away" Like it's-
Yeah.
But, but we're, we're phase shifting right now, and so you'll-
Yeah
... see much more of this being announced and-
Yeah. Um, I- I don't know how to make this happen, but, like, I think you win when Mistral, Meta, DeepSeek, and Qwen adopts you and, like, ship you natively, right? How does, how does that happen at that point?
I don't wanna talk that far into the future.
BizDev.
But we, we, we, we may have a-
I don't think it's that far
... we may have an industry-leading state-of-the-art model launching first on Mac soon.
I'll stay tuned for that, but also let us know.
Yeah. So but that hasn't happened, so assume it doesn't happen.
I think, like, that's when it really tips because then they're... everyone's like, "Okay, if it's good enough for them, it's good enough for us," right?
Mm-hmm.
And then-
Yep
... you get the rest of the industry.
Yeah.
Um-
But, but, but again, I mean, I'm, I'm in it for a long game, right?
Yeah.
And, and I realize, again, like, the stuff I work on takes... it takes time for people to process. And so what we need to do and what I want us to do and what I ask the team to do is keep making things better and better and better and better and better.
Mm-hmm.
And there's, like, an S curve of technology adoption. And so I think it's great that there's a small number of crazy early adopters that were using our stuff in February-
Yeah
... before it was open source, and, like, like, it made no rational sense. It ver- barely worked. But it was amazing, and I'm very thankful for those people. And then, of course, we open source it, and we start teaching people, and you get a much bigger adoption group.
You make it free, go adopt and, and go and, and then as you say, there's more validation that'll be coming soon. And, like, each of these things is a knee in the curve, but what it also does is it gives us the ability to, like, uh, fix bugs and improve things and add more features and, and roll out new capabilities.
How does this feel rolling this out as compared to, like, Swift?
Oh, uh, well, so let me reinterpret your question of-
Okay
... uh, given you've done a few interesting things in the past, uh, what have you learned and what are you not doing again?
Yeah. That actually is a better question than the one I asked, so thank you. Because it... Swift is too narrow almost.
The character of Swift, so just... 'cause I assume most people don't know about this. The character of Swift was I started as a nights-and-weekends project in 2010, hacked on it alone nights and weekends for a year and a half, eventually told management at Apple about it.
Their heads exploded. Like, "Why do we need a something? Objective C's good enough. Why do you need this?" Um, got approval to have a couple more people get involved in it.
You, you were on a fellowship or an internship at, at Apple at the time?
No, I was leading the developer tools team.
You, okay, got it.
So yeah.
Sorry.
I was, I was leading a-
I think you were, you were doing something... Yeah. Okay.
Yeah, yeah, no, I was leading a huge team, and I... Let's just say this was not my day job. But so it started in 2010. It launched publicly by Apple in 2014, okay? And by the time it launched in 2014, only about 250 people in the world knew about it, most of whom were in my team.
About 200 and something of them were in my team, and then it was senior execs, marketing, Tim Cook, like, et cetera, right? And this was the category of people that knew about Swift. So we had built it in secret, literally an NDA, you know, within Apple, to know about it, right?
When we launched it, part of the requirement was that you had to be able to submit apps to the App Store in Swift, right? And that was a requirement input put on me. And so it's like, cool, that sounds great.
And so we launched it and said it's a 1.0. So you're launching a 1.0 brand-new programming language. Nobody has ever seen it before. No internal user, like one demo app. It was a fricking nightmare.
Yeah.
And so it was a nightmare for the community because, I mean, fortunately, a lot of people were excited and wanted to adopt it, and a lot of people did adopt it right away. But it was not battle-hardened and had tons of bugs.
We should've launched it as a 0.5, right?
Mm.
And so it took another year for it to become pretty good and then two years for it to become quite good, in my opinion. Also, none of the software engineers at Apple knew about it, and so they, their heads explode, and they said, "Wait a second.
Why are you replacing Objective C? I joined Apple because I love Objective C. Why didn't you ask me my opinions about the new programming language," right? And so there's that whole dynamic. And so-
Oh, was there a company mandate that they had to write Swift from now on?
No.
Oh, okay.
But, but still, it's like, "Wait a second. This isn't the company I thought I joined," right?
Mm-hmm.
And stuff like this, right? And so, and so there's this huge amount of turmoil and drama and nonsense that came out of that. And so, okay, fast-forward to Mojo, lessons learned. Hey, uh, one, don't have a hot start.
And so we launched Mojo a long time ago, before it even made sense, and we called it a 0.1. And so that... how's that honesty in advertisement, right?
Right. Yeah, yeah.
It's like, like, "This is 0.1. Please don't use it, but if you're interested-
Uh-huh
... we'd love your feedback," right? And so soft start, go. Second thing that's very different is that in Swift, we had one demo app, and so you have, you know, a very, I think, high-powered team building a language.
They had done lots of credible stuff but was building a language for iOS developers, and the compiler's written in C++, right? And so, yeah, there's sympathy for the user, but not a lot of understanding and a lot of, not a lot of learning internally when we launch.
In the case of Mojo, guess what? Modular is Mojo's first customer. Like, we have more Mojo code in our repository than any other language, right? And, and-
That's awesome
... it's open source. Like, and we, we open source like 650,000 lines of Mojo code, right?
Yeah.
This is, this is a lot. And so we suffer, and we drive the features and the improvements based on our needs. We also im- appreciate the community, and we have a whole bunch of contributions coming in. And somebody just optimized my string, you know, that I, you know, to get, get rid of a bit out of my string implementation, which was suboptimal, and so that was super awesome.
But driving it that way makes sure it's real. It's grounded. It's on the use case. It's not... We're trying not to overpromise. Like, even when you're asking me what it's useful for earlier, right? I didn't say it's a replacement for Python.
I said it's, it's a go fast language. Someday it may be a pretty credible Python alternative, but for right now, it's good at a specific class of use cases. And if you're interested in those use cases, like making GPUs go brr- Mojo's awesome.
Like, but if you want a replacement for Rust end to end, then give us six months
Yeah. Yeah. I mean, uh, not... You're, you're a force of nature. I think there's a lot of mystery around, like, what is going on in, with Apple's AI initiatives.
Mm-hmm.
And I think the consumers suffer. Like, at, at the end of the day, like, the end users are, like, waiting for this, and it's not happening.
Uh, unfortunately, anything I know is massively out of date.
No, no, no. But what-
I mean, they've-
Yeah
... changed and reorged and grown and culture, and it's a very successful company. And so I think that they probably feel success, and, um, they're having trouble adapting to changes in the industry, and that's pretty typical of a lot of big companies.
Yeah.
And so I can't speak to the specific causes.
Uh, speaking of G- like, Google, I, obviously, I think they were, they were one of the earliest... Well, you know, what are we talking about? They invented transformers and a lot of other things.
Yeah.
TensorFlow, remember that?
Yeah, exactly.
Huge. Yeah. Um, NTPUs.
I credit Google with making AI open source. Well, they did not have to open source TensorFlow.
Yeah.
That was an incredible decision. Full kudo to Jeff Dean and, like, many of the other people that were involved in this because they said, "You know what's actually the most important thing for Google? Is for AI to go faster."
Yeah.
"How do we do that? We open source TensorFlow rather than making it some proprietary internal thing," which they had a previous system called Disbelief. And so that is a huge moment that set the stage for PyTorch to be open source and for the research to be open and for all of these things because they decided the value system was AI go faster.
The transformer paper being published, like, so many contributions from Google came from that.
Okay.
I, I don't think Google gets enough credit for that.
Well, yeah. Like, why is it better for Google for AI to be open source rather than, you know, Google owns it?
Uh, well, so I can't tell you if, like, the bet worked.
Yeah.
But I can tell you that that was a bet. But from my outsider now perspective, 'cause I haven't been at Google for over five years somehow, time flies, the bet makes sense when you have an amazing team of researchers and SWEs that can go incorporate this into your products.
And so Google does have billions of users. It has all the product services, it has all the different applications, and it has an incredible density of talent. And I think that, uh, Google's recent announcements, so just after Google IO and things like this-
Yeah, we were both there. Yeah
... yeah, the, it's like Google's actually working-
It's-
... I think. It's pretty impressive. And for a while they were dealing with organizational drama and Google Brain versus DeepMind and some of this stuff, and I, I can't speak to what they've done, but seems like they're a much more unified team.
They're executing well, they're getting research into product, and so feels like a different Google to me.
It... Yeah, it totally does. It used to be that, that there was just two of everything in Google, and you didn't know which one to use, and like-
Killed by Google
... they all-
Right. Yeah
... they all deprecated in a year. So like, yeah, I think they've, they've gotten the memo.
Yeah.
Whatever.
Yeah. And, uh, the other thing that's super impressive to me about them, and just me fanboying Google, right?
Sure. Yeah, yeah, yeah.
I mean, after railing-
Absolutely. Yeah
... the trillion dollar companies, but the, uh, the, uh, the things that they announce, they're actually shipping.
Yeah.
So much in AI is-
This is more Apple shade.
I was specifically saying Apple. This is, like, very common in AI is, like, here's, here's... And Modular's done this in the past too. This is why... So I, I gave this, uh, very deep tech talk. I sent you a link to the GP-
The GP mode. Yeah
... GP mode talk.
Yeah.
And the slide two was warning-
You can actually use this. Yeah
... this is not vaporware.
Yeah, yeah.
Everything here you can reproduce.
Yeah, yeah.
These claims you can download. This is actually real. Like, here's links to the source code, right?
I was wondering why you stressed that so much. I'm like, "Who, who hurt you?" You know? Like I know.
I mean, there's so many claims. Nobody knows what is real anymore.
Yeah.
Right? And I mean, there's lit- literally been product demos where, you know, it's like some electric semi rolled downhill instead of working under its own power-
Mm-hmm. Yeah
... and like, nobody knows what's real.
I knew they started to work. There's, like, a WhatsApp chat. We're, like, playing soccer at the Google field during the week. And about six months ago, some- the admin posted, it's like, "Hey, not enough people are showing up anymore to play soccer at lunch."
"What is going on?" And I think that's when-
Ah
... people started-
That's your Google indicator
... working again.
Yeah.
Yeah.
There you go. Um-
I mean, Sergey Brin was at IO-
Yeah
... and, like, he's, he's definitely, he's, like, working again, and it's, it's awesome.
Yeah. So I, I have mad respect for that, right? And so my values are aligned with people who ship stuff.
Yep.
'Cause that, that's what impacts the world.
Let's talk about open source a little more. There's the more recent maybe open source thing, which is DeepSeek, obviously.
Yeah.
And I think specifically in your case, you know, they worked at the PTX layer of the GPU, which is, like, even lower and more proprietary than CUDA.
Yeah.
I'm curious how, both in terms of, like, obviously the impact was, like, huge-
Yeah
... but maybe impact on how much people should actually try and move away from this proprietary thing. Because now the next-
Yeah
... from, from my understanding is, like, the next set of chips, all the code is, like, useless.
DeepSeek & Inference53:17
Well, so it's, it's not-
Basically
... widely known, but Blackwell is not compatible with Hopper.
Right.
Hopper kernels don't always-
Yeah
... run on Blackwell, for example, right? Um, but so your question is, like, what does it mean for the industry or what we learn or?
Well, it's like why, why is it so important?
Yeah.
Like, why the DeepSeek-
Yeah
... the DeepSeek example is, like, so important of, like, they need to navigate all this, like, proprietary stuff-
Yeah
... just to make it work.
So I'll give you my lived experience 'cause DeepSeek came out in December, which is when I, and probably you, noticed it, right?
Yep.
But then the world had a big wake-up call, and NVIDIA stock price went down and all that stuff like a month later, right? So here's my explanation of what happened, okay? What the DeepSeek team did was really impressive research.
They pushed MLA forward, like, they, which is a form of attention, and they pushed, uh, low-precision training forward. They pushed a whole bunch of stuff forward. They reverse engineered some PTX instructions that weren't well known at the time, right?
And so a lot of people were just like... And, and it was a Chinese team, right? Which put Americanism, it threatened Americanism, right? And things like this. And so what I found really exciting about it was they pushed the research forward, and they did these incredible things.
They showed the world that it was possible, and they opened it, and they published it.
Yeah.
And they actually taught the world about it because I don't know why they chose to do that, but it's because they believe in openness and-
Mm-hmm
... AI moving forward, right? And so the thing that I found striking is that the world's reaction to that was more striking to me than, uh, the actual models.
Ban or whatever. I don't know.
Yeah.
Mm-hmm.
Because, so first of all, there's the Chinese-American drama, which Geopolitics is not exactly my strong point, so I w- I, I get it, then I push that aside, right? But the other thing I found really interesting is that people said, "Wow, okay, only DeepSeek is able to go down to the PTX level."
But that is standard. That is what all of the leading teams do. In the case of Modular, we go literally, like, we only work at that level 'cause we get rid of all of CUDA, right? And so we've replaced the entire stack, and so we only do that.
But that's-- But a lot of teams that care about performance will actually go down and use the PTX Tensor Core instruction foo, which isn't really documented. Like, you kinda have to figure it out and look at Cutlass code.
Like, it's, it's really... NVIDIA doesn't make it as easy as they should to do this. Um-
Wonder why.
Well, I think it's 'cause they're, they're breaking Blackwell, and so they didn't want people to actually use it, and so they have their own issues, right? But good luck with that. There's a lot of smart people out in the world.
Ugh.
Um, and so to me, I thought it was really interesting because there wasn't a lot of awareness of how that level of the stack worked. For me, I thought it was a great wake-up call where people were like, "Oh, wow, if you work at this level of the stack," you know, a level of the stack I have inhabited for decades now, "you have power over compute.
You can understand and solve problems. You can drive research forward in ways that nobody else can do." And this is what the trillion-dollar companies do. It's not just DeepSeek. But I think DeepSeek was a huge wake-up call, and it drew attention to that layer of the stack because it wasn't just, like, throwing layers on top of PyTorch or s- you know, or VLLM, or it was, like, doing that fundamental work, and I thought that was incredible.
Now, the challenge with it, and the challenge with the way DeepSeek did it and the way that everybody else does all this work, is that it's completely specific to one GPU. It's not just that you're working at the PTX level, it's that you're writing code that really can only work on that one GPU, and that means that when Blackwell comes out, you have to throw it away and write a new one, right?
But here's the open secret. That's what VLLM is. Go look at VLLM. They have different kernels for A100 than H100. They're now trying to catch up with Blackwell, and so state-of-the-art for these systems since GenAI. Before that, there were fancy AI compilers like XLA and that kind of stuff in the trad AI world, uh, provided some scalability.
But in GenAI, it's a rewrite all the things when a new piece of hardware comes out. And so this is what Mojo is solving, right? And this is where, again, we can't turn the incremental cost of a new piece of hardware to zero, but we can massively reduce it.
And so this is a really exciting time, I think, uh, for us that we've demonstrated now, but also what it means for the future.
I still also wonder, they're hiring or internal training that they manage to have a small team that does this in the same way that you do.
I think they have a pretty significant team. I don't-- have no insider information, of course. But it's-
Yeah
... it's not like a five-person team. It's like-
Yeah, but-
... hundreds of people. Okay. Yeah, so it's-
Yeah. You know, it's not 1,000.
Mm-hmm.
Right? It's amazing, especially, like, I mean, I, presumably there's some language barrier, but e-even ignoring that, just getting that amount of t-talent density in, in one company is, is not a-
Yeah. Well, no
... significant task.
Well, I agree with that. Building a company is hard. I have no visibility on how they built DeepSeek.
Yeah.
But, um, all I can say is thank you to DeepSeek for publishing their work.
Yeah.
'Cause they didn't have to do that, and I think it left temporarily a lot of people flat-footed. Right? It, it was kind of embarrassing for certain groups, and I think a lot of people paused and were like, "Oh, crap, what do I do about this?"
Yeah.
But I think that what it did is it pulled forward progress in AI by, like, six months.
Yeah. They didn't just do GPU-level stuff. They also, like, had a file system-
Oh, yeah.
If you, like-
Yeah
... looked at, like-
That's incredible
... do you see... Obviously, that's not Modular's bread and butter, but, like, do you see a potential there for someone else to, I don't know, like, take that and run with it?
I've been very obsessed with the inference problem for a couple years.
Exactly, yeah.
And so I don't know the best way to solve that problem for training. And so-
Like, like Google internally has a great-
Indeed.
You know? Right.
Yeah.
And, and so, like, where is that for the rest of the world?
Yeah. I, I'm not-
Okay
... honestly just not the right person to answer that question-
All right
... 'cause I have my own obsessions.
Yeah. Uh, well, talking about inference, reasoning models and inference time compute-
Yeah
... very big topic. Does anything change or, or does nothing change because it's just more inference?
It depends on change from what, right? So we, uh, so when we started Modular over three years ago, we made the, like, ridiculously weird bet at the time to focus on inference instead of training.
Yeah.
And again, at the time, this-- I'm used to this. People are like, "What's wrong with you? Everybody obviously knows that training is the thing, and training, training, training, and people building these massive clusters and all this spend is on training, et cetera, et cetera, et cetera."
And I said, "Well, yeah, I, I understand why you see that, but I've lived this at Google."
Yeah.
You know, Google's like five years ahead of the rest of the industry in many ways.
The most scaled inference, yeah.
Yeah. Uh, training scales the size of your research team. Inference scales the size of your customer base. Right? And so the thing that happens between those is research gets into production. And so the gap is research getting production.
Once you do that, suddenly it scales like crazy. So I didn't plan for GenAI. I did not plan for inference time compute and things like this. But that bet on the production use case, 'cause it scales with the number of applications of AI, not just the hot research team, right, which, you know, is pretty important.
It's just not something I've focused on for the last few years. That was controversial. And so I think now the whole world's flipped, and I think the world gets it, right? But, but that was really because of lived experience.
And so when you come to these new techniques, right, another controversial thing, going back to the CPU thing, is why are you starting with CPUs? The answer at the time was pre-processing, post-processing, full system integration, networking. You need a CPU to feed the GPU, like very standard things.
But now you say KV caches.
Mm-hmm. Yeah.
Like your eviction policy runs on a CPU, like that radix hashing algorithm and block hashing and all that stuff happens, like, primarily CPU. That's really important for performance because if you have latency in these steps, like, you're not keeping your GPU utilized, right?
Again, this comes back to the rewrite in Rust and things like this.
Yeah.
All the agentic stuff and things like this, I didn't predict that, but I'm not surprised, and I think we're gonna see more.
Yeah.
And so tight integration, optimization across boundaries, like, these are things I believe in, and I think this is how you move the world forward.
It's amazing how, like, um, basically every... Mo- all the smart programmers I know always focus on where the bottleneck is.
Yeah.
And it's always, it always leads you to the right answer. Like, if you just are very clear-eyed about that.
And, uh, I don't know, it's, it's nice to see that happening. The other talking point I'll, I'll, uh, mention for you, 'cause you didn't, you didn't bring it up, but it- it's something that other people are talking about, is that actually now because of the need requirement for train of thought and reasoning models, inference is now part of training.
Okay. Yeah. Because you need to inference- RL has always done that ... to RL. Yeah. Yeah. Yeah. And, and so, like, it's, it's getting... That actually is a plus in favor of what you're doing anyway. Yeah, absolutely. Because it's the same code.
Well, and so, uh, so my, my experience with RL systems were, like, at scale, DeepMind-style AlphaGo and things like this, right? Yeah. So it's been a couple years ago, but setting those things up was incredibly complex, 'cause now you're dealing with cluster-scale orchestration, you're batching across all these agents.
Like, and, and again, none of the systems are set up for that, right? You've got PyTorch if you wanna train a model, but nothing is set up to do this. And so, again, I can't speak to all RL systems everywhere, but they ended up being like duct tape and bailing wire and super crazy- ...
stuff, and it was incredible. Again, the, it's incredible what a team of experts who knows the full stack and what they can achieve, but it shouldn't be that hard, right? So we're... Modular's not focused on solving that problem yet.
Um, maybe- But we have the- ... maybe we'll get there ... building blocks. We have the building blocks. And, and again, we're, we're not focused on solving training yet. Maybe we'll get there. I have some ideas on logical steps to do that, but I wanna make sure we're grounded and we solve things all the way end to end, we make people happy, people talk about our stuff in their, from their voice, it's not just me talking about our stuff.
Yeah. And so, so we're in that phase where you'll hear people talking about our stuff soon. I think people already are, but, like, they will even more, and I think I appreciate you coming on the, the podcast, uh, to talk about that more.
We wanna turn to some personal- Yeah ... stuff. Um, so I think that was a great Modular story. I think now on the personal side, there's a couple- Yeah ... things. So when you first joined us, I think I saw it was September 2023, so that was a year and a half- Wow.
Leader's Routine1:02:31
... into building the company. Yeah. Something like that. Now three years and a quarter. And what are, like, personal learnings, both from obviously being a leader in a company- Yeah ... having a growing team that you're managing, having a lot of responsibilities that are not technical- Yeah ...
anymore. Kind of run us through some of that. Yeah, so for me, um, I've built large teams from scratch before- Right ... but they've all been at established companies. Um, I've worked at a startup before, but it was somebody else's startup, and so this is the first startup that I've founded and then built a significant team and built a product that takes years to build.
And so a lot of the lessons I learned were super valuable and allowed, uh, me to achieve some of the stuff, but it's also very different, right? And so one of the things that's very different is it's very personal.
Mm. Right? So I've lost people from our team before, like many times. I've had to fire people and, and people have quit and gone, you know, Apple to Google or whatever, right? And but that was never personal in the same way it is at Modular, right?
And so this is something where the intellectual side of my brain knows, like, if somebody leaves, it makes sense. They had a life change. Like, I don't wanna get in the way of their family or the, you know, the whatever, right?
I mean, I intellectually know that, but on the other hand, it hurts a little bit. Mm. And so I think I'm getting better at handling some of that kind of stuff. The other thing is that when you're growing a team from, like, 0 to 100 people at a big company, you have all of the infrastructure, and the infrastructure's mature, right?
And so you're... Even if you grow to a team of 100 people, you're still tiny, you know, compared to an Apple or a Google or a company like this. Like, you're still tiny in proportion to the size of the overall scale, and so they've already got all the recruiting and all the other stuff, and all the legal and finance and all that kind of stuff going.
They've got the manager training, they got all this stuff going. And so in a startup, you're sometimes get in some hot messes where it's like, okay, well, we need to do a reorganization and things like this. And so that's been, um, good learnings.
I think we've scaled into that well. And, uh, another thing, coming back to the people telling you it's impossible, actually, here's probably the most important thing, is that I'm used to being told that things are impossible, and then I'm pretty bullheaded, and I have a formula.
And so, like, there's a path to success, and I can explain if you want, but the, uh... I'm used to the feeling when it's, you know, that year one and it's just a completely skunk works, let's, like, prove a thing phase.
I'm used to that, okay, cool, now we're telling people about it, but it's still not good enough, and people tell you all your stuff sucks. I'm used to that. Like, okay, well, you get into that window where all the stuff is almost there, people are sweating, it's really hard.
There's, like, all these, like, seemingly impossible things. It's a significant team, but the pieces don't line up yet. And so we went through some of that last fall, and people were like, "Oh my gosh, maybe it will never work," and you get this, like, anxiety.
Mm-hmm. And I think the thing that I didn't appreciate going through that, and this is also why I'm so excited to be in this phase, that actually impacts other people's, like, their thought processes, 'cause they haven't been through it.
And they, and they like, "Trust me, we'll get there." "Well, when? Like, exactly what?" Like, you get these, like, super analytical engineers who want to understand everything, and they're really expert in their part of the stack, and they don't really have- Yeah ...
that established trust with the other department and the other department and the other department, and they're all lining up. And so definitely some learnings from that. But- Yeah ... but this is where, again, you get out of that R&D phase, and you get into the execution phase, and it's like, okay, well, engineers are really good at taking an imperfect thing that works end to end and making it better and better and better and better and better.
And so it's, it just, like, feels fundamentally different at Modular now than it did in any of those previous phases. Yeah. What's your, uh, day-to-day like? You have a lot of meetings. Do you have just a few? Yeah, I, I have a lot of...
You're talking about my, my, my lived life. Um, a normal weekday for me is, um, I wake up about 7:15. My wife and I get the kids out the door, and she drives them to school at about 8:00.
I usually do a half hour to 45-minute walk with my dogs, get exercise, strenuous, up and down a hill, heart rate goes up, which is good. Listen to podcasts, for example, yours. So that's, you know, why, why I have time to actually follow exciting, cool things that other people are doing.
Get to work at 9:00, and so work 9:00 to 6:00 or 7:00 or something like that. And most of that is meetings, and so that's me trying to solve whatever the problems are of the day. Get home, dinner with the kids.
I insist on eating with my family. Hang out with them until bedtime, and then crush through for like two, three more hours until I pass out. And so... And then do that regularly.
The second work day.
And then on the weekend it's amazing, 'cause I get a lot of time to work. Not all day, 'cause I do things with kids and stuff, but I get a lot of time and there's no meetings.
Yeah.
And so that's amazing.
That's when stuff actually gets done.
Yeah, exactly.
I think the, the key here is some kind of strategic review time that you lock off. Because, like, I think people often say, you know, you're- there's, there's times when you're working in the business-
Yeah
... and then there's times you're working on the business, where you step out for a bit. Often for, for founders it's when you do a board meeting. But I, I wonder if there's anything that's very meaningful for you.
Do you have a coach? Do you have-
Yeah
... something like that?
Yeah, so I guess there's two people I really owe a lot to. Um, or two, two categories. Like, so one is my co-founder, Tim.
Yeah.
So Tim and I formed the company together. Uh, we walk every week, every Friday, catch up, and it's like that zoom out, and try not to make it tactical. And so make sure that we can bounce crazy ideas off, and I have a lot of crazy ideas, believe it or not.
He does, too. But then ground ourselves on execution and, and go. The other thing is this combination between my wife, who's a sounding board, and the executive team. And I, I'll admit, get a little bit crazy and wanna solve the industry problem, and the exec team pulls it back to, "Okay, next quarter, let's make sure we have-" "...
a clear plan. Let's make sure we can communicate. Let's decide what the actual priorities are. Okay, we can have three priorities MAX for the whole team-"
Mm-hmm.
"... not, not 50 or something." Right?
Otherwise they're not priorities.
Yeah, exactly.
Yeah.
Everything can't be the top priority, otherwise nothing is, right? And things like this. And then my wife, who keeps me sane and is, you know, an amazing l- life, life coach.
How much do you get your wife involved on, like, the actual work? Like, does she know-
Uh, zero. Zero
... like, you don't, uh... Yeah.
Yeah, she has her own thing going on. Uh, my wife runs the LVM Foundation, and so she's got a bunch of things going on, plus kids and everything else, too.
She's like, "I don't wanna hear about these kernels."
Well, but she's, she's great for helping me.
Yeah.
So I'm more IQ than EQ, and so working with humans is not... Is, it's an acquired skill, not a natural skill, and so I think this is something where once, uh, you know, I often end up at this place where some weird thing is happening.
"What the heck is going on?" It's like, "Chris, it's obvious. Like, they're saying this, but this is what they actually mean." And I'm like, "Oh, I never thought about that." You know?
You mentioned coding agents.
Yeah.
What do you guys use internally? What do you like?
Yeah, so, I mean, as of this recording, I personally use Cursor. And so Cursor is great for... I mean, it's the best thing I've seen, and I don't spend a lot of time dabbling with things, and so... But Cursor is, um...
So I write a lot of C++ and Mojo code. The key thing for working on Mojo code in AI coding tools is make sure you have a lot of code in the context window.
Yeah.
And so Cursor and these tools can index really well.
I was thinking you open source Mojo code. Actually, that's pretty good for training.
That's one of the reasons we did this.
Yeah.
Yeah. One of the many reasons we did this, and so this is what we saw at our hackathon, is people could go zero to hero with Mojo because you can just put, you just index this entire huge code base-
Yeah
... and it's, it's phenomenal, and then learning a new language is actually easy when the AI is doing a lot of the mechanical stuff for you.
Yeah.
And also it looks like Python, so you can read it. But it just, like, massively scales the on-ramp.
Mm.
And so for new language adoptions, AI is... Let me convince you, AI is cool. Um, and so-
I, I want to just go ahead and, like, actually, like, ask the data sets people at the big labs what they need-
Yeah
... uh, and then just, like, feed it-
I've asked them all to index
... feed it to them.
Yeah, just take the code, please. It's-
Well, no, no
... it's all good. It's-
Label it for them
... it's Apache 2. Like, just go. Yeah. Yeah, we're adding markdown files, so there's a cloud.md, and so we're doing, we're doing some of the basic stuff.
Yeah, yeah.
Um, people within the company are dabbling with cloud code and some of the stuff, and so I don't have personal experience with that. But, um, but for me, I found that it's mostly... It, it is very useful, but it's really about boilerplate.
And so it's not about inventing new algorithms and stuff like this, and it's probably 'cause I'm not building a React component-
Yeah
... to match this.
I mean, so the labs are touting that they are training or benchmarking their models for writing CUDA kernels, right?
Yeah.
I wonder if you see a noticeable performance improvement when you change model to model.
It's-
Uh, I don't know if you... You probably don't benchmark them.
I haven't, I haven't looked at that-
Yeah
... specific use case. Um, but, uh, people often ask me, "Hey, Chris, why are you building a new programming language when AI's gonna write all the code?" Similarly, if you look at a lot of this, like, let's generate a CUDA kernel, you start to ask and wonder, or at least I do a lot, what is the purpose of code?
What I've reflected on and the way I currently think about it, subject to change obviously, 'cause we're all learning, I had to take a step back and I say, "Well, code isn't really about telling the computer what to do.
Code is about humans being able to understand what code does." And so someday when we get AGI or ASI or something like this, maybe it can be completely opaque and I really don't have to know. We're not there yet, so...
And I don't know when that will happen. But in the meantime, like, I wanna be able to look at what the code actually does, and I live in a world of constraints. I need to know I have a product which has all these features.
If I go add another feature, what happens? Is it gonna hit my latency budget? Is it gonna crash, run out of memory? Is it gonna do, is it gonna cost too much? Like, I need to be able to reason about this.
And so to me, I look at a lot of these coding tools, you know, scaled beyond where they currently are now, but into the foreseeable future as saying, like, okay, well, it's like hiring another engineer onto your code base or onto your team.
And fundamentally, coding and s- software engineering is a team sport. And like, you have product managers, you have engineers, you have a lot of things, and if you automate all the engineering of code, maybe you get to product managers only and marketing only or something, theoretically.
But you still wanna reason about what the code does. And so in that, in that-
This is like a, yeah, interpretability argument.
Yeah. And so in that world, like, what is actually the most important thing? The most important thing is you can express everything the hardware can do, 'cause you don't wanna have some capability, some cost or some, some boundary you can't penetrate.
Mojo does that. The second is you want readable code that you can actually understand, right? And so you don't want assembly language or something like that. You want something high level and expressive and easy to understand. And so, like, this is where I think Mojo is really unique.
And then the AI coding tools I see as a straight value add for adoption. 'Cause I mean, I've already seen what people that have never touched any of this stuff can do, and it's just incredible. You put somebody that's intelligent, they know the use case, and now you put these tools in their hands, and they can do amazing things already.
And so I, I find it super empowering. And, and y- again, come back to me wanting to see humans being empowered with new technology and being able to upskill and be able to learn. You know, you're not a GPU programmer today, but, you know, tomorrow hopefully there'll be 10 times as many people programming GPUs.
I think that'll make the world a better place. And so I don't think we're going back to the world of CPUs.
Mm-hmm.
I think GPUs will only get more important, and if we can help people do that, I think it's great.
AI Coding & Projects1:13:24
You mentioned some of the research on inference and the-
Mm-hmm
... arXiv papers. How do you keep up with arXiv?
Yeah. Well, s- we're a Slack shop, and so we have a paper channel. It's, I think-
Oh
... I think there's, uh, I don't know, somewhere between, uh, three to 10 papers a day that come through and, and so we have amazing, smart people that do that. I also follow, uh, the Reddit communities and things like this.
I'm an old-school person that uses RSS still.
Mm-hmm.
And so if you have an RSS feed, then it's way easier for me to follow you.
Which reader?
And, and I use Feedly.
Feedly.
Yeah. But, um, arXiv has the ability to follow specific groups, and so I do that for various groups, and so that's, that's another technique.
Any notable papers come to mind that you wanna give a mention to, or, or p- authors, you know, that you're really watching, where any time they publish you're like, "I'm reading that one."
I can't give you a well-considered answer. Um-
Okay.
... just, just this morning, a person from Microsoft published a paper. Uh, forget his name offhand, but, uh, he just published a really cool paper about auto-generating, uh, flash attentions and doing blocker or something. It had some cute name with block in it.
And so anyways, I mean, there, there's a ton of things going on.
Okay, so that kind of... Yeah, I, you know, it's interesting to just see what you pay attention to. Last but not least, the question that everybody wants the answer to: Have you finished building the Lego robotics table for your kids-
Ah
... that you mentioned last time, and what's the... Do you have a new project that you're working on?
I've massively screwed about this. So we still use the Lego robotics table.
Nice.
So this is a big 4x8 sheet of plywood with a bunch of, uh, 2x3s around the edge. But a 4x8 sheet of plywood is pretty hard to work with, and so it breaks apart into 3D composable modular sections, and so it's super great and the kids are getting way better at programming and Legos and all this kinda stuff.
Um, gosh, what's my most recent project? I was just building swords in the shop with kids on a band saw. And so you take a piece of wood, you have a band saw-
Wow
... give it to a kid, and you say, "Don't cut your fingers off." Turns out that a band s-
Easy to do.
And, and I mean, I'm, I'm obviously joking a little bit.
Okay.
They get a lot of, they get a lot of oversight, but, uh, a band saw is actually a very safe tool. And so the reason for that is that a band saw, which if you, probably a lot of people have never seen a band saw.
You can do a Google search for it. You have two wheels, and then you have, uh, a blade that goes around the wheels, and you've got a table. And the cool thing about a band saw is that you can crank down the opening towards the blade so it's just, just as big as the piece of wood, and also it pulls the wood into the table.
And so as, because it pulls the wood into the table, there, the risk of, like, something called kickback, and there's a lot of other things like this, is very low. And so you can basically tell a kid, "Look, you see your fingers?
Keep them away from the blade." And there's no, like, sudden jerking or other things that if you, you know, use the wrong technique, you can get yourself into real trouble. And if you keep the guard all the way down, then you can't get, like, an arm in there or something like that.
Yeah.
But, I don't know.
Did you make your whole garage woodworking station?
Yeah. Yeah, I'm a cars outside kinda guy. So that's-
Oh
... again, my Thank, thank you to my wife for tolerating my odd behaviors. I mean, I love building things, right? And so this is fundamentally, when you talk about, you know, what makes me tick, is I love the joy of discovery, right?
And so whether it's building an amazing team or building a new table. You know, built, like, dining room table or things like this. Uh, you can look at my website. I'm not a very good-
It is
... web designer.
No, but-
I have some woodworking projects on there. But I love building software. I love building and solving the problems that come to this. I'm not super great at building things that are ro- mechanical. And so maybe I'd be good at building one chair, but I'm not gonna build eight chairs to go around a dining room table.
Right.
That would just drive me crazy. And so the discovery and the learning is what, what motivates me.
You can build the machine that builds the chairs.
There you go. That sounds great, right?
Yeah.
Spend 10 years building a thing-
Yeah
... so you got three weeks, crank it out. Yeah.
Yeah.
Um, yeah. Awesome.
Yeah. Any call to actions? Like, are you hiring? Like, uh-
Hiring & Wrap1:17:05
Yeah. We're, we're, we're hiring a small number of elite nerds. And so if you care about GPU programming, you care about AI models, you care about inference, you care about Kubernetes and cloud-scale stuff, please check us out. We expect to grow a lot more later this year.
The, uh, uh, other thing is that we have a ton of open source code. And so if you've heard about Mojo, but you looked at it a year ago, guess what? Everything's completely different now. And so if you are interested in a lot of these things, if you're interested in learning about GPUs, we have a ton of content that will teach you about GPU programming, GPU puzzles, and things like this.
People are now picking up Mojo and putting in a lot of, like, elite GPU, uh, came out today, and there's a whole bunch of other people that are taking the stuff and putting it out there. And I think it's just such an exciting time because I think that lots more people should be programming GPUs.
I think this is a huge opportunity for the industry. And of course, if you're enterprise and you're having trouble scaling your AI, you know, let us know. We can, we can help.
Awesome. Thank you so much for coming on here.
Yeah.
Inspiration as always.
Yeah. Well, thank you for having me.






