LALatent SpaceMar 6, 2024· 1:37:38

Open Source AI is AI we can Trust — with Soumith Chintala of Meta AI

Soumith Chintala, creator of PyTorch and engineering lead at Meta AI, argues that open source AI is essential for distributing opportunity and trust. He details PyTorch's complexity—1,000 operators needed for generality—and explains synthetic data as a vehicle for imparting symbolic knowledge where humans already have good symbolic models. He highlights a coordination problem in open source: feedback is lost because frontends like Ooba and Ollama lack feedback buttons, and proposes a centralized sinkhole to collect high-quality feedback. Beyond text, he is excited about robotics, where hardware remains a bottleneck, and Osmo's work to digitize smell, which he compares to images in the 1800s.

  1. 0:00Intro
  2. 4:25PyTorch Design
  3. 22:38Inference & Benchmarks
  4. 29:51Synthetic Data
  5. 42:00Meta AI
  6. 58:56Open Source AI
  7. 1:21:16Beyond Text

Powered by PodHood

Transcript

Intro0:00

Alessio0:00

Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host, Swyx, founder of Osmo AI.

Swyx0:10

Hey, and today we have in studio Soumith Chintala. Welcome.

Soumith Chintala0:12

Thanks for having me.

Swyx0:14

Uh, on one of your rare visits, uh, from New York, where you live.

Soumith Chintala0:17

Yeah.

Swyx0:18

Um, you were-- you got your start in computer vision, um, at NYU with, um, Yann LeCun. That was a very fortui-fortuitous start. I was actually listening to your interview on the Gradient podcast. So if people want to know more about, like, the history of Soumith, uh, the history of PyTorch, uh, they can go to that podcast.

We won't spend that much time there. Uh, but I just was marveling at your luck, or I, I don't know if it's your luck or your drive to find AI early and then find, like, the right quality, uh, mentor, because I guess Yann, Yann really sort of introduced you to that world.

Soumith Chintala0:52

You're talking about extrinsic success, right? Like, a lot of people just have drive to do things that they think is fun, and a lot of those things might or might not be extrinsically perceived as, as, like, good and successful.

I think I just happened to like something that is now, like, one of the coolest things in the world or whatever. But if I happen... You know, the, the first thing I tried to become was a, was a 3D, uh, VFX artist, and I was really interested in doing that, but I turned out to be very bad at it, so I ended up not doing that further.

But even if I was good at that, whatever, and I ended up going down that path, I probably would have been equally happy. It's just, like, maybe, like, the perception of, oh, is this person successful or not might be different.

But I think, like, after a baseline, like, your happiness is probably more correlated with your intrinsic stuff.

Swyx1:56

Yes. Um, I think Dan Pink has this book on Drive, uh, that, that I often refer to about the power of intrinsic motivation versus extrinsic and how long extrinsic lasts. It's, it's not very long at all.

Soumith Chintala2:07

Yeah. Yeah.

Swyx2:07

Um, but anyway, now you are, you know, an investor in Runway, so in a way you're working on VFX.

Soumith Chintala2:13

Yes. I mean, in a very convoluted way.

Swyx2:16

It, it, it reminds me of the, um, Ed Catmull. I don't know if you guys know, but, uh, you know, he actually tried to become an animator in his early years and failed or didn't get accepted by Disney and then went and created Pixar and then got bought by Disney and created Toy Story.

Soumith Chintala2:30

Yeah.

Swyx2:31

Um, so you joined Facebook in twenty fourteen, um, and, uh, eventually became, uh, creator and maintainer of PyTorch.

Soumith Chintala2:37

Yep.

Swyx2:37

Um, and, uh, and there's, there's a long story there. You can refer to on the Gradient. Uh, but you also, like, um, you-- I think maybe people don't know that you're also involved in more sort of hardware and cluster decisions at FAIR, um, and we can dive into more details there because-

Soumith Chintala2:49

Yeah

Swyx2:49

... we're all, all about hardware this month. Um, and, uh, yeah, yeah, and then finally, I, I don't know what else. Like, what else should people know about you on, on a personal side or professional side?

Soumith Chintala2:58

I, I think, uh, open source is definitely, like, a big passion of mine, probably forms a little bit of my identity at this point. Um, I am irrationally interested in open source, right? It's like one of those things that I attribute to, um...

I think open source has that fundamental, uh, way to distribute opportunity in a way that, uh, is very powerful. Like, I grew up in India. Um, I didn't have internet for, for a while. Uh, and in college, actually, I didn't have internet, uh, except for, like, GPRS or whatever.

Um, so just having op... Like, and, like, knowledge w- knowledge was very centralized, but, like, I saw that evolution of knowledge slowly getting decentralized, and that ended up helping me learn quicker and faster for, like, zero dollars. And I think that was a strong reason why I ended up where I am.

So, like, that, like, the open source side of things, I always push regardless of, like, what I get paid for. Like, I think I would do that as a passion project on the side.

Swyx4:14

Yeah. That's wonderful. And we, we, we'll talk about the challenges as well that open source has.

Soumith Chintala4:17

Yeah.

Swyx4:18

Uh, uh, mo- uh, open models versus closed models. Uh, but maybe you wanna talk a-

Soumith Chintala4:22

Yeah

Swyx4:22

... little, touch a little bit on PyTorch before we move on to sort of Meta AI in general.

PyTorch Design4:25

Soumith Chintala4:25

Mm-hmm.

Alessio4:26

Yeah. We kind of touched on PyTorch in a lot of episodes. So we had George Hotz from TinyGrad. Um, he called PyTorch a, a CISC and TinyGrad a, a RISC. I would love to get your thoughts on PyTorch design direction as far as, um, I, I know you talk a lot about kind of having a, a happy path to start with and then making complexity hidden away-

Soumith Chintala4:48

Yeah

Alessio4:48

... but then available to the, to the end user. One, one of the things that George mentioned is I think you have, like, two hundred fifty primitive operators in, in PyTorch. I think TinyGrad has four. So, um-

Soumith Chintala4:58

Yeah

Alessio4:59

... how, h-how do you think about some of the learnings that maybe he's gonna run into that you already had in the past seven, eight years almost-

Soumith Chintala5:07

Yeah

Alessio5:07

... of, of running PyTorch?

Soumith Chintala5:09

Yeah. I think, uh, everyone starts... Uh, there, there's different models here, but, like, I think it's two mo- two different models that people generally start with. Either they go, like, "I have a grand vision, and I'm gonna build, like, a giant system that achieves this grand vision, and maybe one is, or, like, you know, super, like, complex, feature complete," whatever.

Or other people say they will get incrementally ambitious, right? And they say, "Oh, we'll start with something simple, and then we'll slowly layer out complexity in a way that, um, optimally applies Huffman coding or whatever."

Swyx5:42

Mm-hmm.

Soumith Chintala5:42

Like, you know, where the density of, uh, users are and what they're using, I would want to keep it in the, like, easy, happy path, and where the more niche advanced use cases, I'll still want people to try them, but they need to take additional frictional steps.

Um- George, I think, uh, just like we started with PyTorch, George started with a, like the incrementally ambitious thing. I remember, um, TinyGrad used to be like we would be limited to 1,000 lines of code, and I think-

Alessio6:17

Mm-hmm

Soumith Chintala6:17

... now it's like 5,000. So I think there is no real, like, magic with- to which why PyTorch has a kind of complexity. I think it's, like, probably partly necessitated and partly, uh, because we built with the technology available under us at that time.

If you had to rewrite Py- PyTorch is like 190,000 lines of code-

Alessio6:41

Mm

Soumith Chintala6:41

... or something at this point. I think if you had to rewrite it, we would probably think about ways to rewrite it in, like, a vastly simplified way, for sure. But a lot of that complexity comes from the fact that the ha- like, um, in a very simple explainable way, you have memory hierarchies.

Uh, you have, CPU has, like, three levels of caches, and then you have DRAM and SSD, and then you have network. Um, similarly, GPU has several-

Alessio7:16

Mm

Soumith Chintala7:16

... levels of memory, and then you have, like, different levels of network hierarchies, NVLink plus, like, uh, InfiniBand or RoCE or something like that, right? And the way the flops are available on your hardware, they are available in a certain way, and your computation is in a certain way, and you have to retrofit your computation onto both the memory hierarchy and, like, the flops available.

When you're doing this, uh, it is actually, like, a fairly hard mathematical problem to, uh, do this, do this, um, setup, like you find the optimal thing. And finding the optimal thing is like, what is optimal? The optim- what is optimal depends on, like, the input variables themselves.

So like, okay, like, what is the shape of your input tensors and like, um, what is the operation you're trying to do and like, uh, various things like that. Finding that optimal configuration, uh, like, and writing it out in code, um, is not the same for every, every op- every input, uh, configuration you have.

Like b- like for example, just as the shape of the tensors change, let's say you have three input tensors into like a, a, um, a sparse dot product or something like that. The, the shape of each of these input tensors will vastly change how you do this optimally placing, uh, th- this operation onto the hardware in a way that will get you maximal throughput.

So, uh, a lot of our complexity comes from writing out like hundreds of configurations, uh, for each single PyTorch operator, uh, a- and templatizing these things and like symbolically, like, like generating the, the final CUDA code or like CPU code.

Um, there's no way to avoid it because mathematically we haven't found symbolic ways to do this, uh, that also keep compile time near zero. Um, you can write a very simple framework, uh, but then you also should be willing to eat the long compile times-

Alessio9:35

Mm-hmm

Soumith Chintala9:35

... of like searching for that optimal performance at runtime. But that's the, the trade-off.

Alessio9:40

Yeah.

Soumith Chintala9:40

There's no, like... I don't think, like, unless we have, like, great breakthroughs, like George's vision is achievable. Like, or like he should be thinking about a narrower problem, such as, "I'm only gonna make this for like, work for self-driving car ConNets," uh, or like, "I'm only gonna make this work for like LLM transformers of the Llama style."

Like, if you start narrowing the problem down, you can make a vastly simpler framework.

Alessio10:08

Mm-hmm.

Soumith Chintala10:08

But, uh, if you don't, if you need the generality to power all of the AI research that is happening and keep like zero compile time and, you know, all these other factors, I think it's, it's not easy to avoid the-

Alessio10:21

Yeah

Soumith Chintala10:21

... complexity.

Alessio10:23

Um, that's interesting, and we kind of touched on this with, uh, with Chris Lattner when he was on the podcast. If you think about frameworks, they have the model target, they have the hardware target, they have different, different things to think about.

He mentioned when he was at Google, TensorFlow, it's trying to be optimized to make TPUs go brrr, you know-

Soumith Chintala10:41

Yeah

Alessio10:41

... a- and go as fast. Um, I think George is trying to make especially AMD stack be better than-

Soumith Chintala10:47

Yeah

Alessio10:47

... than ROCm. How come PyTorch has been such a Switzerland versus just making- ... Meta hardware go, go brrr?

Soumith Chintala10:54

First, Meta is not in the business of selling hardware. Meta is not in the business of, uh, cloud compute. Um, we kind of, uh... The way Meta thinks about funding PyTorch is it's just like we're funding it because it's net good for Meta to fund PyTorch because PyTorch has become a standard and, and a big open source project.

And generally, it, uh, gives us a timeline edge, it gives us like various, like, leverage, like, and all that to, w- uh, like within our own work. Um, so why is PyTorch like more of a Switzerland rather than being opinionated?

Like, I think the way we think about it is not in terms of Switzerland or not. We actually, the way, like, we articulate it to all hardware vendors and software vendors and all who come to us being like, "We want to build a back end in core for PyTorch and ship it by default," is like, we just only look at our user side of things.

Like, if users are using a particular piece of hardware, then we want to support it.

Alessio12:05

Mm.

Soumith Chintala12:06

We very much Don't want to king make the hardware side of things. Um, so f- as like the MacBooks, uh, have GPUs and as that stuff started getting increasingly interesting, we pushed Apple to push some engineers and work on the MPS support, and we spent significant time from, like, Meta-funded engineers on that as well.

Because a lot of people are using the Apple GPUs, and there's demand. So, like, we kind of mostly look at it from the demand side.

Swyx12:38

Mm-hmm.

Soumith Chintala12:38

We never look at it from, like, "Oh, which hardware should we start t-taking opinions on?"

Swyx12:45

I-Is, is there a future in which, uh, because Py- um, Mojo or Modulus, Mojo is kind of a superset of Python, is there a future in which, um, PyTorch might use Mojo features optionally?

Soumith Chintala12:56

I think it depends on how well integrated it is-

Swyx12:59

Uh-huh

Soumith Chintala12:59

... uh, into the Python eco-ecosystem. So if Mojo is like a pip install and, uh, it's readily available and users feel like they can use Mojo so smoothly within their workflows, within, uh, you know, in a way that just is low friction, um, we would definitely look into that.

Like, in the same way, like, PyTorch now depends on Triton, like OpenAI Triton, and we weren't, we didn't-- we never had a conversation that was like, "Huh, that's like a dependency. Should we just build a Triton of our own or should we, like, use Triton?"

Like, it, it almost doesn't... Like, those conversations don't really come up for us. It's, the conversations are more like, "Well, does Triton have like 10,000 dependencies, and is it hard to install?" Like, we almost don't look at these things from, like, a strategic leverage point of view.

We look at these things from, like, a user experience point of view. Like, is it easy to install? Is it-

Swyx13:59

Yeah

Soumith Chintala13:59

... like, smoothly integrated? If so, we should consider... And does it give enough benefits for us to, like, start depending on it? If so, yeah, we should consider it. That's how we think about it.

Swyx14:08

You're, you're inclusive by default if, as long as it meets, like, the, the minimum bar of-

Soumith Chintala14:11

Yeah.

Swyx14:12

Yeah. Um, but like, uh, may-maybe I phrased it wrongly. Maybe it's more like, okay, like, what problems would you look to solve-

Soumith Chintala14:17

Okay

Swyx14:18

... right, um, that, that you have right now?

Soumith Chintala14:20

I think it depends on what problems Mojo will be useful at. Um-

Swyx14:26

It's more a performance, mainly a performance pitch, uh, some amount of cross-compiling pitch.

Soumith Chintala14:31

Yeah, I think, like, the performance pitch for Mojo was like, "We're gonna be per-per-performant, uh, even if you have, like, a lot of custom stuff. Like, you can write arbitrary custom things and, like, they will be performant." And that value proposition is not clear to us from the PyTorch side to consider it for PyTorch.

So PyTorch exposes, like, it's actually not 250 operators, like 1,000 operators. PyTorch exposes about 1,000 operators, and people kind of write their ideas in the 1,000 operators of PyTorch. Um, Mojo's like, "Well, maybe, like, it's okay to completely sidestep those, like, 1,000 operators of PyTorch and just write it in a more natural form.

Just write, like, raw Python, like write for loops or whatever," right? So from the consideration of how do we intersect PyTorch with Mojo, like, I can see one use case where you're like, you have custom stuff, uh, for some parts of your program, but mostly it's PyTorch, and so we can probably figure out how to, like, make it easier for, say, torch.compile to, like, smoothly also consume Mojo subgraphs.

Uh, and like, you know, the, the interoperability being actually usable, that I think is valuable. But, like, Mojo as a fundamental front-end would be replacing PyTorch, not like augmenting-

Swyx16:06

Yeah

Soumith Chintala16:06

... PyTorch. So in that sense, I don't see a synergy in more deeply, like, integrating-

Swyx16:12

Okay

Soumith Chintala16:12

... Mojo.

Swyx16:13

So call out to Mojo whenever they have written something in Mojo and, uh, there's some performance-

Soumith Chintala16:18

Yeah

Swyx16:19

... related thing going on. Uh, and then since you mentioned Map- uh, Apple, um, what do you, what should people think of PyTorch versus MLX?

Soumith Chintala16:27

I mean, MLX is early, and I know the, the folks well. Ani, uh, used to work, uh, at FAIR, and I chatted, you know, I used to chat with him all the time. He used to be based out of New York as well.

Um, the way I think about MLX is that MLX is specialized for Apple right now. Um, it has a happy path because it's like in a, it, it, it's defined its product in a narrow way. At some point, MLX either says, "We will only be supporting Apple, and we will just focus on enabling..."

You know, this is a framework if you use your MacBook, but once you, like, go server-side or whatever, that's not my problem and I don't care. Um, or MLX, uh, it enters like the server-side, uh, side of things as well.

Like, it, one of these two things will happen, right?

Swyx17:26

Mm.

Soumith Chintala17:26

If the first thing will happen, like MLX's overall addressable market will be small, but it probably do well within that addressable market. Uh, if it enters the second phase, they're gonna run into all the same complexities that, uh, we have to deal with.

They will not have any magic wand, and they will have vastly more complex work to do. They probably wouldn't be able to move as fast in certain ways.

Swyx17:53

Like having to deal with distributed compute.

Soumith Chintala17:55

Distributed, uh, NVIDIA and AMD GPUs, like, just like having a generalization of a, the concept of a back end, how they treat compilation, uh, with plus overheads. Right now, they deeply assume, like, the whole MPS graph thing. Um- So they need to think about all these additional things if they end up expanding to, uh, onto the server side, and they'll probably build something like PyTorch as well, right?

Like, you know, eventually that's where it will land. And I think there they will kind of fail on the, like, lack of differentiation.

Swyx18:31

Yeah.

Soumith Chintala18:31

Like, the-- it wouldn't be obvious to people why they would want to use it. Um-

Swyx18:36

Yeah. I mean, there are some cloud companies offering M1 and M2 chips on, on servers. Um, I feel like it might be interesting for Apple to pursue that market, but it's not their core-

Soumith Chintala18:46

Yeah.

Swyx18:46

-strength.

Soumith Chintala18:46

I mean, if Apple can figure out their interconnect story, maybe, like, then it-

Swyx18:51

Wow.

Soumith Chintala18:51

-it can become a thing. Yeah.

Swyx18:53

Honestly, that's more interesting than the cars.

Soumith Chintala18:54

Yes. I think, like... I mean, the moat that Nvidia has right now, I feel like, is that they're, they have the interconnect that no one else has. Like, AMD GPUs are pretty good. Um, I'm sure there's various silicon that is not bad at all, but, like, the interconnect, um, like NVLink, is uniquely awesome.

So I'm sh- like, and, like, I'm sure the other hardware providers are working on it, but-

Swyx19:22

I feel like when you say it's uniquely awesome, you have some appreciation of it that the rest of us don't. Um, the rest-- I mean, the rest of us just, like, you know, we hear marketing lines, but what, what do you mean when you say, um, Nvidia is very good at networking?

Uh, obviously-

Soumith Chintala19:33

Um-

Swyx19:33

-they made the acquisition maybe, like, 15 years ago.

Soumith Chintala19:34

Just, like, the bandwidth it offers-

Swyx19:37

Okay

Soumith Chintala19:37

...and the latency it offers. I mean, like, TPUs also have a good interconnect, but you can't buy them, so you have to go to Google to use it.

Swyx19:46

Got it.

Alessio19:47

Who are some of the other fair PyTorch, uh, alumni that are building cool companies? I know you have Fireworks AI, Lightning AI, uh, Lepton.

Soumith Chintala19:55

Lepton.

Alessio19:56

Y-Yancheng you knew since college when he was building Caffe, um-

Soumith Chintala20:00

Yeah. So Yancheng and I used to be framework rivals, like-

Alessio20:04

Oh, yeah.

Soumith Chintala20:05

...Caffe, Torch. Um, I mean, we, we were all a very small, close-knit community back then. Um, Caffe, Caffe, Torch, Tiano, uh, Chainer, um, Keras, um, various, various frameworks. I mean, it used to be, like, more like 20 frameworks.

I can't remember all the names. CCV by Liu Liu, who is also based out of SF. Um, and I would actually, like... You know, one of the ways, uh, it was interesting is, like, you went into the framework guts and saw if someone wrote their own convolution kernel or they, like, were just copying someone else's.

And there were, like, four or five convolution kernels that were, like, unique and interesting. There was one from this, from this guy out of Russia. I, I forgot, uh, the name. But, like, I remembered who was awesome enough to have, like, written their own con kernel.

Um, and at some point there, like, I, um, I built out these, these benchmarks called ConNet benchmarks-

Alessio21:15

Yep

Soumith Chintala21:16

...that, um, they were just benchmarking all, uh, all the convolution kernels that were available at that time. Uh, and it hilariously became big enough that at that time, AI was getting, like, important, but not important enough that industrial strength players came in to do these kind of benchmarking and standardization like we have MLPerf today.

So a lot of the startups, um, were using ConNet benchmarks in their pitch decks- ...as, like, "Oh, you know, on ConNet benchmarks, like, this is v- how we fare, so you should fund us." I remember Nirvana actually was at the top of the pack because Scott Gray wrote, like, amazingly fast con- convolution kernels at that time.

Um, very interesting, but separate times. But to answer your question, Alessio, um, I think mainly Lepton, Fireworks are the two most obvious ones. Um, but I'm sure the, the, the fingerprints are a lot wider. Um, they're just people who worked within the PyTorch Caffe 2 cohort of things and now end up at various other places.

Inference & Benchmarks22:38

Alessio22:38

Yeah. Um-

Swyx22:39

I, I think as a, um, both as an investor and a, um, people looking to build on top of their services, um, it's a, um, uncomfortable slash, like, I don't know what I don't know pitch, uh, when-- 'cause I've met Yancheng, and I've met, uh-

Soumith Chintala22:56

Lin Chao.

Swyx22:57

Lin, yeah. I've, um, you know, I've met, I've met these folks, and they're like, "You know, we, we're deep in the Py- on the, um, PyTorch ecosystem, and we serve billions of inferences a day or whatever at, at, at Facebook, and now we can do it for you."

And I'm like, "Okay, that's, that's, like, great. Like, what, what should I be wary of or cautious of when, when, when these things happen?" Because I'm like, obviously this, this experience is extremely powerful and, and, um, and valuable.

I just don't know what I don't know. Like, what, what should people know about, like, these sort of new, um, inference-as-a-service companies?

Soumith Chintala23:28

At that point, you would be in-investing in them for their expertise, uh, of one kind. So if they, if they've been at a large company, uh, but they've been doing amazing work, you would be thinking about it as like, okay, like what these people bring to the table is that they're really good at, like, GPU programming or understanding the complexity of serving models once, once it hits a certain scale.

Like, you know, various expertise, like from the infra and, like, AI and GPUs point of view. What you would obviously want to figure out is, like, whether their understanding of the external markets is clear, whether they know and understand how to think about Running a business, like understanding how to be disciplined about making money or, you know, various things like that

Swyx24:25

Oh, maybe I'll, I'll put it like it's... I actually, I will de-emphasize the investing bit-

Soumith Chintala24:29

Yeah

Swyx24:29

... and just more as a potential customer.

Soumith Chintala24:31

Oh, okay.

Swyx24:32

Like, it's more like, okay, like, you know, you're PyTorch god- gods, of course. Like, what, what, what should I, what else should I know?

Soumith Chintala24:39

I mean, I would not care about who's building something if I'm trying to be a customer. I would care about whether-

Swyx24:46

The benchmark

Soumith Chintala24:47

Yeah, I'm-

Swyx24:47

Okay

Soumith Chintala24:47

... I, I use it, and it's like usability and-

Swyx24:51

Speed

Soumith Chintala24:51

... reliability and speed, right?

Swyx24:53

The quality as well.

Soumith Chintala24:54

Yeah, I... If someone from some random unknown place came to me and say, "U- use our stuff, it's great," like I sh- and I have the bandwidth, I probably will give it a shot, and if it turns out to be great, like I'll just use it.

Swyx25:10

Okay, great. Um, and, and then maybe one, one more thing about benchmarks, since we already brought it up, uh, and you brought up Coffeenet benchmarks. Um, there was recent, some recent drama around Anyscale. Um, the Anyscale released their own benchmarks, and obviously they look great on their own benchmarks, but, um, maybe didn't give the other, uh...

I, I feel, I feel like there are two lines of criticism. One, which is they, they didn't test sort of apples for apples on the kind of endpoints that, uh, the other providers that they are competitors with, um, you know, on their benchmarks and, you know, that is due diligence baseline.

And then the second would be more just like optimizing for the right thing. Um, you had some commentary on it. I'll just kind of let you riff.

Soumith Chintala25:48

Yeah, I mean, in summary, basically my criticism of that was Anyscale built these benchmarks for end users to just understand what they should pick, right? And that's a very good thing to do. I think what they didn't do a good job of is give that end user an underst- a full understanding of what they should pick.

Like, they just gave them like a very narrow slice of understanding. I think they just gave them, um, latency numbers and, um, that's not sufficient, right? Like you, you need to understand your total cost of ownership at some reasonable scale, not like, oh, like one API call is like one cent, but like 1,000 API calls are like 10 cents or like, you know, like you don't...

Like people can misprice to cheat on those benchmarks. So you want to understand, okay, like how much is it gonna cost me if I actually subscribe to you and do like a million API calls a month or something?

And then you want to understand, uh, the, the latency, uh, and reliability, not just from like one call you made, but like an aggregate of calls you made over several, various times of the day and times of the week.

Swyx27:05

Yeah.

Soumith Chintala27:06

And, uh, the nature of the workloads. Like it's like, is it just like some generic single paragraph that you're sending that is cacheable or like, is it like-

Swyx27:15

Mm-hmm

Soumith Chintala27:15

... testing of real work, right, real world workload? I think that kind of rigor, like in presenting that benchmark wasn't there. It was mu- much more narrow sliver of what should have been a, a good benchmark. That was my main criticism, and I'm pretty sure if before they released it, they, uh, showed it to their like other stakeholders who would be caring about this benchmark because they are present in it, they would have easily just pointed out these gaps.

Swyx27:46

Yeah.

Soumith Chintala27:46

And I think they didn't do that, and they just like-

Swyx27:49

Yeah

Soumith Chintala27:49

... released it. So I think those were the two main criticisms, and I think they were fair, and Robert took it well and he-

Swyx27:55

He, he took it very well.

Soumith Chintala27:56

Yeah.

Swyx27:56

Yeah. And we'll, we'll have him on at some point, and-

Soumith Chintala27:58

Yeah

Swyx27:58

... we'll, uh, discuss it. But I think it's important for, I think the market being maturing enough that people start caring and competing on these kinds of things-

Soumith Chintala28:05

Yeah

Swyx28:05

... means that we need to establish what best practice is.

Soumith Chintala28:08

Right.

Swyx28:08

Because otherwise everyone's gonna play dirty.

Soumith Chintala28:10

Yeah, absolutely. Uh, my view of the LLM inference market in general is that it's, it's like the laundromat, uh, model. Like you're, the margins are gonna drive down towards the bare minimum. Like it's gonna be all kinds of arbitrage between how much you can get the hardware for and then how much you sell the API, and how much like latency your customers are willing to let go.

Like you need to figure out how to squeeze your margins. Like what is your unique thing here? Like, uh, I think like Together and Fireworks and all these people are trying to build some faster CUDA kernels and faster like, you know, hardware kern- kernels in general.

But those moats only last for a month or two. Like these ideas quickly propagate, um-

Swyx28:55

Even if they're not published?

Soumith Chintala28:58

Even if they're not published, like the idea space is small.

Swyx29:03

Okay.

Soumith Chintala29:04

So even if they're not published, s- the discovery rate is gonna be pretty high. It's not like we're talking about a combinatorial thing that is really large. You're talking about like Llama style LLM models, and we're gonna beat those to death, like on like- ...

a few different hardware SKUs, right? Like it's not even like we have a huge diversity of hardware you're going to aim to run it on. Now, when you have such a narrow problem and you have a lot of people working on it, like the rate at which these ideas are gonna get figured out is gonna be pretty rapid.

Swyx29:37

Is it like a standard bag of tricks? Like the, the standard one that I know of is, you know, fusing operators and-

Soumith Chintala29:42

Yeah, it's the standard bag of tricks on like figuring out how to like improve your memory-

Swyx29:48

Yeah

Soumith Chintala29:48

... bandwidth and all that. Yeah.

Swyx29:50

Okay. Interesting.

Alessio29:51

Um, any ideas instead of things that are not being beaten to death- ... that people should be paying more attention to?

Synthetic Data29:51

Swyx29:57

One thing I was like, you know, you have 1,000 operators, right? Like w- what's the most interesting use- usage of PyTorch that you're, that you're seeing maybe outside of this little bubble?

Soumith Chintala30:05

So PyTorch, uh, it's very interesting and scary at the same time, but basically it's used in a lot of exotic ways, like from the ML angle. Like, okay, like what kind of models are being built? And you get all the way from like state space model and all these things to like Stuff that like Nth order differentiable models like, like neural, neural ODs and stuff like that.

Um, I think like there's one set of interestingness factor from like the, the ML, uh, side of things, and then there's the, the other set of interesting factor from the applications point of view. It's used in Mars rover simulations to drug discovery to Tesla cars, and there's a huge diversity of like applications which it, it, in which it is used in.

So in terms of the most in- like, I think like in terms of most interesting application side of things, I think

I am scared at how many interesting things that are also very critical and really important it is used in. Um, if, if you're, if like if... I think the scariest was when I went to, uh, visit CERN at some point, and they said they were using it, PyTorch, and they were using GANs at the same time for like particle physics research.

And I was scared more about the fact that they were using GANs than they were using PyTorch, because at that time I was like a researcher focusing on GANs. The diversity is probably the most interesting. The-- how many different things it is being used in, I think that's the most interesting to, to me from the applications perspective.

From the models perspective, I, I think I've seen a lot of them. Like the, the, the really interesting ones to me are where we're starting to combine search, uh, and symbolic stuff with, with differentiable models. Um, I, I think, uh, like the whole AlphaGo style model so-

Swyx32:24

Mm

Soumith Chintala32:24

... is one example, and then I think we're attempting to do it for LLMs as well with like various reward model and then search. I mean, I don't think PyTorch is being used, uh, in this, but like the whole alpha geometry thing was interesting because again, it's an example of combining the symbolic models with, with, uh, the gradient-based ones.

Uh, but there are stuff like alpha geometry that PyTorch is used at, um, especially when you intersect biology and chemistry with, with, um, ML. Like in those areas, you, you want stronger guarantees on the output. Um, so, so yeah, maybe from the ML side, those things to me are very interesting right now.

Swyx33:09

Yeah. People are very excited about the alpha geometry thing, and it's kind of like, uh, for me, it's theoretical. It's, you know, it's great. You can solve some Olympiad questions. I'm not sure how to make that bridge over into the real world applications, but I'm sure-

Soumith Chintala33:22

Well, okay

Swyx33:22

... people are smart enough, they will figure it out.

Soumith Chintala33:23

Let me give you an example of it. You know how like the whole thing about synthetic data will be-

Swyx33:30

Yes

Soumith Chintala33:30

... the next rage in LLMs is a thing?

Swyx33:32

It already is a rage.

Soumith Chintala33:34

Which I think is, uh, fairly misplaced in how people perceive it. People think synthetic data is some kind of magic wand that you wave and it's gonna be amazing. Synthetic data is useful in neural networks right now because we as humans have figured out a bunch of, uh, symbolic models of the world or made up certain symbolic models because of human innate biases.

Uh, so we've figured out how to ground particle physics in a 30 parameter model, and it's just very hard and, uh, to compute as in like, it's like it takes a lot of flops to compute, but like it only has 30 parameters or so.

I mean, I, I'm not a physics expert, but like it's a very low-

Swyx34:28

Mm-hmm

Soumith Chintala34:28

... rank model. We built mathematics as a, as a field that basically is very low rank. Um, language, like un- a deep understanding of language, like the whole syntactic parse trees and like just understanding how language can be broken down in, into a formal symbolism is something that, that we figured out.

So we basically, as humans, have accumulated all this knowledge on these subjects, either synthetic-- I mean, we created those subjects, you know, in our heads or like we've grounded some real world phenomenon into a set of symbols. But we haven't figured out how to teach neural networks symbolic world models directly.

The only way we have to teach them is generating a bunch of inputs and outputs and gradient descending over them.

Swyx35:19

Mm-hmm.

Soumith Chintala35:20

So in areas where we have the symbolic models, but we, we like, you know, and we need to teach all like the, the, the, the knowledge we have that is better encoded in the symbolic models, what we're doing is we're generating a bunch of synthetic data of a bunch of input-output pairs, and then giving that to the neural network and asking it to learn the same thing that we already have a better-

Swyx35:45

Mm-hmm

Soumith Chintala35:45

... low rank model of in gradient descent in a much more overparameterized way. Outside of this, like where we don't have good symbolic models, like synthetic data obviously like doesn't make any sense. So synthetic data is not a magic wand where it'll work in all cases, in every case, you know, whatever.

It's just where- We as humans already have good symbolic models of. We can-- We need to give, impart that knowledge to neural networks, and we figured out the synthetic data is a vehicle to, like, impart this knowledge to.

So but people, becau- uh, because maybe they, they don't know enough about synthetic data as a notion, but, like, they hear, like, "You know, the next wave of data revolution is synthetic data." They think it's k- some kind of magic where we just, just create a bunch of random data somehow.

They don't think about how, and then they think, like, that's just the revolution, and I think that's maybe, like, a gap in understanding most people have in this hype cycle.

Swyx36:46

Yeah. Well, it's a relatively new concept, so.

Soumith Chintala36:48

Yeah.

Swyx36:49

Oh, there's two more that I- I'll push-- I'll, I'll put in front of you, and then you can see if you-- see what you respond. Um, one is, um, you know, I have this joke that it's, uh, you know, it's only synthetic data if it's from the Mistral region of France.

Otherwise, just a sparkling distillation, which is, uh, which is what news research is doing. Like, they're distilling GPT-4 by creating synthetic data from GPT-4, like creating mock textbooks inspired-

Soumith Chintala37:11

Right

Swyx37:11

... by Phi-2, and then, uh, fine-tuning open source models like, like Llama.

Soumith Chintala37:15

Yeah.

Swyx37:16

Um, and so should we call that synthetic data? Should we call it something else? I don't know. But it's, you know...

Soumith Chintala37:20

Yeah. I mean, the outputs of LLMs, are they synthetic data? They probably are, but I think it depends on the goal you have. If your, if your goal is, like, you're creating synthetic data by-- with the goal of trying to distill GPT-4's superiority into another model, I guess you can call it synthetic data.

But it also feels, like, disingenuous because your goal is, like, I need to, like, copy the behavior of GPT-4, and-

Swyx37:54

It's also-

Soumith Chintala37:54

... how do I do that?

Swyx37:55

... uh, not just the behavior, but dataset. Um, so like-

Soumith Chintala37:58

Yeah

Swyx37:58

... I, I've, I've often thought of this as dataset washing. Like, you need one model at the top of the chain.

Soumith Chintala38:02

Yeah, yeah.

Swyx38:03

Um, you know, unnamed French company- ... that has the, you know, makes a model that has all the data in it that we don't know where it's from, but it's open source. Hey, and then we distill from that.

Soumith Chintala38:11

Yeah.

Swyx38:11

And it's great.

Soumith Chintala38:13

Yeah.

Swyx38:14

Um, so th- but th- they also-- To be fair, uh, they also use, uh, larger models as judges or for preference ranking, right? Like, so-

Soumith Chintala38:21

Yes

Swyx38:21

... th- that is, uh, I think in a very, very accepted use of syn-synthetic data.

Soumith Chintala38:25

Correct. I think, uh, it's a very interesting time where we don't really have good social, like, uh, social models of what is acceptable, uh, in term-- de- de- depending on how many bits of information you use from someone else, right?

It's like, okay, you use, like, one bit. Is that okay?

Swyx38:50

Mm-hmm.

Soumith Chintala38:50

Yeah, that's accepted to be okay. Okay, what about if you use, like, 20 bits? Is that okay? I don't know. What if you use, like, 200 bits? Like, I don't think we as society have ever been, uh, in this conundrum where we have to be like, where is the boundary of copyright or where is the boundary of socially accepted understanding of copying someone else?

Swyx39:16

Yeah.

Soumith Chintala39:16

Like this-- Like, we haven't been tested this mathematically before, in my opinion, so.

Swyx39:21

Yeah. And whether it's transformative use.

Soumith Chintala39:23

Yes.

Swyx39:23

Um, so I-- Yeah. I, I think this New York Times OpenAI case is gonna go to the Supreme Court-

Soumith Chintala39:28

Yeah

Swyx39:28

... and we'll have to decide it because ultimately-

Soumith Chintala39:29

I think it'll be very interesting

Swyx39:30

... never had to deal with it before. Uh, and then fi- finally, for synthetic data, the, the thing that I'm personally exploring is, um, solving this, um, very stark paradigm difference between RAG and fine-tuning, um, where you can kind of create synthetic data off of your, um, retrieved documents-

Soumith Chintala39:46

Yeah

Swyx39:46

... and then fine-tune on that. That's kind of synthetic. Um, all you need is variation or, uh, diversity of, of samples-

Soumith Chintala39:54

Right

Swyx39:54

... for you, for you to fine-tune on, and then you can fine-tune new knowledge into your, your dataset, uh, your, your model.

Soumith Chintala39:58

Yeah.

Swyx39:59

Um, I don't know if you've seen that as a direction for synthetic data.

Soumith Chintala40:03

I think that is, that is-- Like, you're basically trying to create very-- Like, what you're doing is you're saying, "Well, language, I, I know how to parameterize language to an extent."

Swyx40:15

Yeah.

Soumith Chintala40:15

And I need to teach my model variations of this input data so that it's resilient or invariant-

Swyx40:23

Yes

Soumith Chintala40:23

... to the, to language users of that data.

Swyx40:26

Yeah, it doesn't overfit on-

Soumith Chintala40:27

Yeah

Swyx40:27

... the raw source documents.

Soumith Chintala40:27

So I think that's 100% like synthetic, right? You understand-- Like, the key is, like, you create variations of your documents, and you know how to do that because you have a symbolic model or like-

Swyx40:38

Okay

Soumith Chintala40:38

... some implicit symbolic model of language.

Swyx40:42

Okay.

Alessio40:43

Do you think the, the issue with symbolic models is just the architecture of the language models that we're building? I, I think, like, the-- maybe the thing that people grasp is, like, the inability of transformers to deal with numbers because of the tokenizer.

Um, is it a fundamental issue there too, and do you see alternative architectures that will be better with symbolic understanding?

Soumith Chintala41:06

I am not sure if it's a fundamental issue or not. I think we just don't understand transformers enough.

Alessio41:13

Mm-hmm.

Soumith Chintala41:13

Uh, I, I don't even mean transformers as an architecture. I mean, like, the use of transformers today, like combining the tokenizer and transformers and the dynamics of training, like when you, when you show math-heavy questions-

Alessio41:28

Mm-hmm

Soumith Chintala41:29

... versus not. I don't have a good calibration of whether I know the answer or not. Um, I-- You know, there's common criticisms that are like, "Well, you know, transformers will just fail at X." But then when you scale them up to sufficient scale, um, they actually don't fail at that X.

Alessio41:48

Mm-hmm.

Soumith Chintala41:49

I think this is, this is entire subfield where they're trying to figure out these answers called, like, the science of deep learning or something. So we'll, we'll, we'll get to know more. I, I don't know the answer.

Meta AI42:00

Swyx42:00

Got it. Um, let's touch a little bit on, uh, just Meta AI and, uh, you know, stuff that's going on there. Maybe-- I, I don't know how- ... deeply you're personally involved in it, but, uh, you're our first guest from Meta AI ...

which is really fantastic. And Llama 1 was, uh, you know, uh, you know, you, you, you are such a believer in open source. Llama 1 was more or less like the real breakthrough in, in open source AI. Um, uh, the most interesting thing for, for, for us on covering on this sto- in, in this podcast was, uh, the death of Chinchilla, as, as people say.

Soumith Chintala42:30

Mm.

Swyx42:30

Um, any, any interesting insights there around like the scaling models for s- for open source models or smaller models or whatever that, that design decision was when, when you guys were doing it?

Soumith Chintala42:41

So Llama 1 was Guillaume Lample and team. Um, there was, uh, OPT before, which I have-- I think I'm also very proud of, um-

Swyx42:52

That's true

Soumith Chintala42:53

... because we bridged the gap in understanding of, um, how complex it is to train these models to the world. Like, until then, no one really in gory detail published-

Swyx43:07

The logs.

Soumith Chintala43:07

Yeah. Like, like why is it complex? And everyone says like- ... "Oh, it's complex," but no one really talked about why it's complex.

Swyx43:16

Mm.

Soumith Chintala43:17

Um, so I, I, I think OPT was cool. Uh, we probably-

Swyx43:20

Also, also I met Susan, and she's very, very outspoken.

Soumith Chintala43:23

Yeah. We probably, I think s- uh, didn't train it for long enough, right? Like, you know, that's, that's kind of obvious at, in retrospect.

Swyx43:32

For a 175B?

Soumith Chintala43:34

Yeah.

Swyx43:34

But, uh, we trained it accord- uh, you trained it according to Chinchilla at the time, or, or...?

Soumith Chintala43:39

I, I can't remember-

Swyx43:40

Okay

Soumith Chintala43:40

... the details, but I think it's a commonly held belief at this point that like, well, if we trained OPT longer, it would actually end up being better. Um, Llama 1, I think was, yeah, Guillaume Lample and team.

Guillaume is fantastic, uh, and went on to build Mistral. I wasn't too involved in that side of things, uh, so I don't know what you're asking me, which is like, well, like how did it, they think about scaling laws and all of that.

Um, Llama 2, I was more closely involved in. Um, I helped them a reasonable amount with like their, um, infrastructure needs and stuff. Llama 2, I think was more like, let's get to, to the evolution. At that point, we kind of understood what we were missing from the industry's understanding of LLMs, and we needed more data, and we needed more to train the models for longer.

And we made, I think, a few tweaks to the architecture, and we scaled up more. Um, and like that was Llama 2. I think Llama 2, you can think of it as like after Guillaume left, the team kind of rebuilt their muscle around, um, around Llama 2.

And Hugo, I think, who's the first author, is fantastic, and I think he, he did play a reasonable big role in Llama 1 as well, and he overlaps between Llama 1 and 2. So-- and Llama 3 obviously, hopefully, will be awesome.

Swyx45:19

Mm-hmm.

Alessio45:20

Um, just one question on Llama 2, and then we'll try and fish Llama 3 spoilers out of you. Uh, in the Llama 2 paper, the, the loss curves of the 34 and 70B parameter, they, they still seem kind of steep.

Soumith Chintala45:33

Mm.

Alessio45:33

Feel like they, they could go lower. How, from an infrastructure level, how do you allocate resources? Like, uh, could they have just gone longer, or were you just like, "Hey, this is all the GPUs that we can burn, and let's just move on to Llama 3 and then make that one better"?

Soumith Chintala45:47

Instead of answering specifically about like that Llama 2 situation or whatever, I'll tell you like how we think about things.

Alessio45:54

Mm-hmm.

Soumith Chintala45:54

Generally, we're-- we have-- I mean, Mark released some numbers, right? He, I mean-

Swyx46:02

So let's, let's, uh, cite those things again.

Soumith Chintala46:04

Yeah.

Swyx46:04

Uh, s- uh, all I remember is like 600K GPUs.

Soumith Chintala46:07

Uh, that is by the end of this year, and 600K H100 equivalents.

Alessio46:11

Equivalent. Yeah, okay.

Soumith Chintala46:12

Uh, with 250K H100s and including all of the, our other GPU or accelerator stuff, it would be 600 and something, uh, K, um, aggregate capacity.

Alessio46:27

Mm-hmm.

Soumith Chintala46:28

That's a lot of GPUs. We'll talk about that separately. But, um, the way we think about it is we have a, we have a train of models, right? Llama 1, 2, 3, 4. Um, and we have a bunch of GPUs.

I, I don't think we're short of GPUs. Like-

Alessio46:44

Yeah, no, I wouldn't say so.

Soumith Chintala46:46

Yeah. So I think the-- it's, it's all a matter of time. I think time is the biggest bottleneck.

Alessio46:53

Mm.

Soumith Chintala46:53

It's like, when do you stop training the previous one, and when do you start training the next one? And how do you make those decisions? Um, the, the data, do you have net new data, better clean data for the next one in a way that it's not worth like really focusing on the previous one?

It's just a standard iterative product. You're like, when is the iPhone 1?

Alessio47:15

Mm.

Soumith Chintala47:15

Uh, when, like when do you start working on iPhone 2? Where is the iPhone-- like so on, right?

Alessio47:20

Yeah.

Soumith Chintala47:20

Um, so mostly the considerations are time and generation rather than GPUs, in my opinion.

Alessio47:28

So one, one of the thing with the scaling laws, like Chinchilla is like optimal to balance training and inference costs. I think at Facebook scale or Meta scale, you would rather pay a lot more maybe at training and then save on inference.

How do you think about that from a infrastructure perspective? Uh, I, I think in your tweet you say you can try and guess on like how we're using these GPUs. Can you just give people a better understanding? It's like-- because I've already seen a lot of VCs say Llama 3 is being trained on 600,000 GPUs, and that's obviously not true, uh, I'm sure.

Um, how do you allocate between the research like FAIR and, uh, the Llama training, the inference on Instagram suggestions that get me to scroll-

Soumith Chintala48:09

Right

Alessio48:09

... like the AI-generated stickers on WhatsApp- ... and all that?

Soumith Chintala48:13

Yeah. Um, we haven't talked about any of this publicly, uh, but like as a broad stroke, it's like how we would allocate resources of any other kinds at any company. Um, you, you run a comp- you, you run like a VC portfolio.

Like, how do you allocate, um, your-- uh, how do you allocate your investments between different companies or whatever? You kind of make various trade-offs, and you kind of decide, "Should I invest in this project or this other project, or how much should I invest in this project?"

It's very much like a s- a, a zero-sum of trade-offs. And it also comes into play, like, you know, how is your-- how, how are your, like, clusters configured-

Alessio49:00

Mm-hmm

Soumith Chintala49:00

... like overall? Like, you know, what you can fit of what size and what cluster, and so on. So broadly, there's no magic sauce here. Like, I mean, I think the details would add more spice, but also wouldn't add more understanding.

Alessio49:17

Mm.

Soumith Chintala49:17

Uh, it's just gonna be like, "Oh, okay, I mean, this looks like they just think about this as I would normally do."

Alessio49:24

Right. So even the GPU-rich run through the same struggles- ... of having to, to decide where to allocate things.

Soumith Chintala49:31

Uh, yeah. I mean, like, at some point, uh, I forgot who said it, but it's like you kind of fit your models to the amount of-

Alessio49:42

Mm-hmm

Soumith Chintala49:42

... compute you have. If you don't have enough compute, you figure out how to make do with smaller models. But, like, no one, as of today, I think would feel like they have enough compute. I don't think, like, I've heard any company within the AI space, uh, be like, "Oh yeah, like we feel like we have sufficient compute, and we couldn't have done better."

Uh, so like that, that conversation-

Alessio50:12

Mm

Soumith Chintala50:12

... I don't think I've heard from any of my friends at other companies. Um-

Alessio50:16

Stella, Stella from Eleuther sometimes says that because she has a lot of donated compute.

Soumith Chintala50:21

Yeah.

Alessio50:22

Um, and she's trying to put it to interesting uses. But, uh, for some reason, she's decided to stop, uh, making large models. Uh, so there-

Soumith Chintala50:29

I mean, that's a, that's a cool high-conviction opinion that might pay out, right?

Alessio50:36

Yeah.

Soumith Chintala50:36

I mean, she's taking a path that most people don't care to take about in this climate, and she probably will have very differentiated ideas.

Alessio50:45

Yeah.

Soumith Chintala50:46

Um, I mean, think about the correlation of ideas in AI right now. It's so bad, right? Like- So everyone's fighting for the same pie. Um, uh, in some weird sense, like that's partly why I don't really directly work on LLMs.

I used to be a gen- gen-- like I used to do image models and stuff.

Alessio51:07

GANs. Yeah, mm-hmm.

Soumith Chintala51:08

And I actually stopped doing GANs because, uh, GANs were getting so hot that I didn't have any calibration of whether, like, my work would be useful or not. Because, oh yeah, like someone else did the same thing you did.

It's like there's so much to do, I don't understand why I need to like, uh, fight for the same pie. So like, you know, I, I think like Stella's decision is very smart.

Alessio51:33

And h- how do you reconcile that with how we started the discussion about, uh, intrinsic versus extrinsic kind of like a accomplishment-

Soumith Chintala51:42

Yeah

Alessio51:42

... success? H- how should people think about that when, especially when they're doing a PhD or like early in their career? Seems like-- I, I think at NeurIPS, I walked through, uh, a lot of the posters and whatnot.

There seems to be mode collapse i- in a way in the research.

Soumith Chintala51:56

Yeah.

Alessio51:56

A lot of people working on, on the same things. Is it worth for like a PhD to not take a bet on something that is like maybe not as interesting, you know, just because of funding and, you know, visibility and whatnot?

Or, uh, yeah, what, what suggestions would you give?

Soumith Chintala52:10

I think there's a baseline level of compatibility you need to have with the field. Uh, basically, you need to figure out if you will get paid enough to eat, right?

Alessio52:22

Mm.

Soumith Chintala52:22

Like, and like whatever reasonable normal lifestyle you want to have as a baseline. So you at least have to pick a problem within the neighborhood of like fundable. Like- ... you, you wouldn't want to be doing something so obscure that people are like, "Ah, I don't know."

Like you can work on it.

Alessio52:43

Would, would a limit on fundability-- I'm just like observing something like three months of compute, right? That's the top line. That's the, like, max that you can spend on any one project.

Soumith Chintala52:54

But like I, I think that's very ill-specified, like how much compute?

Alessio52:57

Yeah.

Soumith Chintala52:58

Right? So I think, uh, I think the, the notion of fundability is broader. It's more like, "Hey, are these family of models within the acceptable set of you're not crazy or something," right? Like even something like neural ODEs, which is a very like boundary-pushing thing, or like state-space models or whatever.

Like all of these things I think are still in fundable territory, right? You're talking about, "I'm gonna do one of the neuromorphic models, um, and then apply, like, image classification to them or something," then it becomes like a bit questionable.

Again, it depends on your motivation. Maybe if you're a neuroscientist, it actually is feasible. But if you're like a AI engineer, like the audience of this podcast, then it's less quest- you know, it's more questionable. So I, I think like the way I think about it is like you, you need to figure out how you can be in the baseline level of fundability just so that you, you can, you can just live.

And then after that, really focus on intrinsic motivation and, um- Depends on your strengths, like how you can play to your strengths and your interests at the same time. Like you-- Like I try to look at a bunch of ideas that are interesting to me, uh, but also try to play to my strengths.

Swyx54:31

Mm-hmm.

Soumith Chintala54:32

I'm not gonna go work on theoretical ML. Um, I'm interested in it, but when I want to work on something like that, I try to partner with someone who is actually a good, like, theoretical ML person, and see if I actually have any value to provide.

And if they think I do, then I come in. So I think you'd want to find that intersection of ideas you like and that also play to your, your strengths, and I'd go from there. Everything else, like actually finding extrinsic success and all of that, I think is-- the way I think about it is, like, somewhat immaterial.

Swyx55:07

Mm-hmm.

Soumith Chintala55:08

And when you're talking about building ecosystems and stuff, like slightly different considerations come into play. But that, that's a, that's a different conversation.

Swyx55:17

Yeah. Um, I, I should-- We-we're, we're gonna pivot a little bit to just talk, talk, talk about open source AI. Um, but one, one more thing I wanted to establish for Meta is, like this 600K number, just kind of rounding out the discussion, uh, that's for all Meta.

Um, so including-

Soumith Chintala55:32

Right

Swyx55:32

... your own inference needs, right? It's not just about training.

Soumith Chintala55:34

It's, I, it's for all-- It's, it's gonna be the number in our data centers-

Swyx55:39

Yeah

Soumith Chintala55:39

... for all of Meta, yeah.

Swyx55:39

Yeah. So like, you know, there, there's a decent amount of workload serving Facebook and Instagram and, you know, whatever. Um, and then, uh, is there interest in like your own hardware?

Soumith Chintala55:50

We already talked about our own hardware. Um, it's called MTIA.

Swyx55:56

Yeah.

Soumith Chintala55:56

Our own silicon. Uh, I think we've even showed like the standard photograph of you holding- ... like the chip that doesn't work. I mean, like as in the chip that you basically just get like-

Swyx56:12

As a test run.

Soumith Chintala56:13

Yeah, a test chip-

Swyx56:14

Yeah

Soumith Chintala56:14

... or whatever. Um, so we are working on our silicon, and we'll probably talk more about it when the time is right. But-

Swyx56:25

Like what gaps do you have that, you know, the market doesn't offer?

Soumith Chintala56:29

Okay. I mean, this is easy to answer. So basically, uh, remember how I told you about the whole sweet-- like there's, there's this memory hierarchy and like sweet spots and all of that? Fundamentally, like when you build a hardware, like you, you make it general enough that a wide set of customers and a wide set of workloads can use it effectively while trying to get the maximum level of performance they can.

Um, the more specializ-specialized you make the chip, the, the more, uh, hardware efficient it's going to be.

Swyx57:02

Mm-hmm.

Soumith Chintala57:03

The more power efficient it's gonna be, the more easier it's going to be f- to find like the software, um, like the kernels to write to just map one, that one or two workloads to that hardware, and so on.

Um, so it's pretty well understood across the industry that if you have a sufficiently large, large volume enough workload, you can specialize it and get some efficiency gains, like power gains and so on. So the way you can think about everyone building-- every large company building silicon, like I think, um, a bunch of the other large companies are building their own silicon as well, is they-- each large company has a sufficient enough set of verticalized workloads, uh, that have a pattern to them-

Swyx57:59

Mm-hmm

Soumith Chintala57:59

... that, say, a more generic accelerator like an NVIDIA or an AMD GPU does not exploit. So there is some level of power efficiency that you're leaving on the table by not exploiting that, and you have sufficient scale, and you have sufficient, uh, s- forecasted stability that those workloads will e-exist in the same form, that it's worth spending the time to build out a chip, uh, to, to exploit that sweet spot.

Like obviously something like this is only useful if you hit a certain scale and that your, like, forecasted prediction of those kind of workloads being in the same kind of specializable ex- you know, exploitable way is true. Um, so yeah, that's, that's, that's why we're building our own chips.

Swyx58:55

Mm-hmm. Amazing.

Alessio58:56

Awesome. Um, yeah, I know we-we've been talking a lot on, on a lot of different topics, and g-going back to open source, you had a very good tweet. You said that a single company's closed source effort really limits against people's imaginations and needs.

Open Source AI58:56

Alessio59:10

How do you think about that? How do you think about all the impact that some of the Meta AI work, uh, in open source has been doing, and maybe directions of the whole open source AI space?

Soumith Chintala59:20

Yeah. Um, in general, I think-- first, I think it's worth talking about this in terms of open, uh, and not just open source because like with the whole notion of model weights- ... no one even knows what source means for these things.

Uh, but like just, just for the discussion, when I say open source, you can assume it's just I'm talking about open.

Alessio59:42

Mm-hmm.

Soumith Chintala59:42

And then there's the whole notion of like licensing and all that, like, you know, what happens.

Alessio59:46

Commercial.

Soumith Chintala59:47

Commercial, non-commercial, commercial with clauses and all that. I think like at a fundamental level, uh, the most benefited value of open source is that you make the distribution to be very wide. Like it's just available with no friction, and like you can-- people can do transformative things, um, in a way that's very accessible.

Like- Maybe like it's open source, but it has a commercial license and I'm a student like in India. I don't care about the license. I just don't even understand the license. But like the fact that I can use it and do something with it is very transformative to me.

Like I, I got this thing in a very accessible way. Um, and then like, so it's very j- very-- various degrees, right? Like, and then like if it's open source, but it's a, like a, actually like a commercial license, then a lot of companies are gonna benefit from like gaining value that they didn't previously have, that they maybe had to pay a closed source company for it.

So open source is just a very interesting tool that you can use in various ways. So there's, again, two kinds of open source. One is like some large company doing a lot of work and then open sourcing it, and that kind of effort is not really feasible by, say, like a band of volunteers doing it the same way.

So there's both a capital and operational expenditure that the large company just decided to, um, ignore and give it away to the world, uh, for some benefits of some kind. Uh, they're not as tangible as like direct revenue or something.

Mm-hmm. So in that part, Meta has been doing incredibly good things. Um, they fund a huge amount of the PyTorch, uh, development. They've open sourced Llama and those family of models, um, and several other fairly transformative, uh, projects.

Uh, FAISS is one, Segment Anything- Mm. -Detectron, Detectron 2, uh, DensePose. I mean, it's- Seamless. Yeah, Seamless. Like it's just like the list is so long that, you know- ... we're, we're not gonna cover. So like I think Meta comes into that category where like we spend a lot of CapEx and OpEx, and, um, we have a high talent density of great- Mm ...

AI people, and we open our stuff. And the thesis for that, uh, I remember when FAIR was started, the common thing was like, "Wait, why would Meta wanna start a open AI lab? Like, what, what, what exactly is the benefit like from a commercial perspective?"

Mm. And for then, like the thesis was very simple. It was like, AI is currently rate limiting Meta's ability to do things. Um, our ability to, um, build various product integrations, moderation, various other factors. Like AI was the limiting factor.

Mm. And we just wanted AI to advance more, and we didn't care if the IP of the AI was uniquely in our possession or not. For us, like however the field advances, that accelerates like Meta's ability to build a better product.

So we just built like an open AI lab, and we- Mm ... said, "If this helps accelerate the progress of AI, that's strictly great for us." But like very easy rationale, right? Still the same to a large extent with like the Llama stuff, and, uh, it's, it's a bit more, um...

I think it's, it's the same values, but like, you know, the argument, like it's a, it's a bit more nuanced. Um, and then there's the second kind of open source, which is, "Oh, you know, we built this project nights and weekends, and we're very smart people, and we open sourced it, and then we built a community around it."

This is like the Linux kernel- Mm ... and, uh, various software projects like that. So, um, I think about open source b- like both of these things being beneficial and both of these things being different. Um, they're, they're different and beneficial in their own ways.

Uh, the second one is really useful when there's an active arbitrage to be done. Um, if, if someone's not really looking at a particular space, uh, because it's not commercially viable or whatever, like a band of volunteers can just coordinate online and do something and then make that happen.

Uh, and that's great. Um, I wanna cover a little bit about like open source LLMs maybe. Mm-hmm. Um, so open source LLMs have been very interesting because I think we were trending towards a, an increase in open source in AI from 2010, uh, all the way to like 2017 or something.

Like where more and more pressure, um, within the community was to open source their stuff- Mm ... so that their methods and stuff get, get adopted. And then the LLMs revolution kind of, uh, took the opposite effect. Um, OpenAI stopped open sourcing their stuff.

Mm-hmm. And, uh, DeepMind kind of, you know, didn't. Like, you know, all the other, Cloud and all these other like providers, they, they didn't, uh, open source their stuff. And it was not good, uh, in the sense that first, like science done in isolation probably will just form its own bubble where like people believe their own bullshit or whatever, right?

So there is, there is that problem. Uh, and then there was the other problem, which was the accessibility part. Like, okay, uh, I again always go back to like, I'm a student in India with no money. Mm. Um, what is my accessibility to any of these closed source, closed source models?

Uh, at, at some scale I have to pay money. Um, that makes it a non-starter and stuff. And there is also the control thing. I strongly believe the best, um, if you want human-aligned stuff, you want all humans to give- Mm-hmm ...

feedback, and you want all humans to have access to that technology in the first place. Um, and I actually have seen b- because, you know, living in New York, whenever I come to Silicon Valley, I see a different cultural bubble.

Uh, like all the friends I hang out with talk about some random thing, like Dyson spheres or whatever, you know, that's a thing. And most of the world doesn't know or care about any of this stuff. Like it's like definitely like a bubble.

Mm-hmm. And bubbles can form very easily. And when you make a lot of decisions because you're in a bubble, uh, they're probably not globally optimal decisions. So I think like open source, the distribution of open source powers a certain kind of non-falsifiability- Mm-hmm ...

that I think is very important. Um, so I think, uh, on the open source models, like it's going great in the fact that LoRA, I think, came out of the necessity of open source models, uh, needing to be fine-tunable in some way.

Um- GPT. Yeah. And- Without-- Yeah, without a lot of GPUs ... and I think DPO also came out of like, uh, like, um, the academic open source- Yes. Mm-hmm ... side of things. So

why, like, do any of the closed source labs all-- Did, did any, did any of them already have LoRA or DPO internally? Maybe, but like that does not advance- Mm-hmm ... like humanity in any way. It advances like some company's probability of doing the winner takes all, uh, that I talked about earlier in the podcast.

So, um, I don't know, it just feels fundamentally good. Like when people try to, you know, people are like, "Well, like what are the ways in which it is not okay?" And this, this might be a little controversial, but like I find a lot of arguments, uh, based on whether like closed source models are safer or open source models are safer, very much related to whether what kind of cultural, uh, culture they grew up in, what kind of society they grew up in.

If they grew up in a society that they trusted, then I think they take the closed source argument. Mm-hmm. And if they grew up in a society that they couldn't trust, where the norm was that you, ah, you didn't trust your government, obviously, like it's corrupt or whatever- ...

then I think like the open source argument is what they take. I think there's a deep connection to like people's innate, um, innate biases from their childhood and their trust in society and governmental aspects that push them towards one opinion or the other.

And I'm definitely in the camp of, um, open source is definitely going to actually have better outcomes for society. Closed source to me just means that centralization of power, which- Mm-hmm ... you know, is really hard to trust.

Um, so I think it's, it's, it's, it's going well in so many ways. Um, there's not, like s- we're actively disaggregating the centralization of power to just like two or three providers. We are, I think, benefiting from like so many people using these models in so many ways that aren't allowed by like, say- Right ...

like Silicon Valley left-wing, um, um, tropes. Like some of these things are good or bad, but like they're not culturally accepted universally in the world. Um, so those are things worth thinking about. And I think open source is not winning in certain ways.

Uh, like these are all the things in which like, uh, as I mentioned, it's actually being very good and beneficial and winning. I think one of the ways in which it's not winning, at some point I should write a long form post about this, is I think it has a classic, uh, coordination problem.

I mean, open source in general always has a coordination problem. Right. If there's a vertically integrated provider with more resources, um, uh, they will just be better coordinated than open source. And so now open source has to figure out how to have coordinated benefits.

And the reason you want coordinated benefits is because these models are getting better, uh, based on human feedback. Um, and if you see with, with open source models, like if you go to like Reddit, LocalLLAMA subreddit, like- Mm-hmm ...

there's so many variations of models that are being produced from, say, nose research to-- I mean, like there's like so many like variations built by so many people. And one common theme is they're all using these fine-tuning or human preferences datasets that are very limited, and like someone published them somewhere, and like they're, they're, they're not sufficiently diverse.

And you, you look at the other side, like say frontends like Ooba or like, um, Hugging Chat or, uh, Ollama, they don't really have like feedback buttons. Like- Mm-hmm ... all the people using all these frontends, they probably want to give feedback, but there's no way for them to give feedback.

So these models are being built, they're being, uh, arbitrarily measured, and then they are being deployed into all these open source frontends, uh, or like apps that are closed source- Yeah ... they're serving open source models. And these mo- these frontends don't have-- They are not exposing the ability to give feedback.

So we're just losing all of this feedback Maybe open source models are being as used as GPT is at this point in like all kinds of way, in a very fragmented way. Like in aggregate, all the open source models together are probably being used as much as GPT is, maybe, you know, close to that.

But the amount of feedback that is driving back into the open source ecosystem is like negligible, maybe less than 1% of like the usage. Um, so I think like some-- like the, the, the blueprint here I think is you'd want someone to create a sinkhole for the feedback, some centralized sinkhole, like maybe Hugging Face or someone, uh, just funds like, okay, like, we-- I will make available a call to log a string along with like, you know, a, a bit of information of positive or negative or something like that.

And then you would want to send pull requests to all the like, um, open source front ends like Uber and all being like, "Hey, we're just integrating like a feedback UI." And, and, and then work with like the closed source people as also being like, "Look, it doesn't cost you anything.

Just like have a button." And then the sinkhole will have a bunch of this data coming in, and then I think a bunch of open source researchers should figure out how to filter the feedback into only the like high quality one.

I'm sure like it'll be exploited by spam bots or whatever, right? Like this is like the perfect way to inject your advertising product into like the next-

Alessio1:14:03

Buy Coca-Cola. Yeah.

Soumith Chintala1:14:05

So, uh, there needs to be some level of that. Uh, that in the same way, I'm sure like, like all the closed providers are doing today, like OpenAI, Claude, like with the-- like the feedback that comes in, I'm sure they are figuring out if that's legit or not.

So that kind of data filtering needs to be done. And that f- loop has to be set up, and this requires that central sinkhole and that like data cleaning effort both to be like there. They're not there right now.

They're not there right now, I think for, for capital reasons, but also for coordination reasons. Okay, if that central sinkhole is there, who's gonna go coordinate all of this integration across all of these like open source front ends?

But I think if we do that, if that actually happens, I think that probably has a real chance of the open source models having a runaway effect against, um, OpenAI with their current like daily active users rumored. Um, probably doesn't have a chance against Google because, you know, Google has Android and Chrome and Gmail and Google Docs and everything, you know.

So people just use that, uh, a lot. Uh, but like I think like there's a clear chance we can take at, um, truly winning open source.

Alessio1:15:35

Do you think this feedback is helpful to make open source models better or to get to like open source AGI? Because in, in a way, like OpenAI's goal is to get to AGI, right? So versus I think in open source, we're more focused on personal better usage or like commercial better usage.

Soumith Chintala1:15:51

Yeah. I think that's a good question, but I think like largely, I, I actually don't think people have a good understanding of AGI, and I don't mean definition level. I mean, people are like, "Okay, we're gonna-- AGI means it's powering 40% of world economic output," or some, some, something like that, right?

But what does that mean? So do you think electricity is powering 40% of world economic output or is it not? Like generally, the notion of like powering X percent of economic output is not defined well at all for me to understand like how to know when we got to AGI or how, how to measure whether we're getting to AGI.

Like, you know, you can look at it in terms of intelligence or task automation or whatever. And I think that's what we are doing right now. We're basically integrating like the current set of AI technologies into so many real world use cases where we find value that if some new version of AI comes in, we can find-- like we can be like, "Ah, this helps me more."

Um, in that sense, I think like the whole process of like how we think we got to AGI will be continuous and not like, not discontinuous like how I think, uh, the question is posed. So I think the open source thing will be very much in line with, um, getting to AGI because open source has that like, uh, natural selection effect.

Like if a better open source model comes, really no one says, "Ha, I don't wanna use it because there are ecosystem effect. I'm logged into my ecosystem," or like, "I don't know if I like the models," you know, whatever.

It's just a very pure direct thing. So if there's a better model that comes out, then it will be used. Uh, so I, I, I definitely think it, it, it has a good chance of achieving how I would think about as a continuous, um, path to what we might define as AGI.

Swyx1:18:15

Um, for the listeners, I would actually mention, uh, a couple other maybe related notes on just, uh, uh, this, this very interesting concept of, uh, feedback sinkhole for, for open source to really catch up, um, in terms of the, the overall Google versus OpenAI debate.

Um, Open Assistant was, was led by, uh, Yannic Kilcher, who recently ended his effort. And I think the criticism there was, like, the kind of people that go to a specific website to give feedback, uh, is not representative of real-world usage, and that's why the, the models trained on Open Assistant, uh, didn't, didn't really seem like they have caught on in the open source world.

Um, the two leading candidates in my mind are LMSYS out of UC Berkeley, um, who have the LSS-- LMSYS Arena, which, um, you know, is being touted as one of the only ways-- only reliable benchmarks anymore. Like, I kinda, kinda call them nonparametric benchmarks 'cause there's nothing to cheat on here except for ELO.

Uh, and then the other one is, um, um, OpenRouter-

Soumith Chintala1:19:06

Yeah

Swyx1:19:06

... which is Alex Otala's thing. I don't know if you've, uh, talked to any of these people.

Soumith Chintala1:19:10

I obviously know all, all of the efforts that you talked about. I haven't talked to them directly about this yet, but, uh, the way I think about it is the way these models are going to be used is always going to be way more distributed than centralized.

Um, like, which is the power of the open source movement. Like, the, the UI within which these models are going to be used is going to be decentralized. Like, it's-- these models are gonna be integrated into, like, hundreds-

Swyx1:19:42

Mm-hmm

Soumith Chintala1:19:42

... and thousands of projects and products and all of that, right? And I think that is important to recognize. Like, like, the LMSYS leaderboard is the best thing we have right now to understand whether a model is better or not versus another model, but it's also biased and only having a sliver of view into how people actually use these models.

Like, the-

Swyx1:20:06

Mm

Soumith Chintala1:20:06

... the people who actually end up coming to the LMSYS leaderboard and then using a model only use it for certain things. Like, like G- like GitHub Copilot style usage is not captured-

Swyx1:20:20

Mm-hmm

Soumith Chintala1:20:20

... in, say, like LMSYS thing. And so many other styles, like the Character AI style, uh, things-

Swyx1:20:26

Mm

Soumith Chintala1:20:26

... is not captured in LMSYS.

Swyx1:20:27

Which, which OpenRouter could do. They don't do it n- right now, but-

Soumith Chintala1:20:30

Yeah.

Swyx1:20:30

Yeah.

Soumith Chintala1:20:30

So, like, I, I think, like, yeah, my, my point is, like, the way these models are going to be used is going to be always a large surface area, and I think we need to figure out how to provide the infrastructure to integrate with all these, like, ways in which it's being used.

Even if you get, like, the top hundred front ends that this, this-- the, the model-- like the open source models are used through, subscribe to like the sinkhole, I think that's already like a substantial thing. I think, like, thinking one or two things built by themselves get a lot of data I think is not gonna happen.

Swyx1:21:14

Yep. Fair enough.

Beyond Text1:21:16

Alessio1:21:16

Um, before we let you go, uh, can we do just a quick beyond text, uh, segment? So, uh, you're an investor in Runway, which is-

Soumith Chintala1:21:24

Oh, yeah

Alessio1:21:25

... a video generation. You're an investor in 1X-

Soumith Chintala1:21:27

Mm-hmm

Alessio1:21:27

... which is a humanoid assistant. Osmo, which is focused on using AI for smell recognition and synthesis. Uh, you advise a bunch of robotics projects at, at NYU. Maybe-

Swyx1:21:38

And he, and he builds his own home robot.

Soumith Chintala1:21:40

Yeah.

Alessio1:21:41

Yeah, yeah. E- exactly. On a more, yeah, maybe open-ended thing, what are like the things that you're most excited about beyond like text generation and kind of the more mundane usage?

Soumith Chintala1:21:50

Yeah, I mean, in general, I have more things that I'm generally excited about than I can possibly do. Uh, investing is one way to try to clear those urges. Um, I'm generally excited about robotics, uh, being a possibility, home robotics being like five to seven years away into commercialization.

Alessio1:22:18

Mm.

Soumith Chintala1:22:18

I think, like, it's not like next year or two years from now, but like, I think five to seven years from now, I think, like, a lot more robotics companies might pop out. Um, there's not a good consensus on whether hardware is a bottleneck or AI is a bottleneck in robotics right now.

Um, my view is actually hardware is still the bottleneck.

Alessio1:22:40

Mm.

Soumith Chintala1:22:41

And AI is also a little bit of bottleneck, but like I don't think there's any, like, obvious, um, breakthroughs we need. I think it's just work. So I'm generally excited about robotics. I spend a lot of time, a lot of personal time.

I spend like every Wednesday afternoon at NYU working with Lerrel Pinto and team, and just getting towards my like home robot that just-

Alessio1:23:06

Mm-hmm

Soumith Chintala1:23:06

... does my dishes and stuff. Um-

Swyx1:23:08

What's the status of it? Like what, what, what does it do for you now?

Soumith Chintala1:23:10

As of today, uh, we just deployed, um, a couple of months ago, we deployed our, our home robotics stuff into like several tens of New York City homes and like tried to make it do a bunch of tasks.

And we're basically starting to build out a framework that gets to a certain level of robustness on fairly simple tasks, like, you know, picking this cup and putting it somewhere else, or like, um, taking a few pieces of cloth on the ground and put it somewhere else.

Alessio1:23:47

Mm.

Soumith Chintala1:23:47

Uh, or open your microwave. Like various like, like baseline tasks like that, um, with low sample complexity. So like our-- the key thing-- I, I think one of the things people don't spend enough time in robotics is like the user experience, uh, which I think we s- in the research I do at NYU, we spend a huge amount of time on.

I think the key there is sample complexity has to be really low. Uh, a lot of the current robotics research, if you see, they're like, "Oh yeah, we collected like fifty demos, and now it's able to do this task," or, "We collected like three hundred demos," or like it's the sample-- the number of samples you need for this thing to do the task is really high.

So we're focusing a lot on You c— you show it like two or three times, and that's sufficient for it to actually, like, do the task. Um, but it comes with, like, less generalization.

Swyx1:24:39

Mm-hmm.

Soumith Chintala1:24:40

Right? Like you-- there's some initial conditions that have to be true for it to do the task. Uh, um, so we're making progress. That's very interesting in general, the space. Um, I don't think people in this space have settled on the hardware.

Swyx1:24:55

Mm-hmm.

Soumith Chintala1:24:56

Like, you know, how the hardware looks like for it to be truly useful in the home or whatever, um, how— or the UX or the, like, AI, ML stuff needed to, to make it sample efficient and all of that.

Uh, but I think, like, lots of work is happening in the field.

Swyx1:25:15

Mm-hmm. Yeah. One, one of my friends, Carlo at Berkeley, he worked on a project called M3L-

Soumith Chintala1:25:20

Uh-huh

Swyx1:25:20

... which is two CNNs, one for tactile feedback and one for image.

Soumith Chintala1:25:24

Yeah.

Swyx1:25:25

Uh, w-when you say hardware, is it running all these things on the edge, or is it just, like, the, the actual servos and, uh, the-

Soumith Chintala1:25:33

Um, yeah. By hardware, I mean, like, the actual, like, servos, like, uh, the, the motors, servos, the, the— even, like, the sensors. Um, I think we have incredible vision that still, like, is, is so much better compared to—in the field of view and in resolution compared to, um, any of the cameras we can buy.

We have— Our, our skin is, like, all available touch sensing.

Swyx1:26:05

Mm-hmm.

Soumith Chintala1:26:05

And we have, like, some of the most efficient, you know, s-some of the most, uh, high capacity, um, motors that can l-lift large loads, you know, in, in, like the dexterity of a hand and stuff. So in terms of, um, hardware, I mean, like, in terms of those capabilities-

Swyx1:26:26

Mm-hmm

Soumith Chintala1:26:26

... like, you know, we haven't figured out, um, how to do a lot of this stuff. Um, I mean, Tesla has been making incredible progress. Um, 1X, uh, I think, announced their new thing that looks incredible. Um, some of the other companies, Figure and, like, others are, are doing great work.

But we're really not anywhere close to, like-

Swyx1:26:47

Mm-hmm

Soumith Chintala1:26:47

... the hardware that we feel like we need. And there's obviously—the other thing I wanna call out is, um, a lot of what people show, um, works, but, like, has to be fixed all the time. And, like, that's the other thing we, we are incredible at.

Like, we, we don't need any maintenance, or, like, the maintenance is part of us.

Swyx1:27:11

Right.

Soumith Chintala1:27:11

Um, if you buy a product, an elec- an e-electronics product of any kind, you, you buy a PS5, you don't say, "Oh yeah, my PS5 breaks, like, every six days, and I have to, like, do some reasonable amount of work on it."

Swyx1:27:24

Mm-hmm.

Soumith Chintala1:27:24

But, like, that's robotics. Like, if, if it's not industrial robotics, where it's very controlled and specialized or whatever, like, you're talking about reliability, like, in those ranges. So I think people don't talk about the reliability thing enough. Like, when I mean, like, you know, we need—we're gonna enter the commercialization phase, I mean, like, we're gonna start thinking about, "Okay, now we have this thing, and we need to figure out how to get reliability high enough-

Swyx1:27:49

Mm-hmm

Soumith Chintala1:27:49

... to deploy it into homes and, like, just sell it to people in, like, Best Buy or something."

Swyx1:27:54

Yeah.

Soumith Chintala1:27:54

So that's the other factor that, that we have to make a lot of progress on.

Swyx1:27:59

I, I, I just realized that Google has a play in this with, like, PaLM E and stuff, and OpenAI obviously, uh, has a long history of doing this stuff. Um, d- is there anything at Meta?

Soumith Chintala1:28:11

Uh, I-

Swyx1:28:11

No, no robotics stuff at Meta.

Soumith Chintala1:28:13

I used to... Uh, we, we have a small ro-robotics program at Meta out of FAIR. I actually used to do it at FAIR-

Swyx1:28:19

Okay

Soumith Chintala1:28:19

... a little bit before I moved into Infra and focused on my, my Meta time on a lot of, like, other infrastructural stuff. Um, so yeah, Meta's, uh, robotics program is a lot smaller. Uh-

Swyx1:28:32

Seems like it would be a fit.

Soumith Chintala1:28:34

I think, like-

Swyx1:28:34

You know, personal computing.

Soumith Chintala1:28:36

You can think of it as, like, Meta's-- Like, Meta has a ridiculously large device strategy, right? Like, you know, this is our, our reality labs stuff.

Swyx1:28:45

Mm-hmm.

Soumith Chintala1:28:46

Like, you know, we're going at it from VR and AR, and, you know, we showcase a lot of that stuff. I think for Meta, like, the robot is not as important as, like, the-

Swyx1:28:58

Screen real estate

Soumith Chintala1:28:59

... physical devices, physical devices kind of stuff.

Swyx1:29:01

Yeah. Yeah.

Soumith Chintala1:29:01

For sure.

Swyx1:29:02

Yeah. Um, okay, I want to touch on Osmo a bit.

Soumith Chintala1:29:04

Yeah.

Swyx1:29:04

Because, uh, very unusual company to, um, the stuff that we normally discuss.

Soumith Chintala1:29:08

Yeah.

Swyx1:29:08

Um, not robotics. Uh, sense of smell.

Soumith Chintala1:29:10

Yeah.

Swyx1:29:11

Um, the, the—my, my—the original pitch I heard from the founder, maybe you can correct me, is that the—you realize that you can smell cancer.

Soumith Chintala1:29:18

Yeah.

Swyx1:29:19

Uh, is that in-intuitive? Is that what you get? Or is that-

Soumith Chintala1:29:21

Yeah, I mean, first, like-

Swyx1:29:23

... like the potential that you see?

Soumith Chintala1:29:24

... uh, the very interesting reason I invested in Osmo is because Alex Vilenski, the founder of Osmo, also was, like, a, um... Before PyTorch, there was Torch, and Alex Vilenski actually worked on Torch.

Swyx1:29:41

Mm-hmm.

Soumith Chintala1:29:41

He's actually, like, a frameworks guy. Like, you know, he built, uh, this thing called Tangent from Google, um, like another, like, autodiff framework and stuff. Like, so I know him from that side of things. And then where—like, I also-- Like, he is a neurobiologist by training.

Um, he just happens to also love-

Swyx1:30:02

Mm-hmm

Soumith Chintala1:30:02

... like, neural networks and, like, hacking on those frameworks. So incredibly smart guy, one of the smartest people I know. Uh, so when he was going in this direction, I thought it was incredible that, like, smell is something that we haven't even started to scrape-

Swyx1:30:21

Yeah

Soumith Chintala1:30:21

... in terms of digitization. When we think about audio or images or video, they're, like, so advanced that we have the concept of color spaces. We have the concept of, like- Frequency spectrums. Like, you know, we figured out how ears process, like, uh, frequencies in mel spectrum or whatever, like logarithmically scaled.

Images were like RGB, YUV. Like, we have so many different kinds of parameterizations. We have formalized these two senses ridiculously well. Um, touch and smell, nada. We're like, we're like where we were with images in, say, in 1920 or maybe even the 1800s, right?

That's where we're at, and Alex has this incredible vision of, like, having a smell sensor just eventually just be part of your daily life. Like, as of today, you don't really think about, like, when you're watching an Instagram reel or something, "Huh, like, I also would love to know what it smelled like," you know, when you're watching a reel of a food or-

Swyx1:31:31

Mm-hmm

Soumith Chintala1:31:31

... or something. You don't because we really haven't, as a society, got that muscle to even understand what a smell sensor can do. I think the more near-term effects are obviously going to be around m- things that provide more obvious, uh, utility in the short term, like maybe smelling cancer or, like, repelling mosquitoes better or, you know, stuff like that.

Swyx1:31:57

Yeah. Yeah. More recently, he's been talking about, like, categorizing perfumes, obviously-

Soumith Chintala1:32:00

Yeah, exactly

Swyx1:32:01

... 'cause that's a, that's a market that you can pursue.

Soumith Chintala1:32:02

Yeah. Like, I mean, think about how you can customize a perfume to your own liking in the same way you can customize a shoe or something, right? Um, so that-- Like, that's how I think all the near-term stuff, I think if, um, he's able to figure out a near-term value for it, they as a company can sustain themselves to then eventually, like, try to make progress on the long term, which is really an unchartered territory.

Swyx1:32:32

Mm-hmm.

Soumith Chintala1:32:33

Like, think about it. Fifty years from now, it would be pretty obvious to, like, kids of the generation to just, like, you know, I guess I was saying s- I was gonna say scroll a reel on their phone.

Swyx1:32:44

Mm-hmm. Smell it.

Soumith Chintala1:32:44

Maybe phones wouldn't be there. They're just, like, you know, on their glasses, they're watching something.

Swyx1:32:50

Yeah. I think VR would be-

Soumith Chintala1:32:51

And then, like, they immediately get, like, a smell sense off that remote experience as well. Like, we, we haven't really progressed enough in that dimension, um, and I think they have a chance to do it.

Swyx1:33:06

Awesome.

Alessio1:33:07

Um, awesome. I mean, we touched on a lot of things. Anything we're, we're missing? Anything you wanna direct people to or...

Swyx1:33:13

Yeah, call to action.

Alessio1:33:15

Yeah.

Swyx1:33:15

Call to-- call for research, call for startups.

Soumith Chintala1:33:18

I don't really have a lot of calls to action because usually I think intrin- like, people should be intrinsically, like, figuring it out.

Alessio1:33:25

That's a good-

Swyx1:33:26

Look inside yourself.

Soumith Chintala1:33:27

Yeah.

Alessio1:33:29

That's good. Um, awesome. Thank you so much for coming on.

Soumith Chintala1:33:32

Yeah, for sure.

Alessio1:33:32

This was great.

Swyx1:33:33

Thanks, Soumith.