LALatent SpaceNov 25, 2025· 1:00:39

After LLMs: Spatial Intelligence and World Models — Fei-Fei Li & Justin Johnson, World Labs

World Labs co-founders Fei-Fei Li and Justin Johnson argue that spatial intelligence is the next frontier beyond LLMs, with their product Marble generating editable 3D worlds from text or images using Gaussian splats. They claim language is a lossy channel for describing the rich 3D/4D world, while spatial intelligence handles everything from picking up a mug to inferring DNA's double helix. Marble already serves gaming, VFX, and film, with precise camera control and real-time rendering on phones and VR headsets. They discuss scaling compute from AlexNet to today's million-fold increase, why world models can 'soak up' modern GPU clusters better than language alone, and the gap between pattern-fitting and causal understanding in physics. Fei-Fei worries more about under-resourced academia than open vs. closed models, advocating for national AI compute clouds and open benchmarks like the BEHAVIOR dataset. Justin notes transformers are natively set models, not sequence models, opening architectural possibilities for new spatial data.

  1. 0:00Intro
  2. 1:29Origins
  3. 4:54Open Science
  4. 11:30Wacky Ideas
  5. 13:55Dense Captioning
  6. 21:02Spatial Intelligence
  7. 29:57Marble
  8. 38:49Use Cases
  9. 42:22Deep Dive
  10. 58:06Outro

Powered by PodHood

Transcript

Intro0:00

Justin Johnson0:00

I think the whole history of deep learning is, in some sense, the, the history of scaling up compute.

Fei-Fei Li0:04

When I graduated from grad school, I really thought the rest of my entire career would be towards solving that single problem, which is A lot of AI as a field, as a discipline, is inspired by human intelligence. We thought we were the first people doing it.

It turned out that was also simultaneously doing it.

Justin Johnson0:27

So Marble, like, basically, one way of looking at it, it's the system-- it's a generative model of 3D worlds, right? So you can input things like text or image or multiple images, and it will generate for you a 3D world that kind of matches those inputs.

So while Marble is simultaneously a world model that is building towards this vision of spatial intelligence, it was also very intentionally designed to be a thing that people could find useful today. Um, and we're see- starting to see emerging use cases, um, for in gaming, in VFX, um, in, in film, where I think there's a lot of really interesting stuff that Marble can do today as a product, and then also set a foundation for the, for the, for the grand world models that we want to build going into the future.

Alessio1:12

Hey, everyone. Welcome to the Late in Space Podcast. This is Alessio, founder of Kernel Labs, and I'm joined by Swyx, editor of Late in Space.

Swyx1:19

And we are so excited to be in the studio with Fei-Fei and Justin of, uh, World Labs. Welcome.

Fei-Fei Li1:24

We're excited too.

Swyx1:26

I nearly said Marble.

Justin Johnson1:27

Yeah, thanks for having us.

Fei-Fei Li1:29

Yeah.

Swyx1:29

I think there's a lot of interest in world models, and you've done, you've done a little bit of publicity around spatial intelligence and all that. Um, I guess maybe one of the part of story that is a real opportunity for you to tell is how you two came together, uh, to start building World Labs.

Origins1:29

Fei-Fei Li1:43

That's very easy because-

Swyx1:45

Yeah.

Fei-Fei Li1:45

... Justin was my former student.

Swyx1:46

Yeah.

Fei-Fei Li1:47

So Justin came to my-- I, you know, in my... The other hat I wear is a professor of computer science at Stanford. Justin joined my lab when? Which year?

Justin Johnson1:57

Uh, twenty twelve. Actually, the, the semester that I-- the quarter that I joined your lab was the same quarter that, that, uh, AlexNet came out.

Fei-Fei Li2:03

Yeah.

Swyx2:04

Mm.

Fei-Fei Li2:04

Yeah. So Justin is, uh, my first-

Swyx2:06

Were you involved in the whole announcement, uh, drama, I guess?

Justin Johnson2:08

No, no, not at all. But I was, uh, sort of watching all the ImageNet excitement around AlexNet at that, that quarter.

Fei-Fei Li2:14

So he was my-- one of my very best students. And, uh, and then he went on to have a very successful, uh, early career as a professor in Michigan, University of Michigan, Ann Arbor in Meta. And then when we, um, I think around, you know, more than two years ago, for sure, I think both independently, both of us have been looking at the development of the large models and thinking about what's beyond language models and, and this idea of building world models, spatial intelligence, uh, really was natural for us.

So we started talking and decided that we should just put all the eggs in one basket and focus on solving this problem and started World Labs together.

Justin Johnson3:03

Yeah, pretty much. I mean, like I-- after that, seeing that kind of ImageNet era during my PhD, um, I had the sense that the next sort of decade of computer vision was gonna be about getting, getting AI out of the, out of the data center and out into the world.

Um, so a lot of my interests post PhD kinda shifted in, uh, into, into 3D vision, a little bit more in com-- into computer graphics, uh, more into generative modeling. Um, and I was, uh-- I thought I was kind of drifting away from my advisor post PhD, but then when we reunited a couple of years later, it turned out she was thinking of very similar things.

Alessio3:32

So if you think about AlexNet, the core pieces of it were obviously ImageNet. It was the move to GPUs and neural networks. How do you think about the AlexNet equivalent model for world models? In a way, it's an idea that has been out there, right?

There's been, you know, Yann LeCun is maybe like the most-- the biggest proponent, most prominent of it. What have you seen in the last two years that you were like, "Hey, now is the time to do this," and what are maybe the things, um, fundamentally that you wanna build as far as data and kinda like maybe different types of, uh, algorithms or approaches to compute, uh, to make world models really come to life?

Justin Johnson4:05

Yeah, I, I think one is just there is a lot more data and compute generally available. Um, I think the whole history of deep learning is, in some sense, the, the history of scaling up compute. Um, and if you think about, you know, AlexNet required this jump from CPUs to GPUs, but even from AlexNet to today, we're getting about a thousand times more performance per card, um, than we had on AlexNet days.

And now it's common to train, to train models not just on one GPU, but on hundreds or thousands or tens of thousands or even more. So the amount of compute that we can marshal today on, on a single model is, is, you know, about a million fold more than we could have even at the start of my PhD.

So I think language was one of the really im-- one of the really interesting things that started to work quite well the last couple of years. But as we think about moving towards visual data and spatial data and world data, you just need to process a lot more, and I think that's gonna be, um, a good way to soak up this, uh, this new compute that's coming online more and more.

Swyx4:54

Does the model of having a public challenge still work, or should it be centralized with inside of a lab?

Open Science4:54

Fei-Fei Li5:00

I think open science still, uh, is important. You know, AI obviously compared to, um, the ImageNet, AlexNet time has really evolved, right? That was such a niche computer science discipline. Now it's just like civilizational, um, technology. But I'll give you an example, right?

Uh, recently, my Stanford, uh, lab just announced a open, uh, dataset and benchmark called, uh, Behavior, which is for benchmarking robotic learning in, uh, simulated environments. And that is a very clear effort in still keeping up this open science model of doing things, especially as, uh, in academia.

But I think it's important to recognize the ecosystem is, uh, is a mixture, right? I, I think, uh, uh, a lot of the, the very focused work in industry, some of them are more seeing the daylight in the form of a product.

Rather than an open challenge per se.

Swyx6:06

Yeah, and, and that's just a matter of the, like, funding and the business model. Like, you have to see some ROI from it.

Fei-Fei Li6:12

I think it's just a matter of the diversity of the ecosystem, right? Even during the so, so-called AlexNet ImageNet time, I mean, there were closed models, there were proprietary models, there were open models, you know. Um, you know, if...

Or you think about iOS versus Android, right? They're different business model. I, I wouldn't say it's just a matter of funding per se, it's just how the, the market is. They're different plays.

Swyx6:41

Yeah.

Alessio6:41

But do you feel like you could redo ImageNet today with the commercial pressure that some of these labs has? I mean, to me, that's, like, the biggest question, right? It's like, what can you open versus what do-- should you keep inside?

Like, you know, if I put myself in your shoes, right, it's like you raise a lot of money, you're building all of this. If you had the best data set for this, what incentives do you really have to publish it?

And it feels like the people at the labs are getting more and more pulled, uh, in the, um, PhD programs are getting pulled earlier and earlier into these labs. So I'm curious if you think there's, like, an issue right now with, like, how much money is at stake and how much pressure it puts on, like, the more academia open research space, or if you feel like that's not really a, a concern.

Fei-Fei Li7:23

I do have concerns about, less about the pressure, it's more about the resourcing and the imbalance, the resourcing of academia. This is a little bit of a different conversation from World Labs, you know. I have been, the past few years, advocating for, uh, uh, resourcing the healthy ecosystem, you know.

As the, um, founding director, co-director of Stanford's, uh, Institute for Human-Centered AI, Stanford HAI, I've been, you know, working with policy, uh, makers about, uh, resourcing public sector and academic, uh, AI work, right? We work with the first Trump administration on, uh, this bill called National AI Research, uh, Resource, NAIR Bill, which is, uh, scoping out a national AI compute cloud as well as data repository, and I also think that open source, open data sets continue to be important part of the ecosystem.

Like I said, right now, in my Stanford lab, we are doing the open data set, open benchmark on robotic learning called behavior, and many of my colleagues are still doing that. I think that's part of the ecosystem. I think what the industry is doing, some startups are doing, are running fast with models, cr- creating products, is also a good thing.

For example, when Justin was a PhD student with me, none of the computer vision programs worked that well, right? We could write beautiful papers. Justin has some beautiful-

Justin Johnson9:01

I mean, like, I- actually even, even before grad school, like, I wanted to do computer vision, and I reached out to a team at Google and, like, wanted to, I-- potentially go and try to do computer vision, like, out of, out of undergrad.

And they told me, like, "What are, what are you talking about?" Like, "You can't do that. Like, go do a PhD first and come back."

Swyx9:15

What, what was the motivation that, that got you so interested-

Justin Johnson9:18

Oh, 'cause I, I had done some computer vision research in, uh, during my undergrad with, uh, with actually Fei-Fei's PhD advisor.

Fei-Fei Li9:23

There's a lineage here.

Swyx9:24

Ah.

Justin Johnson9:24

Yeah, there's a lineage here.

Fei-Fei Li9:25

Yeah.

Swyx9:26

Wow.

Justin Johnson9:26

Um, so I had to-- I'd done some computer vision even as an undergrad, and I thought it was really cool, and I wanted to keep doing it. Um, so then I, I was sort of faced with this sort of industry academia choice even coming out of undergrad that I think a lot of people in the research community are facing now.

Um, but, but to your question, I, I think, like, the role of, of academia, especially in AI, has shifted quite a lot in the last decade. Um, and it's not a bad thing. Um, it's, it's a sense of-- It's, it's because the technology has, has grown and emerged, right?

Like, five or ten years ago, you really could train state-of-the-art models in the lab, um, even with just, with just a couple GPUs. But, you know, because that technology was so successful and scaled up so much, then you, you can't train state-of-the-art models with a couple GPUs anymore.

And that's not a bad thing. It's a good thing. It means the technology actually worked. Um, but that means the, the expectations around what we should be doing as academics shifts a little bit, and it shouldn't be about trying to train the biggest model and scaling up the biggest thing.

It should be about trying wacky ideas and new ideas and crazy ideas, um, most of which won't work. And I think there's a lot to be done there. And if anything, I'm worried that too many people in academia are hyper-focused on this notion of trying to pretend like we can train the biggest models or, or treating it as almost a vocational training program to then graduate and go to a big lab and then be able to play with all the GPUs.

I think there's just so much crazy stuff you can do around, like, new algorithms, new architectures, like, new systems that, you know, there's a lot you can do as, as one person.

Fei-Fei Li10:46

And also just, um, academia has a role to play in understanding the the- theoretical underpinning of these large models. We still know so little about this. Or extend to the interdisciplinary, you know, Justin calls wacky ideas. There's a lot of, uh, basic science ideas.

There's a lot of blue sky problems. So I agree. I don't think the problem is open versus closed, productization versus, uh, open sourcing. I think the problem right now is that academia by itself is severely under-resourced so that, uh, you know, the, the researchers and the, uh, students do not have enough resources to try these, uh, these, uh, ideas.

Wacky Ideas11:30

Swyx11:30

Yeah. Just for people to nerd snipe, uh, what's a wacky idea that comes to mind when you talk about wacky ideas?

Justin Johnson11:36

Oh, like, I had, I had this idea that I kept pitching to my students, uh, at, at Michigan, which is that I, I really like hardware, and I really like, like, new kinds of hardware coming online. Um, and in some sense, the, the emergence of the neural networks that you-- we use today and transformers are really based around matrix multiplication because matrix multiplication fits really well with GPUs.

But if we think about how GPUs are gonna scale, how, how hardware is likely to scale in the future, I don't think the current system that we have, like the GPU, like, hardware design, is gonna scale infinitely. And that we start to see that even now, that, like, the unit of compute is not, not the single device anymore, it's this whole cluster of devices.

So if you imagine-

Swyx12:11

The node.

Justin Johnson12:11

Yeah, it's a whole node or a whole cluster. But the way we talk about neural networks is still as if they are a monolithic thing that could be coded, like, in one GPU in PyTorch. Um, but then in practice, they get distributed over thousands of devices.

So are there, like, dras- just as, you know, transformers are based around MatMul, and MatMul is sort of the primitive that works really well on, on GPUs. As you imagine hardware scaling out, are there other primitives that make more sense for large-scale distributed systems that we could build our neural networks on?

Um, and I think it's possible that there could be drastically different architectures that fit with the next generation or, like, the, the, the hardware that's gonna come ten or twenty years down the line, and we could start imagining that today.

Swyx12:46

It's really hard to make those kinds of bets because there's also the concept that the hardware lottery, where let's just say, you know, NVIDIA has won, and we should just, you know, scale that out in infinity and write software to patch up any pl- any gaps we have in the, in, in the mix, right?

Justin Johnson13:00

I mean, yes. Yes and no. Like, if you look at the, if you look at the numbers, like, even going from Hopper to Blackwell, like the performance per watt is about the same.

Swyx13:07

Yes.

Justin Johnson13:07

Um, they mostly make the ch- the number of transistors go up, and they make the chip size go up, and they mo- make the, the power usage go up. But even from Hopper to Blackwell, we're kind of already seeing like a, a scaling limit in terms of what is the, what is the performance per watt that we can get.

So, um, that I think, I think there are, there is room to, to do something new. And I don't know exactly what it is-

Swyx13:23

Exactly

Justin Johnson13:23

... and I don't think you can get it done, like, in a three-month cycle as a startup. But I think that's the kind of idea that if you sit down and sit with for a couple years, like, maybe you could come up, come up with some breakthroughs, and I think that's the kind of long-range stuff that is a, is a perfect match for academia.

Swyx13:36

Coming back to the little bit of, of background and history, we have this sort of research note on the scene storytelling work that you did or neural, neural image captioning, uh, that you did with Andre. And I just wanted to hear you guys tell that story about, you know, you, you were, like, sort of embarking on that for your PhD and, and Fei-Fei, you, you, like, having that reaction that you had.

Dense Captioning13:55

Fei-Fei Li13:55

Yeah. So I think that line of work started between me and Andre, and then Justin joined, right? So, um, Andre started his PhD. He and I were looking at what is beyond ImageNet object recognition. And at that time, we, you know, convolution on neural network was, uh, has proven some power in, uh, ImageNet tasks, so, so ConvNet is a great way to represent images.

In the meantime, I think in the language that the space, a, a early sequential model is called LSTM, was also being experimented. So Andre and I were just talking about, this has been a long-term dream of mine. I thought it would take a hundred years to, to solve, which is telling the story of images.

When I graduated from grad school, I really thought the rest of my entire career would be towards solving that single problem, which is given a picture or given a scene, tell the story in natural language. But things evolved so fast.

Uh, when Andre and I, uh, when Andre started, we're like maybe combining the representation of convolution on neural network as well as the, the, the language, uh, sequential model of LSTM. We, we might be able to learn, uh, through training to match, uh, um, caption with images.

So that's when we started that line of work, and I, I don't re- it was 2014 or 2015.

Justin Johnson15:28

It was, uh, CVPR 2015 was the-

Fei-Fei Li15:30

Right

Justin Johnson15:30

... captioning paper.

Fei-Fei Li15:31

So it was, uh, our first paper that, uh, uh, Andre got it to work that was, you know, given an image. The image is represented with ConvNet. The language model is the LSTM model, and then we combine it, and it's able to generate one sentence.

And that was one of the first time. It was pretty-- I, I think I wrote it in my book. Uh, we thought, uh, we were the first people doing it. It turned out that Google at that time was also simultaneously doing it.

And, uh, a reporter, it was John Markoff from, uh, uh, New York Times, was breaking the Google story, but he by accident heard about us, and then he realized that we really independently got there together at the same time.

So he wrote the story of both the Google research as well as Andre and my research. But after that, I think Justin was already in the lab at that time.

Justin Johnson16:27

Yeah, yeah. I remember the group read- the, the group meeting where Andre was presenting some of those results and explaining this new thing called LSTMs and RNNs that I had never heard of before. And I thought like, "Wow, this is really amazing stuff.

I wanna, I wanna work on that." So then he had the paper at 20- CVPR 2015, um, the first image captioning results. Then after that, we started working together, and we did a-- we first we did a paper actually just on language modeling-

Fei-Fei Li16:48

Yeah

Justin Johnson16:48

... um, back in, uh, 2015. ICLR 2015.

Fei-Fei Li16:52

Yeah.

Justin Johnson16:52

Yeah. Yeah, I, I should have stuck with language modeling. That turned out- ... to be pretty lucrative in retrospect. But we did this language modeling paper together, um, me and Andre, uh, in, uh, 2015, where it was, like, really cool.

We trained these little R- these little RNN language models that could, you know, spit out a couple sentences at a time and poke at them and try to understand what the neurons inside the neural net- inside the networks were doing.

Fei-Fei Li17:12

Yeah, I remember you guys were doing analysis on the different, like, memory and-

Justin Johnson17:15

Yeah, yeah. It was, it was really-

Fei-Fei Li17:16

... and all that

Justin Johnson17:16

... it was really cool.

Fei-Fei Li17:17

Yeah.

Justin Johnson17:17

And even at that time, we had these results where you could, like, look inside the LSTM and say, like, "Oh, this thing is reading code." So one of the, like, one of the datasets that we, that we trained on for this one was the, the, um, the Linux source code, right?

'Cause the whole, the whole, the whole thing is, you know, open source, and you could just download this. So we trained an RNN on this, uh, on this dataset, and then as the network is trying to predict the tokens there, then, you know, try to correlate the kinds of predictions that it's making with the kind of internal structures, uh, in the RNN.

And there we were able to find some correlations between, oh, like, this unit and this layer of the LSTM fires when there's an open paren and then, like, turns off when there's a closed paren, um, and try to do some empirical stuff like that to, to figure it out.

So that was pretty cool, and that was just, like, di- that was sort of like cutting out the CNN from this, uh, language modeling part and just looking at the language models in isolation.

Fei-Fei Li18:02

But then we wanted to extend the, the image captioning work. Yeah, I remember at that time, we even have a sense of space because we feel like captioning does not capture different parts of the image.

Justin Johnson18:14

Right.

Fei-Fei Li18:15

So I was talking to Justin and Andre about Can you go what we, what we end up calling dense captioning, which is, you know, uh, describe the scene in greater details, especially different parts of the scene. So that's-

Justin Johnson18:29

Yeah, and so, so then, then we built a system. So then it was, um, me and Andre and, and Fei-Fei on a paper the following year, CVPR s- 2016, where we built this system that did dense captioning. So you input a single image, and then it would draw boxes around all the interesting stuff in the image and then write a short snippet about each of them.

It's like, "Oh, it's a green bot- water bottle on the table. It's a person wearing a black shirt." And this was a, a really complicated neural network because, um, that was built on a lot of, um, advancements that had been made in object detection around that time, which was a major topic in computer vision for a, for a long time.

And then it was actually, like, one joint neural network that was both, uh, you know, learning to look at individual images... 'Cause they actually, they actually had, like, then three different representations inside this network. One was the representation of the whole image to kinda get the gestalt of what's going on, then it would propose individual regions that it wants to focus on, and then look at, you know, represent each region independently, and then once you look at the region, then you need to spit out text for each region.

So that was a pretty complicated neural network architecture. Um, this was all pre-PyTorch, right?

Alessio19:22

And does it do it in one pass?

Justin Johnson19:24

Yeah, yeah. So it was a single forward pass that did all of that.

Fei-Fei Li19:25

Not only it was doing it in one pass, y- you also optimized inference. Uh, you're doing it on a webcam, I remember.

Justin Johnson19:32

Yeah, yeah.

Fei-Fei Li19:33

Yeah.

Justin Johnson19:33

Yeah, so I, I had built this, like, crazy real-time demo, um, where I had the network running, like, on a server at, at Stanford, and then a web front end that would stream from a webcam and then, like, send the image back to the server.

The server would run the model and stream the predictions back. So I was just, like, walking around the lab with this laptop-

Alessio19:49

Wow

Justin Johnson19:49

... that would just, like, show people this, uh, this, like, this, this network run in real time.

Alessio19:54

And it had identification and labeling as well on it?

Justin Johnson19:55

Yeah, yeah.

Fei-Fei Li19:56

Yes, it was-

Alessio19:56

Oh my God

Fei-Fei Li19:57

... uh, it was, uh, it was pretty impressive 'cause most of our graduate students would be satisfied if they can publish the paper, right? They, they package the, the, the research, put it in a paper, but Justin went a step further.

He's like, "I wanna do this real time web demo." Uh-

Justin Johnson20:12

Well, ac- actually, I don't, I don't know if I had told you this story, but then, um, we had a... There was a conference that year in Santiago, uh, at ICCV fif- it was ICCV 15. Um, and then, like, I had a paper at that conference for something different.

But I had my, my laptop, I was, like, walking around the conference with my laptop showing everybody this, like, real-time captioning demo, and the model was running on a server in California.

Fei-Fei Li20:31

Nice.

Justin Johnson20:31

So it was, like, actually able to stream, like, all the way from California down to Santiago.

Alessio20:35

What latency does-

Justin Johnson20:36

Oh, it, it was terrible. It was like-

Alessio20:38

Right, right.

Justin Johnson20:38

It was terrible.

Alessio20:38

It was delayed a lot.

Justin Johnson20:39

It was, like, one FPS.

Alessio20:40

Right.

Justin Johnson20:40

But the fact that it worked at all was pretty, was pretty amazing.

Alessio20:42

So I was gonna briefly quip that, you know, maybe vision and language modeling are not that different. You know, DeepSeek CR recently, uh, tried the crazy thing of let's language, let's model text from pixels and, and just, like, train on that.

And, uh, it might be the future. I don't know. I don't know if you guys have any takes on whether language is actually necessary at all.

Spatial Intelligence21:02

Fei-Fei Li21:02

I just wrote a whole manifesto on spatial intelligence.

Alessio21:05

Yeah. This is my segue into this.

Fei-Fei Li21:07

Yeah.

Alessio21:08

Yes.

Fei-Fei Li21:09

I think they are different. Um, I do think the architecture, uh, of, uh, these generative models will share a lot of, uh, shareable components, but, uh, I think the deeply 3D, 4D spatial world has a level of structure that is fundamentally different from a purely generative, uh, signal that is one-dimensional.

Justin Johnson21:35

Yeah, I, I think there's something to be said for pixel maximalism, right? Like, there's this notion that language is this different thing, but you-- we see language with our eyes, and our eyes are just, like, you know, basically pixels, right?

Like, we've got sort of biological pixels in the back of our eyes that are processing these things, and, you know, we see text and we think of it as this discrete thing, but that really only exists in our minds.

Like, the physical manifestation of text and language in our world are, you know, physical objects that are printed on things in the world, and we see it with our eyes.

Fei-Fei Li22:03

Well, you can also think it's sound, but even sound-

Justin Johnson22:05

Oh, sure, sure, sure.

Fei-Fei Li22:06

Even sound, you can translate into a cor-

Alessio22:07

You can visualize sound.

Fei-Fei Li22:08

Yeah, you get correlogram, which is a 2D signal.

Justin Johnson22:11

Right. A- and then, like, you actually lose something if you translate to this, like, purely tokenized representations that we use in LLMs, right? Like, you lose the font, you lose l- you, you lose the line breaks, you lose sort of the 2D arrangement on the page.

Um, and, and for a lot of cases, for a lot of things, maybe that doesn't matter. Um, but for some things it does. Um, and I, I think pixels are this sort of l- more, more lossless representation of what's going on in the world, and in, in some ways, a more general rep- general representation that more matches what, what we, what we humans see as we, as we navigate the world.

So, so, like, uh, there's an efficiency argument to be made. Like, maybe it's not super efficient to, like, you know, render your text to an image and then feed that to a vision model.

Alessio22:47

That's exactly what DeepSeek did, right?

Justin Johnson22:48

Yeah.

Alessio22:48

And it, it was, like, kinda worked.

Justin Johnson22:50

Yeah.

Alessio22:52

I think this ties into the whole world model. Like, one of the-- my favorite papers that I saw this year was about inductive bias to pro- for world models. So it was a Harvard paper where they fed a lot of, like, orbital, uh, patterns into an LLM, and then they asked the LLM to predict the orbit-

Justin Johnson23:07

Oh, I know that one

Alessio23:07

... of a planet around the sun, and the model generated looked good, but then if you asked it to draw the force vectors, it would be all wacky. You know-

Justin Johnson23:16

Mm-hmm

Alessio23:16

... it wouldn't actually follow it. So how do you think about what's embedded into the data that you get? And we can talk about maybe tokenizing for 3D world models. Like, what are, like, the dimensions of information? There's the visual, but, like, how much of, like, the underlying hidden forces, so to speak, you need to extract out of this data and, like, what are some of the challenges there?

Justin Johnson23:39

Yeah, I, I think there's different ways you could approach that problem. Um, one is, like, you could try to be explicit about it and say, like, "Oh, I want to, you know, measure all the forces and feed those as training data to your model," right?

Then you could, like, sort of run a traditional physics simulation and, you know, then know all the forces in the scene, and then use those as, as training data to train a model that's now gonna hopefully predict those.

Or you could hope that something emerges more latently, right? That you kind of train on something end to end, and then on, on a more general problem, and then hope that somewhere some- something in, in the internals of the model must learn to model something like physics in order to make the proper predictions.

Um, and those are kind of the two big paradigms that we have more generally.

Fei-Fei Li24:16

But there's no indication that th- those latent, uh, uh, modeling Will get you to a causal law of, uh-

Justin Johnson24:24

Right

Fei-Fei Li24:24

... of space and dynamics, right? That's where today's deep learning and, uh, human intelligence actually start to bifurcate 'cause fundamentally, the deep learning is still fitting patterns.

Justin Johnson24:36

There you sort of get philosophical- ... and you say that, like, we're trying to fit patterns, too, but maybe we're trying to fit, you know, a more broad array of patterns, like, over a, o- with, with a longer time horizon, a different reward function.

Um, but, but, like, basically, the, the paper you mentioned is sort of, you know, that problem, that it learns to fit the specific patterns of orbits, but then it doesn't actually generalize in the way that you'd like. It doesn't have a sort of causal model of gravity.

Alessio24:56

Right, because even in Marble, you know, I was trying it, and it generates these beautiful sceneries, and there's, like, arches in them, but does the model actually understand how- ... you know, the arch is actually, you know, drawing on the center kind of, like, stone and, like, you know, the actual physical structure of it?

And the other question is, like, does it matter that it does understand it as long as it always renders something that would fit the physical model that we imagine?

Fei-Fei Li25:22

If you use the word understand the way you understand, I'm pretty sure the model doesn't understand it. The model is learning from the data, learning from the pattern. Um, yeah, does it matter, especially for the use cases for...

It, it's a good question, right? Like, for now, I don't think it matters because it renders out what you need-

Alessio25:45

Right

Fei-Fei Li25:45

... assuming it's perfect.

Justin Johnson25:47

Yeah, I mean, it depends on the use case. Like, if the use case is I want to generate sort of a backdrop for, for virtual film or production or something like that, all you need is something that looks plausible, and in that case, probably it doesn't matter.

But if you're gonna use this to, like, you know, if you're an architect and you're gonna use this to design a building that you're then gonna go build in the real world- ... then yeah, it does matter that you model the forces correctly 'cause you don't want the thing to break when to actually, actually build it.

Fei-Fei Li26:08

But even there, right, like, even if your model has the semantics in it, let's say, I still don't think the understanding of the signal or the, or the, the output on the model part and the understanding on the human part is a different word, but this gets, again, philosophical.

Justin Johnson26:29

Yeah, I mean, there, there's this trick with understanding, right? Like, these models are a very different kind of intelligence than human intelligence. Um, and human intelligence is interesting because, you know, you know, I, I think that I understand things because I can introspect my own thought process to some extent, and then I believe that my thought process probably works similar to other people's, so that when I observe someone else's behavior, then I infer that their internal mental state is probably similar to my own internal mental state that I've observed.

And therefore, I know that I understand things, so there I assume that you understand something. But these models are sort of like this, this alien form of intelligence where they can do really interesting things. They can exhibit really interesting behavior.

But whatever kind of internal, the equivalent of internal cognition or internal self-reflection that they have, if it exists at all, is totally different from what we do. So it doesn't-

Fei-Fei Li27:14

It doesn't have the self-awareness.

Justin Johnson27:16

Right.

Fei-Fei Li27:16

Yeah.

Justin Johnson27:16

But what that means is that when, when we observe seemingly interesting or intelligent behavior out of these systems, we can't necessarily infer other things about them because their, their model of the world and the way they think is so different from us.

Alessio27:28

So would you need two different models to do the visual one and the architectural generation, you think, eventually? Like, there's not anything fundamental about the approach that you've taken on the model building. It's more about scaling the model and the capabilities of it.

Or, like, is there something about being very visual that prohibits you from actually learning the physics behind this, so to speak, so that you could trust it to generate a CAD design that then is actually gonna work in the real world?

Fei-Fei Li27:58

I think this is a matter of scaling data and, and, and bettering model. I don't think there's anything fundamental that separates these two.

Justin Johnson28:06

Yeah, I would like it to be one model, but I think, like, the big problem in deep learning in some sense is how do you get emergent capabilities beyond your training data? Are you gonna get something that understands the forces while it wasn't trained to predict the forces, but it's gonna learn them implicitly internally?

Um, and I think a lot of what we've seen in other large models is that a lot of this emergent behavior does happen at scale. Um, and will that transfer to other modalities and other use cases and other tasks?

Um, I hope so, but that'll, that'll be a, a process that we need to play out over time and see.

Swyx28:33

Is there a temptation to rely on, um, physics engines, um, that already exist out there that are, you know, basically the gaming industry has saved you a lot of this work or do we have to reinvent things for some fundamental mismatch?

Justin Johnson28:47

I think that's sort of like climbing the ladder of technology, right? Like, in some sense, the reason that you wanna build these things at all is because maybe traditional physics engines don't work in some situations. If a physics engine was perfect, we would have sort of no need to build models because the problem would've already been solved.

So in some sense, the reason why we want to do this is because classical ph- physics engines don't solve problems in the generality that we want. Um, but that doesn't mean we need to throw them away and start everything from scratch, right?

We can use traditional physics engines to generate data that we then train our models on, and then you're sort of distilling the physics engine into the weights of the neural network that you're training.

Swyx29:20

I, I think that's a lot of what if you compare the work of other labs, people are speculating that, you know, Sora had a little bit of that. Genie 3-

Fei-Fei Li29:29

Mm-hmm

Swyx29:29

... had a bit of that. Um, a-a-and G- Genie 3 is, like, explicitly like a video game.

Fei-Fei Li29:33

Mm-hmm.

Swyx29:33

Like, you have controls to-

Fei-Fei Li29:34

Yeah

Swyx29:34

... to walk around in. And I, I, I, I always think, like, it's really funny how the things that we invent for fun actually does eventually make it into serious work.

Fei-Fei Li29:43

Mm-hmm. Yeah. The whole AI revolution started by graphics, uh, chips-

Swyx29:48

Yeah

Fei-Fei Li29:48

... partially.

Swyx29:51

Misusing the GPU for, uh, for generating a lot of triangles into generating a lot of, uh, everything else, basically.

Fei-Fei Li29:56

Yeah.

Swyx29:57

We touched on Marble a little bit. I think you guys chose Marble as, like, kind of your, like, your sort of a little bit coming out of stealth moment, if you can call it that.

Marble29:57

Fei-Fei Li30:03

Mm-hmm. Yeah.

Swyx30:04

Uh, maybe, uh, we can get a concise explanation from you on what people should take away because everyone here can try Marble, but I don't think they might be able to link it to- The differences between what your vision is versus other, uh, I guess generative worlds they may have seen, uh, from other labs.

Fei-Fei Li30:20

So Marble is a glimpse into our model, right? We are a model spatial intelligence model company. We believe spatial intelligence is the next frontier. Uh, in order to make spatial- spatially intelligent models, the model has to be very powerful in terms of its ability to, you know, understand, reason, generate in very multimodal, uh, fashion of worlds, as well as allow the level of interactivity that we eventually hope to be as, you know, complex as how humans can interact with the world.

So that's the grand vision of spatial intelligence, as well as the, the kind of world models we, we, uh, see. Marble is the first glimpse into that. It's the, it's the, the first part of that journey. It's the first-in-class model in the world that generates, uh, 3D worlds in this level of fidelity that is in the hands of the, the public.

It's the starting point, right? Uh, we actually wrote this, uh, tech blog. Justin spent a lot of time writing that p- tech blog. I don't know if you had time to browse it. Uh, I mean, Justin really broke it down into what are the inputs, uh, we can...

multimodal inputs of, uh, Marble, what are the kind of, uh, editability, which is, you know, allows user to be interactive with the model, and what are, uh, the kind of outputs we can have.

Justin Johnson31:43

Yeah. So, so Marble, like, basically, one way of looking at it, it's the system-- it's a generative model of 3D worlds, right? So you can input things like text or image or multiple images, and it will generate for you a 3D world that kind of matches those inputs.

And it's also interactive in the sense that, um, you can interactively edit scenes. Like, I could generate this scene and then say, "I don't like the water bottle. Make it blue instead." Like, take out the table, like, ma- change these microphones around.

And then you can generate new worlds based on these interactive edits and export in a variety of formats. And with Marble, we were actually trying to do sort of two things simultaneously, and I think we, we managed to pull off the balance pretty well.

Um, one is actually build a model that goes towards the grand vision of spatial intelligence. And models need to be able to understand lots of different kinds of inputs, need to be able to model worlds in a lot of situations, need to be able to model counterfactuals of how they could change over time.

So we wanted to start to build models that have these capabilities, and Marble, Marble today does already have hints of all of these. But at the same time, we, we're, we're a company, we're a business. We were really trying not to have this be a science project, but also build a product that would be, be useful to people in the real world today.

So while Marble is simultaneously a world model that is building towards this vision of spatial intelligence, it was also very intentionally designed to be a thing that people could find useful today. Um, and we're see- starting to see emerging use cases, um, for in gaming, in VFX, um, in, in film, where I think there's a lot of really interesting stuff that Marble can do today as a product, and then also set a foundation for the grand world models that we want to build going into the future.

Swyx33:12

Yeah. I noticed one tool that was very interesting was you can record your scene inside.

Justin Johnson33:16

Yes.

Fei-Fei Li33:16

Yes. It's very important. The ability to record means a very precise control of camera p- placement.

Swyx33:25

Okay.

Fei-Fei Li33:25

In order to have precise camera placement, it means you have to have a sense of 3D space. Otherwise, you don't know how to orient your camera, right, and how to move your camera. So that is a natural consequence of this kind of model, and, and this is why this is just one of the examples.

Swyx33:44

Yeah. I find when I play with video generative models, I'm having to learn the, the language of being a director-

Fei-Fei Li33:50

Yeah

Swyx33:50

... because I have to move them, like pan, uh, you know, like-

Fei-Fei Li33:53

Yeah

Swyx33:53

... dolly out.

Fei-Fei Li33:53

But even there, you cannot say pan 63 degrees- ... to the north, right? You just don't have that control. Whereas in Marble, you have precise control in terms of placing a camera.

Justin Johnson34:06

Yeah. I think that's one of the first things people need to understand is like, it's not-- you're not generating frame by frame-

Fei-Fei Li34:12

Yes

Justin Johnson34:12

... which is like what a lot of the other-

Fei-Fei Li34:13

Yeah

Justin Johnson34:13

... models are.

Fei-Fei Li34:14

Yeah.

Justin Johnson34:15

What are... You know, people understand an LLM generates one token. What are, like, the atomic units? There's kinda like, you know, the meshes, there's, like, the splats, the voxels. There's a lot of pieces in a 3D world. What should be the mental model that people have of, like, your generations?

Yeah. I, I think there's, like, what exists today and what could exist in the future.

Swyx34:33

Right.

Justin Johnson34:33

Um, so what exists today is the model natively outputs splats. Um, so Gaussian splats are these like, you know, each one is a tiny, tiny particle that's semi-transparent, has a position orientation in, in 3D space. Um, and the scene is built up from a large number of these Gaussian splats.

Um, and Gaussian splats are really cool because you can render them in re- in real time really efficiently, so you can render on your iPhone, render, render everything. Um, and that's, that's how we get that sort of precise camera control because the splats can be rendered real time on just on, on pretty much any client-side device that we want.

So for a lot of the scenes that we're generating today, uh, that kind of atomic unit is that individual splat. But I don't think that's fundamental. I, I could imagine other approaches in the future that would be interesting.

So there, like, there are other approaches that even we've worked on at World Labs, um, like our recent RTFM model, that does generate frames one at a time. Um, and there, the atomic unit is generating frames one at a time as the user interacts with the system.

Um, or you could imagine other architectures in the future where the atomic unit is a token, where that token now represents, you know, some chunk of the 3D world. Um, and I think there's a lot of different architectures that we can experiment with here over time.

Swyx35:34

I do wanna press on, double-click on this a little bit. The-- My version of what Alessio was gonna say was like what is the fundamental data structure of a world model? Because exactly like you said, like it's, it's either a Gaussian splat or it's like the frame or what have you.

Uh, you also in, in the, the sort of previous statements w- focus a lot on the physics and the forces, which is something over time, uh, which is loosely. I don't see that in Marble. I presume it's not there yet.

Uh, maybe if there was like a Marble two, it would-- you would have movement, or is there a modification to Gaussian splats that makes sense, or would it be something completely different?

Justin Johnson36:07

Yeah, I, I think there's a couple modifications that make sense, and there, there's actually a lot of interesting ways to integrate things here, which is another nice w- place of working in this space. Then there's actually been a lot of research work on this.

Like, when you talk about wacky ideas, like, there's actually been a lot of really interesting academic work on different ways to imbue physics into, into these-

Fei-Fei Li36:22

We can also do wacky ideas in industry.

Justin Johnson36:26

All right. But, but it's then it's like Gaussian splats are themselves little particles. There's been a lot of, um, approaches where you basically attach physical properties to those splats and say that each one has a mass or, like, maybe you treat each one as being coupled with some kind of, um, virtual spring to nearby neighbors, and now you can start to do sort of physics simulation on top of splats.

So one kind of avenue for adding, uh, adding physics or dynamics or interaction to these things would be to, you know, predict physical properties associated with each of your splat particles, um, and then simulate those downstream either using classical physics or something learned.

Or, you know, the, kind of the beauty of working in 3D is things compose and you can inject logic in different places. So one way is sort of like we're generating a 3D scene, we're gonna predict 3D properties of everything in the scene, then we use a classical physics engine to, to simulate the interaction.

Or you could do something where, like, as a result of a user action, the model is now going to regenerate the entire scene, um, in G- in, in splats or some other representation. Um, and that could potentially be a lot more general because then you're not bound to whatever sort of, um, you know, physical properties you know how to model already.

But that's also a lot more computationally demanding because then you'd need to regenerate the whole scene in response to the, to the user actions. But I, I think this is a, this was a, a really interesting area for, for future work and for, uh, adding on to, to, into pot- potential Marble 2, as you say.

Fei-Fei Li37:40

Yeah.

Guest37:41

Yeah.

Fei-Fei Li37:41

There's opportunity for dynamics-

Justin Johnson37:43

Yeah

Fei-Fei Li37:43

... right?

Guest37:44

What's the state of, like, splat's density, I guess? Like, do we, can we render enough to have very high resolution when we zoom in? Are we limited by, like, the amount that you can generate, the amount that we can render?

Like, how are these gonna get super high fidelity, so to speak?

Justin Johnson37:59

The-- You have some limitations, but depending on your target use case. So, like, one of the, one of the big constraints that we have on our scenes is we wanted things to render cleanly on mobile, um, and we wanted things to render cleanly in VR headsets.

So those are-- those devices have a lot less compute than your us- than you have in a lot of other situations. Um, and, like, if you want to get a splat file to render at high resolution, high, like, 30 to 60 FPS on, like, uh, an iPhone from four years ago, then you are a bit limited in, like, the number of splats that you can handle.

Um, but if you're allowed to, like, work on a recent, like, even this year's iPhone or, like, a recent MacBook or even if you have a local GPU or if you don't need, if you don't need that 60 FPS 1080p, like, then you can relax the constraints and, and get away with more splats, and that lets you get higher resolution in your scenes.

Guest38:44

One, uh, use case I was expecting but didn't hear from you was embodied use cases.

Fei-Fei Li38:49

Mm-hmm.

Use Cases38:49

Guest38:49

Are you f- you're just focusing on virtual for now?

Fei-Fei Li38:51

If you go to World Labs home- homepage-

Guest38:54

Yeah

Fei-Fei Li38:54

... there is a particular page called Marble Labs. There we showcase different use cases, and we actually organize them in more visual effect use cases or gaming use cases as well as simulation use cases. And in that we actually show this is a technology that can help a lot in robotic training, right?

This, uh, goes back to what I was talking about earlier in, uh, speaking of data starvation, uh, robotic training really lack data. You know, high fidelity real world data is absolutely very critical, but, uh, it's, you're just not gonna get a ton of that.

Of course, the other extreme is just purely internet video data, but then you lack a lot of the, the controllability that you want to train your, your embodied agents with. So simulation and synthetic data is actually a very important middle ground.

For that, I've been working in this space for many years, and one of the biggest pain point is where do you get the sy- synthetic simulated data? You have to curate assets and, and, and, and build these, uh, compose these complex situations, and in robotics you want a lot of different states.

You want the embodied, uh, agent to, to, uh, interact in the synthetic environment. Marble actually is a really potential for, uh, helping to generate these synthetic, uh, simulated worlds for embodied, uh, agent training.

Guest40:23

Obviously that's, yeah, that's on the, that's on the homepage.

Fei-Fei Li40:25

Yeah.

Guest40:25

It, it'll be there. I just, I was, like, trying to make the link to, as you said, like, you also have to build, like, a business model. The market for robotics obviously is very huge. Maybe you don't need that or maybe we need to build up and solve the virtual worlds first before we go to embodied, and obviously I think that's a-

Fei-Fei Li40:41

That is actually-

Guest40:41

... good stepping stone.

Fei-Fei Li40:42

That is to be decided. I, I-

Guest40:44

Yeah

Fei-Fei Li40:44

... do think that, uh, uh, uh, uh-

Guest40:47

Because everyone else is going straight there. Right?

Fei-Fei Li40:50

Not everyone else- ... but there is a, there is an excitement, I would say. But, you know, I think the world is big enough to, to-

Guest40:58

Yeah

Fei-Fei Li40:58

... have different, uh, um-

Guest41:00

Approaches.

Fei-Fei Li41:01

Yeah, approaches. Yeah.

Justin Johnson41:01

Yeah, I mean, and we always view this as a pretty horizontal technology that should be able to touch a lot of different industries over time. And, you know, Marble is a little bit more focused on creative industries for now, but I think the, the technology that powers it should be applicable to a lot of different things over time, and robotics is one that, you know, is maybe gonna happen sooner than later.

Fei-Fei Li41:18

Also design, right? It's very adjacent to creative.

Justin Johnson41:21

Oh, yeah, definitely. Like, I, I, I think it's like-

Guest41:23

The architecture stuff?

Fei-Fei Li41:24

Yes.

Justin Johnson41:24

Okay. Yeah, I mean, I, I w- I was joking online. I posted this, uh, this video on Slack of like, "Oh, who wants to use Marble to, to plan your next kitchen remodel?" It actually works great for this already.

Just, like, take two images of your kitchen, like, reconstruct it in Marble, and then use the editing features to see what would that space look like if you change the countertops or change the floors or change the cabinets.

And this is something that's, you know, we didn't necessarily build anything specific for this use case. But because it's a, it's a powerful horizontal technology, you kind of get these emergent use cases that, that just fall out of the model.

Fei-Fei Li41:52

We have early beta users using a, um, um, API key that, uh, is already building for interior design, uh, use case.

Guest42:02

I just did my garage. I should have known about this.

Fei-Fei Li42:05

I know.

Guest42:06

I coulda...

Fei-Fei Li42:07

Next time you remodel we can-

Guest42:09

Exactly

Fei-Fei Li42:09

... be of help.

Guest42:10

Well, kitchen is next, I'm sure.

Fei-Fei Li42:12

Yeah.

Guest42:12

Yeah, I'm curious about the whole spatial intelligence space. I think we should dig more into that. One, how do you define it, and, like, what are, like, the gaps between-

Alessio42:22

Traditional intelligence that people might think about, uh, LLMS. When, you know, Dario says, "We have a data center full of Einsteins," that's like traditional intelligence. It's not spatial intelligence. What is required to be spatially intelligent?

Deep Dive42:22

Fei-Fei Li42:37

First of all, I don't understand that sentence, a data center full of Einsteins. That-

Alessio42:42

Ooh.

Fei-Fei Li42:43

I, I just don't understand that. The, the, the-- it's not a deep, uh-

Alessio42:46

It's an analogy. It's an analogy.

Fei-Fei Li42:48

Yeah. Well, so a lot of AI as a field, as a discipline, is inspired by human intelligence, right? Because we are the most intelligent animal we know in the universe for now, and if you look at human intelligence, it's very multi-intelligent, right?

There is a psychologist, I think his name is Howard Gardner, in the 1960s actually literally called multiple intelligence to describe human intelligence, and there is linguistic intelligence, there's spatial intelligence, there is, uh, uh, logical intelligence, and, and, and emotional intelligence.

So for me, when I think about spatial intelligence, I see it as complementary to language intelligence. So I, I personally would not say it's spatial versus traditional because I don't know tradition means-- What does that mean? I do think spatial is complementary to linguistic.

And, uh, and, uh, how do we define spatial intelligence? It's, it's the capability that, uh, allows you to reason, understand, move, and interact in space. And I use this example of the deduction of DNA structure, right? Uh, and of course, I'm simplifying this, uh, story, but a lot of that had to do with the spatial reasoning of the molecules and the chemical bonds in a 3D space to, to eventually conjecture a double helix.

And that ability that humans or, uh, Francis Crick and Watson had, had done, it is very, very hard to reduce that process into pure language. And that's, that's a pinnacle of a civilizational moment. But every day, right, I'm here trying to grasp a, uh, a mug.

This whole process of seeing the mug, seeing the context where it is, seeing my own hand, opening of my hand that geometrically would match the mug, and touching the right affordance points, all this is deeply, deeply spatial. The-- It's very hard.

I'm trying to use language to narrate it, but on the other hand, that narrated language itself cannot get you to, to pick up a, a mug.

Justin Johnson45:09

It's like, yeah, bandwidth constraint.

Fei-Fei Li45:10

Yes.

Justin Johnson45:11

I did some math recently on, like, if you just spoke, uh, all day, every day for twenty-four hours a day, uh, how many tokens do you generate? At the average speaking rate of, like, 150 words per minute, it rough- roughly around, rounds out to about 215,000 tokens per day.

And, like, your, your, your, your world that you live in is so much higher bandwidth than that.

Alessio45:32

Well, I think that is true, but if I think about Sir Isaac Newton, right? It's like you have things like gravity at the time that have not been formalized in language that people inherently spatially understand, that things fall, right?

But then it's helpful to formalize that in some way or like, you know, all these different rules that we use language to, like, really capture something that empirically and spatially you can also understand, but it's easier to, like, describe in a way.

So I'm curious, like, the interplay of, like, spatial and, like, linguistic intelligence, which is like, okay, you need to understand some rules are easier to write in language for then the spatial intelligence to understand.

Fei-Fei Li46:12

Sure.

Alessio46:12

But you cannot, you know, you cannot write, put your hand like this and put it down this amount. So I'm always curious about how you leverage each other together.

Justin Johnson46:21

I mean, if, if anything, like the example of, of Newton, like Newton only thinks to write down those laws because he's had a lot of embodied experience-

Alessio46:27

Yeah. Right. Yeah, exactly

Justin Johnson46:28

...in the world watching things fall. And actually, it's useful to distinguish between the theory building that you're mentioning versus, like, the embodied, like, the daily experience of being embedded in the three-dimensional world, right? So, so, so to me, pa-- spatial intelligence is sort of encapsulating that embodied experience of being there in 3D space, moving through it, seeing it, actioning it.

And as Fei-Fei said, you can narrate those things, but it's a very lossy channel. It's, uh, just like the notion of, you know, being in the world and doing things in it is a very different modality from trying to describe it.

But because we as humans are animals who have evolved interacting in space all the time, like, we don't even think that that's in-- that's a hard thing, right? And then we sort of naturally leap to language and then theory building as mechanisms to abstract above that sort of native spatial understanding.

And in some sense, LLMs have just, like, jumped all the way to those highest forms of abstracted reasoning, which is very interesting and very useful. But spatial intelligence is almost like opening up that bu- black box again and saying, "Maybe we've lost something by going straight to that fully abstracted form of, of language and reasoning and communication."

Fei-Fei Li47:27

You know, it's funny as a, a vision scientist, right? I always find that vision is underappreciated because it's effortless for humans. You open your eyes as a baby, you start to see your world.

Justin Johnson47:40

Yeah, I was gonna say, we're somehow born with it.

Fei-Fei Li47:41

Right. We're almost born with it, but you have to put effort in learning language, including learning how to write, how to do grammar, how to express, and that makes it feel hard. Whereas something that nature spend way more time actually optimizing, which is perception and spatial intelligence, is underappreciated by humans.

Justin Johnson48:03

Is there proof that we are born with it? You said, you said almost born. So it sounds like-

Fei-Fei Li48:08

We-

Justin Johnson48:08

...we actually do learn after we're born.

Fei-Fei Li48:10

When we are born, our visual acuity is less. And our perceptual ability does increase, but we are, most humans are born with the ability to see, and most humans are born with the abil- ability to link perception with motor, um, movements, right?

I, I mean, the motor movement itself is, takes a while to, uh, refine, but, uh... And then animals are incredible, right? Like, I was just in Africa earlier this summer. These little animals, they're born, and within minutes-

Guest48:42

Mm-hmm

Fei-Fei Li48:42

... they have to get going, and otherwise the, you know, the, the, the lions will get them. And in nature, you know, it took 540 million years to, uh, optimize perception and spatial intelligence, and language, the most generous estimation of, uh, language development is probably half a million years.

Guest49:03

Wow.

Fei-Fei Li49:04

Yeah.

Guest49:04

That's longer than I would've gonna say.

Fei-Fei Li49:06

Oh, I'm being very generous.

Guest49:07

Yeah.

Fei-Fei Li49:07

Yeah.

Guest49:09

Yeah, no, I, I was, uh, you know, sort of going through your book, and, uh, I was really realizing, like, that one of the in- interesting links to something that we covered on the podcast is language model benchmarks and how VinoGrant, uh, actually put in all these, like, sort of physical impossibilities that require spatial intelligence, right?

Like, A is on top of B, therefore A cannot fall through B-

Fei-Fei Li49:27

Mm

Guest49:28

... is obvious to us, but to a language model it could happen. I don't know. Maybe it's, like, a part of the, you know, the next token prediction.

Justin Johnson49:35

And that's sort of what I mean about, like, unwrapping this abstraction.

Guest49:38

Yeah.

Justin Johnson49:38

Right? Like, if your whole model of the world is just, like, saying sequences of words after each other, it's really kind of hard to-

Fei-Fei Li49:44

Yeah

Justin Johnson49:44

... like, why, why not?

Guest49:45

It's actually unfair.

Justin Johnson49:46

Right, right.

Guest49:46

So it's-

Justin Johnson49:47

But then the reason it's obvious to us is because we are internally mapping it back to some three-dimensional representation of the world that we're familiar with.

Guest49:53

Mm-hmm. This, the question is, I guess, like, how hard is it, you know, how long is it gonna take us to distill from, like... I, I, I use the word distill, I don't know if you agree with that, to distill from your world models into a language model?

'Cause we do want our models to have spatial intelligence, right? Like, uh, and do we have to throw the, the language model out completely in order to, to do that? Or-

Fei-Fei Li50:12

No.

Guest50:12

No, right?

Justin Johnson50:13

Yeah, I don't think so.

Guest50:14

Right.

Fei-Fei Li50:14

I think they're multimodal. I mean, even our model, Marble, today takes language as a input.

Guest50:18

Right.

Fei-Fei Li50:19

Right. So it's deeply multi- multimodel, and I think in many use cases, these models will work together.

Guest50:26

Yeah.

Fei-Fei Li50:26

Maybe one day we'll have a universal model.

Justin Johnson50:28

I mean, even, even if you do, like, there, there's sort of a pragmatic thing where people use language, and people want to interact with systems using language. Even pragmatically, it's useful to build systems and build products and build models that let people talk to them, so I don't see that going away.

I think there's a sort of intellectual curiosity of, of saying how, like, intellectually how much could you build a model that, that only uses vision or only uses spatial intelligence? I don't know that that would be practically useful, but I think it'd be an interesting intellectual or academic exercise to see how far you could push that.

Guest50:58

I think, I mean, not to bring it back to physics, but I'm curious, like, if you had a highly precise world model and you didn't give it any notion of, like, our current understanding of the standard model of physics, how much of it it would be able to come up with and, like, recreate from scratch, and what level of, like, language understanding it would need.

Because we have so many notations that, like, we kinda use that, like, we created, but, like, maybe we'll come up with a very different model of it and still be accurate, and I wonder how much we're kinda limited by the high...

You know how people say humans always need to be like humans because the world is built for humans.

Justin Johnson51:32

Mm-hmm.

Guest51:33

And in a way, it's like the way we build language constrains some of the outputs that we can get from these other modalities as well. So I'm super excited to follow your work.

Justin Johnson51:42

Yeah, I, I mean, like, there's another in- I mean, you actually don't even need to be doing AI to answer that question. You could discover aliens and see what kind of physics they have.

Guest51:48

Right.

Justin Johnson51:49

Right? And they might have a totally different-

Guest51:50

Well, Fei-Fei said we are so far the smartest-

Justin Johnson51:53

So, so far

Guest51:53

... animal in the universe.

Fei-Fei Li51:54

So far.

Justin Johnson51:54

Right. So, so what do you... So I mean, but that is a really interesting question, right? Like, is our knowledge of the universe and our understanding of physics, is it constrained in some way by our own cognition or by the path dependence of our own technological evolution?

And one way to sort of an- and, like, do an experiment, like, you almost wanna do an experiment and say, like, "If we were to rerun human civilization again, would we come up with the same physics in the same order?"

And I don't think that's a very practical sol- uh, practical experiment to run.

Fei-Fei Li52:17

You know, one experiment I wonder if people could run is that we have plenty of astrophysical data now on the planet, uh, or, or celestial body, uh, movements. Just feed the data into a model and see if Newtonian law, uh, emerges.

Justin Johnson52:34

My, my guess is it probably-

Fei-Fei Li52:36

It's not

Justin Johnson52:36

... probably won't.

Fei-Fei Li52:37

It, it, that's my guess. It's not. The abstraction level of Newtonian law is at a different level from what these language, uh, LLMs represents.

Justin Johnson52:47

Yeah.

Fei-Fei Li52:48

So I wouldn't be surprised that giving enough celestial movement data, an LLM would actually predict pretty accurate movement trajectories. Let's say I invent a planet, uh, uh, surrounding a, a, a star, and giving enough data, my, my model would tell you, you know, on day one where it is, day two where it is.

I wouldn't be surprised. But F equals ma, or, or, you know, action equals reaction, that's just a whole different abstraction level. That's beyond just the today's LLM.

Guest53:23

Okay, what world model would you need to not have it be a geocentric model? Because if I'm training just on visual data, it makes sense that you think the sun rotates around the Earth, right? But obviously that's not the case.

So how would it learn that? Like, I'm curious about all these, like, you know, forces that we talk about. It's like sometimes maybe you don't need them because as long as it looks right, it's right. But, like, as you make the jump to, like, trying to use these models to do more high-level tasks-

Alessio53:53

How much can we rely on them?

Justin Johnson53:55

I think you can need kind of a different learning paradigm, right? So like, you know, there's a bit of conflation here happening where saying is it LLMs and language and symbols versus, you know, human theory building and human, human physics, and they're very different because an LL- like, the, the human objective function is to understand the world and thrive in your life, and the way that you do that is by, you know, sometimes you observe data, and then you think about it, and then you try to do something in the world, and it doesn't match your expectations, and then you want to go and update sort of your, your, your, your bu- your understanding of the world online.

And people do this all the time constantly, like whether it's, you know, I think my keys are downstairs, so I go downstairs and I look for them and I don't see them, and, oh, no, they're actually up in my bedroom.

So we're, we're... like, because we're constantly interacting with the world, we're constantly having to build theories about what's happening in the world around us and then falsify or add evidence to those theories, and I think that that kind of process writ large and scaled up is what gives us F equals MA in Newtonian physics, and I think that's a little orthogonal to, you know, the modality of model that we're training, whether it's language or, or, or spatially.

Swyx54:58

The way I put it is almost like this is almost more efficient learning because you have a hypothesis of here are the different possible worlds that are granted by my available data, and y- then you do experiments to eliminate the worlds that are not possible, and you resolve to the one that's right.

To me, that's also how I also have theory of mind, which is like I'm-- I have a few th- theses of what you're thinking, what you're thinking, um, and I try to create actions to resolve that or, or check my intuition as to what you're thinking, you know.

And, and obviously our LLMs don't do any of these.

Fei-Fei Li55:31

Uh, theory of mind possibly also will break into even emotional intelligence, which today's AI is really not touching at all, right?

Swyx55:41

And when we really, really need it. Uh, you know, people are starting to depend on these things, uh, probably too much and, and that's a, that's a whole topic of-

Fei-Fei Li55:49

Yeah

Swyx55:49

... of a other debate. I do have to ask because a lot of people have, like, said this to us. How much do we have to get rid of? Uh, you know, is, is, uh, is sequence to sequence modeling out the window?

Is attention out the window? Like, how much are we re-questioning everything?

Justin Johnson56:02

I think, uh, I think you stick with stuff that works, right? I, I, I think like-

Swyx56:06

So attention is still there.

Justin Johnson56:07

I, I think attention is still there. I think there's a lot-- like, you don't need to fix, fix things that aren't broken. Um, and like it's, it's-- there's a lot of hard problems in the world to solve, but let's focus on one at a time.

Um, I, I think it is pretty interesting to think about new architectures or new paradigms or, or drastically different learning ways to learn. Um, but you don't need to throw away everything just because you're working on new modalities.

Fei-Fei Li56:27

I think sequence to sequence is actually, um, in world models, I think we are going to see algorithm or architecture beyond sequence to sequence.

Justin Johnson56:38

Oh. Oh, but, but here actually I think there's, there's a little bit of, um, you know, technological confusion and, and transformers already solved that for us, right? Like transformers are actually not a model of sequences. A transformer is natively a model of sets, um, and that's very powerful.

But because, um, a lot of the transformers grew out of earlier architectures based around recurrent neural networks and RNNs definitely do have like a, a, a built-in architectural-- like they do model one-dimensional sequences.

Swyx57:02

Okay.

Justin Johnson57:02

But transformers are just object, uh, models of sets, and they can model a lot of-- th- those sets could be, you know, 1D sequences. They could be other things as well.

Swyx57:09

Do you literally mean set theory? Like-

Justin Johnson57:11

Yeah, yeah. So, so, um, um, like-

Swyx57:12

Galois, whatever.

Justin Johnson57:13

Yeah, yeah. Yeah, yeah. So a, a transformer is actually not a model of a sequence of tokens. A transformer is actually a model of a set of tokens.

Alessio57:19

Yeah.

Justin Johnson57:19

Right? The only thing that gives the, the injects the order into it in, in, in the trans- in the standard transformer architecture, the only thing that differentiates the order of the things is the positional embedding that you give the tokens, right?

So if you, if you choose to give a, a sort of 1D positional embedding, that's the only like mechanism that the model has to know that it's a 1D sequence. But all the, all the, like, operators that happen inside a transformer block are either token wise, right, so they either you have an FFN, you have QKV projections, like you have norm- per-token normalization.

All of those happen independently per token. Um, then you have interactions between tokens through the attention mechanism, but that's also sort of, um, it's, it's permutation equivariant. So if I permute my tokens, then the attention operator gets a permuted output in exactly the same way.

So it's actually a, a natively an architecture of sets of tokens.

Swyx58:02

Literally a transform-

Justin Johnson58:03

Yeah

Swyx58:04

... in, in a math term.

Alessio58:06

I know we're out of time, but, uh, we just wanna give you the floor for some call to action either on people that would enjoy working at World Labs, what kind of people should apply, what research people should be doing outside of World Labs that would be helpful to you or anything else on your mind.

Outro58:06

Fei-Fei Li58:21

I do think it's very exciting time to, um, to be looking beyond just language models and think about the, the boundless possibilities of, uh, spatial intelligence. So we are actually hungry for talent ranging from very deep researchers, right, thinking about the problems like Justin just described, you know, training large models of, uh, world models.

Uh, we are hungry for engineers, good engineers building systems, you know, from training optimization to inference to, uh, product. And we're also, uh, hungry for good business, uh, you know, um, product, uh, thinkers and go-to-market and, you know, business talents.

So, so we are hungry for, for talent. We-- especially now that we have, uh, uh, exposed the model, uh, to the world through Marble, I think we have a great opportunity to work with even a bigger pool of talent to solve both the, the model problem as well as, uh, deliver the best product to, uh, to the world.

Justin Johnson59:25

Yeah. I think I'm also excited for people to try Marble and do a lot of cool stuff with it. I think it has a lot of really cool capabilities, a lot of really cool features that fit together-

Fei-Fei Li59:32

That's true

Justin Johnson59:32

... really nicely. Um-

Fei-Fei Li59:33

In the car coming here, Justin and I were saying people have not totally discovered the-- okay, it's only twenty-four hours. Have not totally discovered some of the advanced mode of editing, right? Like turn on the advanced mode, you can, like Justin said, change this color of the bottle, you know, change your floor and change the trees and-

Swyx59:53

Well, I, I actually tried to get there, but when it says create, it just makes me create a completely different world instead of-

Fei-Fei Li59:58

You need to click on the-

Swyx1:00:00

Yeah, yeah

Fei-Fei Li1:00:00

... advanced mode.

Swyx1:00:01

It's like a UI issue.

Fei-Fei Li1:00:02

We can, we can improve on our UI, UX. So remember to click that.

Justin Johnson1:00:05

Yeah. We need to, we need to hire people who work on the product.

Swyx1:00:09

Uh, but one thing we got that was clear from what you guys are looking for is also intellectual fearlessness-

Fei-Fei Li1:00:13

Yes

Swyx1:00:13

... which is something that I think you guys hold as principle.

Fei-Fei Li1:00:16

Yeah. I mean, we are literally the first people who are trying this, both on the model side as well as, as on the product side.

Alessio1:00:25

Thank you so much for joining us. This was fun.

Fei-Fei Li1:00:26

Thank you, guys.

Justin Johnson1:00:27

Yeah, thanks for having us.

Fei-Fei Li1:00:28

Yeah.