LALatent SpaceApr 2, 2026· 1:06:48

Moonlake: Interactive, Multimodal World Models — with Chris Manning and Fan-yun Sun

Moonlake AI founders Chris Manning and Fan-yun Sun argue that interactive, multimodal world models require structured symbolic reasoning over pure scale, enabling indefinite multiplayer gameplay and causal consistency that video generation models like Genie and Sora cannot achieve. Their approach uses code engines and physics simulators as cognitive tools, producing reasoning traces that handle geometry, physics, and logic, while a separate diffusion model (Reverie) handles pixel fidelity. They aim to replace traditional rendering and empower creators by allowing human intent to be injected at a symbolic layer. Manning contrasts this with Yann LeCun's JEPA, emphasizing language and abstraction over pixel-level prediction. Moonlake is hiring engineers at the intersection of code generation, computer vision, and graphics.

  1. 0:00Intro
  2. 2:26Why World Models?
  3. 5:37Scale vs Structure
  4. 16:09JEPA Debate
  5. 22:10Reasoning Traces
  6. 34:50Evaluating Worlds
  7. 43:40Diffusion Limits
  8. 48:14Product Vision
  9. 53:30Audio & Roots
  10. 1:01:00Hiring & Name

Powered by PodHood

Transcript

Intro0:00

Chris Manning0:00

I think this whole space is extremely difficult as things are emerging now. And I mean, it's not only for world models, I think it's for everything, including text-based models, right? 'Cause, you know, in the early days, it seemed very easy to have good benchmarks 'cause we could do things like question answering benchmarks.

But, you know, these days so much of what people are wanting to do is nothing like that, right? You're wanting to get some recommendations about which backpack would be best for you for your trip in Europe next month.

It's not so easy to come up with a benchmark, and it's the same problem with these world models.

Host0:44

Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content.

We've been approached by sponsors on an almost daily basis, but fortunately, enough of you actually subscribe to us to keep all this sustainable without ads, and we wanna keep it that way. But I just have one favor to ask all of you.

The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you, and it means absolutely everything to me and my team that works so hard to bring The In-Space to you each and every week.

If you do it, I promise you, we'll never stop working to make this show even better. Now let's get into it.

Okay, we're back in the studio with Moonlake's, uh, two leads. I, I guess there's, there's other founders as well, but, uh, Sun and Chris Manning, welcome to the studio.

Chris Manning1:43

Thank you.

Fan-yun Sun1:43

Thanks a lot. Thanks for having us.

Host1:45

You've guy-- you guys have, uh, you know, come-- burst onto the scene with a really refreshing new take on world models. Um, I would just want to, uh, sort of, I guess, ask how you-- the two of you came together.

Chris, you're a legend in NLP and just AI in, in general. Uh, you're, you're his grad student, I guess.

Fan-yun Sun2:01

Actually, my co-founder.

Host2:02

Oh, yeah.

Fan-yun Sun2:03

I should give a lot of credit to my co-founder, Sharon.

Host2:06

Yeah.

Fan-yun Sun2:06

Um, she was, she was actually working with Professor Fei-Fei Li and Johnson, and then she ended up working with, um, Ron and Chris Manning here. And then so I got connected through to Chris initially, actually, through my co-founder.

Host2:18

What is Moonlake? What, what is, uh... A-actually, I'm also very curious about the name, but, like, why going into world models?

Why World Models?2:26

Fan-yun Sun2:26

So I was working a lot with actually Nvidia Research during my PhD years on essentially generating interactive worlds to train reinforcement learning agents or embodied AI agents. And then there's two observations, one in academia and one in industry.

In industry, like folks at Nvidia are actually paying a lot of dollars to purchase these types of interactive worlds, whether it's for the sake of evaluation or training their robots, um, or policies or models. And then, um, in academia, the same thing is happening.

And more specifically, when I was actually working with Nvidia on the synthetic data foundation model training project, we were actually generating a lot of the synthetic data and showing that, hey, you can actually-- these synthetic data are actually as useful as real world data when it comes to multimodal pre-training.

But then, uh, like I said, there's a lot of dollars being paid out to, like, external vendors or, or, like, other folks to manually curate these types of data. It was very clear to us that, okay, on our way to, let's call it embodied general intelligence, models need to learn the consequences behind their actions, which means that they need interactive data.

And the demand for those types of data are growing exponentially, but everybody's sort of thinking about it from a pure, say, video generation perspective or something else. But we feel like the, the true actually opportunity is actually building reasoning models that can do these things like how humans do these things today.

So that's a little bit on the genesis of Moonlake, and I think the reason I got into world models was partly a philosophical take of the-- on the world where I like, you know, believe in the simulation theory and stuff like that.

But on the other, on the other hand, it's really just like, oh, like, there's an opportunity there that I feel like nobody's doing it the way I think should be done.

Chris Manning4:05

I can say a little bit about that. Yeah, so of-- the o-overall goal is the pursuit of artificial intelligence and, you know, most of my career's been doing that in the language space, and that's been just extremely productive, as we all know the story of the last few years.

I don't have to tell about how much we've achieved with large language models. But although they're being extremely effective for ramping language and general intelligence, it's clearly not the whole world. There's this multimodal world of vision, sound, taste that you'd like to be dealing m-with more than just, um, language.

And then the question is how to do it. Um, and despite, you know, a huge investment in the computer vision space, right? As a research field, computer vision has been for decades far, far larger than the language space actually.

I mean, I think it's fair to say that, you know, vision understanding sort of stalled out, right? You got to object recognition, and then progress just wasn't being made, right? If you look at any of these, um, vision language models, it's the language that's doing ninety percent of the work, and the vision barely works.

And so there's really an interesting research question as to why that is. And at heart, um, the ideas behind Moonlake are an attempt to answer that, believing that there can be a really rich connection between a more symbolic layer of abstracted understanding of visual domains which aren't in the mainstream vision models, which are still trying to operate on the surface level of pixels.

Scale vs Structure5:37

Host5:49

Mm-hmm. I think one of your blog posts you put it as structure, not scale. Is that, uh, a general thesis?

Chris Manning5:57

Yeah. Well, scale is good too.

Host5:59

Yeah, scale is good too.

Chris Manning6:00

Lots of data is good as well.

Host6:01

Structure and scale.

Chris Manning6:01

But nevertheless, you want the structure, yeah, to be able to much more efficiently learn.

Host6:08

Yeah. The other thing I really liked also is you put out an example of what your kind of reasoning traces look like, right, which you wouldn't- Distill is, is the word that comes to mind. I don't even think that's a good, good description, but it would involve, for example, geometry, physics, affordances, symbolic logic, perceptual mappings, um, and what, what have you.

But like that, that is the kind of example that involves, let's call it spatial reasoning, world model reasoning as c- as compared to normal LLM reasoning. Yeah.

Host 26:37

But also like taking it a step back, so how do you guys define world models? You know, a lot of people see like, okay, you can do diffusion, you can do video generation, but, uh, you guys put out quite a few blog posts.

You put out a essay recently, we can even pull it up, about efficient world models. Um, you have a pretty like structural definition here, but for the general audience that don't super follow the space, right? What's, what's the difference in what we see from like a video generation model to a world gen, a simulator?

How do you kind of paint that landscape?

Chris Manning7:04

Yeah. So I think this is actually a little bit subtle because, you know, people look at these amazing generative AI video models, Sora, Veo 3, one of these things, and they think genies. They think, "Oh, this is amazing.

This is sort of, you know, we've solved understanding the world because you can produce these generative AI videos." But the reality is that although the visuals do look fantastic, those visuals actually aren't accompanied by an understanding of the 3D world, understanding how objects can move, what the consequences of different actions are, and that's what's really needed for spatial intelligence.

So I mean, a term we sometimes use is that you need action condition world models, that you only actually have a world model if you can predict, given some action is taken, what is going to change in the world because of it.

And in particular, that becomes hard over longer timescales. So if you're simply, you know, trying to predict the next video frame, that's not so difficult. But what you actually want to do is understand the consequences, likely consequences of actions minutes into the future.

And to do that, you actually need much more of an abstracted semantic model of the world.

Host8:36

Yeah. The question comes where you want to have more structure than is available in just predicting the next token. Um, and typically-- Well, let's, let's call it the experience of the last five years has been that, that is just washed away by scale, right?

Um, so what is the right middle ground here that, uh, you don't ignore the Bitter Lesson, but also you can be more efficient than what we're doing today?

Chris Manning9:02

You know, one possibility is, look, if we just collect masses and masses and masses and masses of video data, this problem will be solved. Um, under certain assumptions, that could be true, but there are sort of multiple avenues in which it could not be true.

The first is what's really essential is understanding the, the consequences of actions, producing an action-conditioned world model. And if you're simply, um, collecting observational video data, which is the easy stuff to collect when you're sort of mining online videos, you don't actually know the actions that are being taken to see how the video is changing.

And so if you're never collecting directly actions and you're having to try and infer them from what happened in the observed video, that's not impossible, but it's very hard, and it's not really established that you can get that to work at any scale yet.

And so there's a lot of premium on collecting action condition video data, which is part of why there's been a lot of interest in using simulation so that you can be collecting data where you do know the actions, which is in quite limited supply.

But there's also in the limit of as much data as you could possibly have, you know, maybe the problem is eventually solvable. But even though we collect huge amounts of text data, text data is always at a great level of abstraction, right?

Language is a human-designed abstracted representation where there's meaning in each token, and it's representing an abstraction of the world, right? As soon as you're describing someone as a professor, and as soon as you're saying that they're condescending, right, you know, these are very abstracted descriptions of the world.

It's not at sort of what you're observing as pixel level. And so to get to that kind of degree of abstraction starting from pixels is orders and magnitude of extra data and processing. And so although, you know, we absolutely want to exploit, get as much data as possible, use the Bitter Lesson, nevertheless, if there are ways in which you can work with five orders of magnitude less data than people working purely from pixels, you're gonna be able to make a lot more progress a lot more quickly, and that's the bet here.

And so you could just say that's only wanting to be able to, you know, do it more efficiently, do it more quickly, do it more cheaply. But I think it's actually more than that. I think one should be making the analogy to how human beings work.

At one level, you know, yes, we have these high-resolution eyes, and we can look and see a scene like a video. But all of the evidence from neuroscience and psychology is that most of what comes into people's eyes is never processed, right?

That you're doing fairly fine-grained-

Host12:27

Foveated

Chris Manning12:28

... processing of e-exactly what you're focusing on. But, you know, as soon as it's away from that of, yeah, there's another guy over there, that you've sort of only processing top-down this very abstracted semantic description of the world around you.

And so, you know, that's what human beings are doing. They're working with semantic abstractions. And so I think it is just the right representation, 'cause we also have other goals. We want to be able to do, you know, real time worlds.

That means there's a limit to how much processing you can do, and we want to do long-term planning and consistency, and again, that favors abstraction. I mean, I guess there was actually a recent blog post that came out from our friends at Physical Intelligence and, you know, they were sort of heading in the same direction.

They were saying, "Oh-"

Host13:20

The Pi model?

Chris Manning13:21

Yeah.

Host13:21

Yeah.

Chris Manning13:21

To maintain a long-term memory of what's happening in the world so we can ha- do longer term, we're actually storing text of what is, um, you know, been happening in the world, right? It's not such a successful strategy of trying to keep it all at a pixel level.

Host 213:40

And yeah, I mean, you can see it in video models like that. Temporal consistency, we're at a scale of train on, you know, all the video data we have. We have it for maybe 30 seconds, a few minutes.

That's not the same as a game state played for half an hour, right? Um, I thought you guys break it down pretty well. You have a, you have a blog post about, uh, building multimodal worlds with an agent.

I don't know if you guys wanna talk about this. This is one of the things I read. I thought-

Host14:04

Yeah, it's, so the thing I talked about with the reasoning chain, yeah.

Host 214:07

So there's, like, different phases to this. It seems like it's more of an agent, a scaffold. Uh, very different approach than just, you know, type in a prompt and you, you don't have the same consistency. It also, like, for people that are listening, you know, I, I would highly recommend reading it.

It breaks down the problem in a different light, right? So, like, what do you need to consider when you're talking about video-- like, world game models, right? How would-- what do you need to consider? What are the factors?

What are the elements? What's the state? So I don't know if you guys have stuff to talk about for this one.

Fan-yun Sun14:37

Yeah. Um, actually, I wanted to add on a little bit-

Host 214:39

Yeah

Fan-yun Sun14:39

... on our previous point. Which is just like-

Host14:41

Change topics quickly.

Fan-yun Sun14:43

I, I do feel like sometimes people confuse, like, oh, like, we're taking an, an ab- a, an method with, with abstraction. That means they don't believe in Bitter Lesson. Like, like, that's just false, right? Like, we are believers of Bitter Lesson, but then I feel like the question that we always discuss is, like, what is the right abstraction level today?

The analogy I like to make is, like, let's just say we can encode and decode, represent all of images, videos, audio in bytes. Then the most Bitter Lesson approached is to train a next byte prediction model as opposed to a next token prediction model, where it's just like, okay, it's natively-

Host 215:15

Yeah

Fan-yun Sun15:15

... multimodal and can just... Um, but it's like, well, yeah. Like, to, to Chris' point, it's like the scale and compute you need to achieve that, um, um... So that's why we always come back to, like, okay, what is the most efficient way to do it?

And, and reasoning models, to, to the point of this blog post, is a showcase of, like, hey, we're actually just, like, reasoning about the world and reasoning about the aspects of the world that c- that matter for me to learn what I want to learn from this world model.

Um-

Host15:43

Yeah, it's like y- you're improving the en- encoder of whatever you're, uh, trying to model, and, like, a better representation would just represent the important things in less space.

Fan-yun Sun15:54

Yeah.

Host15:54

Which would just be more efficient.

Fan-yun Sun15:56

Yeah.

Host15:56

Um, so yeah, I, I fully agree that it is not, um, antagonistic to, uh, Bitter Lesson. I do wanna, wanna mention one more thing. Um, is there any philosophical differences with the JEPA stuff that, uh, Yann LeCun is working on?

I gotta go there. You, you, you're- You're, you're mentioning, like, some latent abstraction. I'm like, "Okay, fine. Let's, let's talk about it," right? Like, it's a elephant in the room.

JEPA Debate16:09

Chris Manning16:17

Yeah, there are philosophical differences. Um, Yann LeCun is a dear friend of mine. Um, but he has never appreciated the power of language in particular or symbolic representations in general. Yann is a very visual thinker. He always wants to claim that he thinks visually, and there are no word symbols or math in his head.

Um, maybe that's true of Yann. It's certainly not the way I think. Um, but at any rate, you know, um, the world according to Yann is the basic stuff of the, the world and of intelligence is visual, and language is just this low bit rate communication mechanism between humans, and it doesn't have much other utility, and it's far inferior to the high bit rate video, um, that comes into your eyes.

And I think he's fundamentally missing a number of important things there, right? Think of this evolutionary argument looking at animals, right? That the closest analogy is the things with chimps, right? So chimpanzees, you know, have fairly similar brains to human beings.

They have great vision systems. They have great memory systems. They've got, you know, better memory than we do of short-term memories. They can plan. They can build primitive tools. But, you know, humans massively ahead in what we understand about the world, what we can plan, what we can build.

And essentially what took off for us was that humans managed to develop language, and that gave a symbolic knowledge representation and reasoning level, which just gave this sort of vaulting of what could be done with the intelligence in brains.

So the philosopher Dan Dennett refers to language as a cognitive tool and argues that, you know, humans, unique among the creatures in the world, have managed to build their own cognitive tools, and language is the famous first example.

But other things like, um, mathematics and programming languages are also cognitive tools. They give you an ability to think in abstractions, in extended causal reasoning chains, and that allows you to do much more. And we use that for spatial representation and intelligence and planning and gameplay as well.

So we believe, and this is, you know, underlying the specific technologies that Moonlake is making, that symbolic representations are powerful, and you want to use it in your understanding of the visual world when you want a causal understanding, when you want to maintain long-term consistency and prediction.

And, you know, as I understand it, that's just not in Yann LeCun's worldview. So I think that's the fundamental philosophical difference. Um, then there's the specific model he's been advancing, JEPA. I mean, that's a reasonable research bet as a direction as to, to head for building out a model of the visual world.

To my mind, it's sort of one reasonable research bet. It's not really established it's the best one that everyone should be following.

Host20:06

At least developed at scale with, at Meta. But it's not just vision, right? Like, I mean, JEPA is a-- you know, just joint embedding prediction can be applied to anything really, and, and people have done it. If the argument is that there is a latent representation or that is, that is probably more, uh, suited to the task, then why not let machines do it for us instead of predefining it at, at all?

And isn't something like a JEPA-shaped thing the right answer? And if not, why not?

Chris Manning20:31

So I think there's a part of JEPA that's right, which is you do want to have a joint embedding that gives you a consistent model of the world. And Yann's argument is you can never get that from auto-regressive language models 'cause they're sort of left to right churning out one token at a time.

I guess this is where we're, um, you know, the research arguments of the field. You know, I'm not actually convinced that's right 'cause although the token production is this auto-regressive, um, process that's heading, you know, left to right...

I guess it doesn't have to be left to right. But anyway, in sequence of tokens. We could have right to left for Arabic. Um, but, um, you know, although that's true, all of the weights of the model that are internal to the transformer, they are a joint model of the model's understanding of the world.

And so I think you can think of the weights of the model as a form of joint representation, and therefore it is plausible to think that that could be the basis of a world model which avoids, um, Yann's objections.

Host21:52

I think I follow, and obviously that will touch on what Moonlake eventually ends up doing as well, right? Like, which it's hard to tell because you put out the end results, but we don't know the inputs that go into it.

So it's, it's, like, you know, that's, that's something that we have to figure out over time.

Chris Manning22:09

Yeah.

Reasoning Traces22:10

Host 222:10

I mean, I guess this kinda breaks down some of the outputs. Do you wanna walk us through it?

Fan-yun Sun22:15

Yeah. So this, this really just walks us through the reasoning traces of like, okay, so let's just say if we wanna build a world. In this context, it's really just a game demo that, that shows the, uh, the variety of interactions that this world model can build.

And yeah, it's really just the reasoning traces of like, okay, if you're prompted to create a bowling game, like, how did it achieve what you saw, that level of causality, interaction, and consistency, right? Um, so yeah, this is almost just like a, an example of like a reasoning trace.

Host 222:45

Very detailed.

Fan-yun Sun22:46

Yeah.

Host 222:46

Very, very detailed. I mean, you gotta... Like, you don't even realize it, right? Like, when a video is generated, what happens when a ball strikes a pin, right? So-

Fan-yun Sun22:53

Yeah

Host 222:53

... first, like, you-- there's audio in that. Like, audio triggers happen, score increments, uh, the world changes. Like, pins have to start dropping. There's a timer that goes on. Um, you know, it's just, like, very similar to how now we're used to reasoning for language models.

There's a whole state of what happened, so geometry, physics, uh, all this stuff. And then-

Fan-yun Sun23:11

Yeah

Host 223:12

... there's kinda that single prompt, so asset, um, physication, all this stuff. It's, it's like a-

Fan-yun Sun23:18

Yeah

Host 223:18

... it's a nice view to see what's going on.

Host23:20

I think Sun is also too polite to point out that, uh, both, like, Google's Genie, uh, demos as well as, uh, World Labs' Marble do not have interactive worlds. Uh,

Fan-yun Sun23:32

That's the benefit of having a reasoning model, right? Like, 'cause you can, you can say, "Oh, like, maybe in this particular context, I want to learn how to bowl." And then you can say, "Okay, then what is it important when it comes to learning how to bowl?

Okay, maybe it's like I need to understand the, the basic of like physics, and I wanna throw it over them. I wanna know that when I-- when it resets, it's, it's a new game, so I know that..." Yeah, basically, you know, you know, you know to pick up the ball.

You know the ball's gonna cause the pins to fall down. You know that what's important to this particular bowling game is to score, and you know that the score corresponds to the number of pins that fell down. Um, so it's just like if it's a model that sort of knows what it looks like, knows what a bowling game looks like, but doesn't actually allows you to practice over and over again and to understand that, oh, like, what it takes to actually get a high score, then it sort of doesn't actually allow you to learn what you set out to learn within the world model, right?

And, and I think this is really just one example of showing, like, the advantages of the approach that we're taking over most, uh, let's call it the zeitgeist is today, uh, when people talk about clinical world models.

Chris Manning24:44

Right. So it sort of seems like the question to ask when there's a world model is Can I not only just wander around the world and look at the beautiful graphics? Can I interact with the objects in the world and see the right consequences of actions?

Host 225:04

And you also understand what the consequences would be if you do something, right? So it's not just like, okay, there's one thing, if I pick it up, something will happen. But, you know, there's, there's 50 options, and I know I can expect, I can infer what would happen if I do any of them, right?

So very different when you can actually see it, play around with it. Um-

Host25:21

There, there's two cheeky elements of that. I mean, the, the, the sort of, I guess, less ambitious one is, um, let- let's really establish for it for listeners, uh, why is this fundamentally different than writing Unity code, right?

Like just creating a model to translate a prompt into Unity code.

Fan-yun Sun25:39

So there is an underlying physics engine.

Host25:41

Yeah.

Fan-yun Sun25:41

Um, in that sense, there's some overlapping things to Unity. But the way we think about it is like physics engine or tools or code are cognitive tools, like borrowing Chris's term, right? Like tools that the model can employ as means to an end.

So today, maybe you say, okay, in this particular context, we care about physics, we care about the long-term causality consequences, then yes, we deploy a-- employ physics engine. And then maybe tomorrow we say, okay, we're, we're training, let's just say drones, where we only care about really fluid dynamics and the visual aspect of the world.

Then, then yeah, maybe we don't actually, the model actually doesn't have to use a physics engine, or maybe it employs other types of representation or physics engine to achieve the task. So yes, writing code for Unity is sort of similar to a tool that our-- a model can employ, but our goal is for model to take a representation conditioned reasoning approach or process-

Host26:42

Yeah

Fan-yun Sun26:42

... internally. Yeah.

Host26:44

Using these things as, uh, just like general tool calls, right? Which I think is very interesting. The other more ambitious one is, uh, some kind of recursive element where it becomes multiplayer, right? Like here, there's single player elements.

You're not modeling any other people involved, and that is a whole other thing.

Fan-yun Sun26:59

But in fact, we can already do multiplayers.

Host27:01

Oh, yeah? Okay.

Fan-yun Sun27:02

Yeah.

Host27:02

I haven't seen any demos.

Fan-yun Sun27:03

So if you, if you just actually just like prompt our, our model to say, "Hey, like configure to multi- multiplayer," then it'll do like this... You'll be able to configure multiplayer-

Host27:11

Great

Fan-yun Sun27:12

... persistency database for you.

Host27:14

Easy.

Fan-yun Sun27:14

Yeah.

Host 227:14

So what, what are like some of the current limitations in where we're at? So there's one approach of like, okay, scale up video predictors. Obviously, there's data issues. Uh, you know, with approaches like this, uh, is it data constraints?

What are like the next steps? Is it real time? Like, so there's one side of, you know, write an agent to write Unity code, but okay, I wanna be streaming a game real time. I wanna have characters being also like agentic.

But where, where do we kinda see this scaling up, right?

Fan-yun Sun27:40

Yeah. There's definitely a data constraint. Like the more data, the, the better this reasoning model can almost basically act as humans to like operate a variety of tools and softwares to build whatever is necessary. And then there's a sort of fidelity constraint, which we're actually solving with a-another model, Reverie, which we can talk about later.

Um, but it's like, well, it's not as easy to get to photorealism with the approach that we're taking. Um, but we think there are better solutions to that, which is we can dive into later, later.

Host 228:13

The one, one thing you note here is it's a diffusion model, right? So there's, there's a few approaches, uh, diffusion, Gaussian splatting. Um, yeah. So Reverie diffusion model, you guys wanna-

Fan-yun Sun28:23

Yeah

Host 228:24

... introduce?

Fan-yun Sun28:24

Yeah, totally. So within our world modeling framework, we think there are two models that we train, right? Like there's the multimodal reasoning model that we just talked about that essentially handles mainly the, the causality, the persistency, and logic determinism, determinism of the world.

And then Reverie is our bet on saying, okay, like while all those model, um, can take care of all these things that we just talked about, its limitations compared to existing, say, video models, is that it doesn't have as high of a pixel fidelity right out the gate, right?

And Reverie is to say, hey, we can actually take whatever persistent representation that we generate with our multimodal reasoning model and learn to restyle it into photorealistic styles or arbitrary styles you want. So this model is almost to say, hey, I'm going to respect the persistency and re- interactivity of the world that you created, but my only job is to make sure that its pixel distribution is close to what we want.

Yeah.

Host 229:29

Yeah. Good example right there.

Host29:30

You kept the KL divergence.

Fan-yun Sun29:33

Oh, where?

Host29:34

No, no. I mean, no, this, this is a, a classic like, um, how you don't stray too far from the source material as you, you kept the KL, which is-

Fan-yun Sun29:41

Oh, yeah

Host29:42

... kinda cool.

Fan-yun Sun29:43

Yeah, yeah.

Chris Manning29:44

I mean, and the difference is, and I mean, Sun was pointing at this, where sort of saying it's in one way a more difficult path, but a better path. That, you know, typically the diffusion models are producing the whole scene, and it looks lovely, but there isn't spatial understanding behind it, which is allowing for the real-time graphics gameplay, the spatial intelligence, understanding the consequences of worlds.

Where this is, um, taking a path where it is assuming an abstracted semantic model of the world, the world state, and then the diffusion model is then being used on top of that to produce the high quality graphics.

Host30:29

Is there an intended practical, uh, or business use for this, or is it like a, like a demonstration of capabilities?

Fan-yun Sun30:36

We actually believe that this is gonna be the next paradigm of rendering. So it's gonna replace how-

Host30:41

Ah

Fan-yun Sun30:42

... rasteri- rasters. It's gonna replace DLSS today because it not only has these pixel prior that's learned from the world such that you can literally play any game in photorealistic styles, which is a lot of people's desire when they do GTA, right?

Like, um-

Host30:54

All the mods, all the people adding perfect lighting and all this.

Fan-yun Sun30:57

You know?

Host30:57

So skins for worlds, let's call it.

Fan-yun Sun30:59

Skins. Let's call it skins for worlds.

Host31:00

I mean, it's also like you can call it skin, you can call it customization. You can play it how you want, right?

Fan-yun Sun31:04

Yeah, exactly. And I think another thing that we really pointed out sp- specifically in this blog is the programmability of it, right? So what this means is that this renderer... Well, historically, renderer is always a derivative of the game state, right?

You're saying, "Okay, here's the game state. I'm rendering out a frame." But here I'm saying, actually, this renderer can be part of the gameplay loop. I can say something along the lines of, "If upon getting 10 apples, I'm gonna-" "...

my weapon of choice, my bullet's gonna turn into apples." And that's, that's possible because we can say we can basically dynamically have certain game state trigger the, the preconditions to the renderer-

Host31:43

Mm

Fan-yun Sun31:43

... such that the rendering is now part of the game loop, too. One thing is to just say, "Okay, it's, it's, it's appearance." But the second thing is also to say there's these novel interactions that are av- possible because this renderer now has actually priors of the world.

Host32:00

It is up to the artist to figure out what to do with it.

Fan-yun Sun32:02

It is up to the creators, yes.

Host32:04

Yeah.

Fan-yun Sun32:04

And I also think that's actually another big argument that we're making, and the reason that we're picking ba- taking the bet we're baking is that a lot of the times, whether it's for embodied AI or gaming, like you want a layer where human can inject their intentions, right?

So for example, let's just say in the context of gaming, it's obviously like my creative intent. But maybe in the context of embodied AI it's like, oh, like I take this foundational policy, and I want to actually fine-tune it to deploy in my house.

So you want to almost say inject-- have a layer where human can say, "Oh, here's the distribution of things I want to create to achieve my goal."

Host32:38

Mm.

Fan-yun Sun32:39

And I think 3D graphics as it, as it is today is basically the layer for people to say, "Hey, what do I care about in this world?" And it allows, um, basically human intent to be expressed in these worlds much more explicitly and distributionally as opposed to just saying, "Hey, I'm gonna generate like arbitrary," and it's like just prompts, you know?

Host33:00

It's one of those things where, like I, I think you, you're gonna build up a series of models, right? This is just one of-- This is probably like the highest utility or he- heaviest, uh, frequency one. I don't know what to call this.

Where like you, yeah, you can immediately drop this in on any game, and you don't need anything else that, that you guys do. But, um, I, I could see, I could see that. I think the, the human intent is something that people are not even used to because we're so used to static worlds or, um, you know, worlds that just don't react or...

I don't know. It's, it-- You're kind of blowing my mind right now with like... Well, I m- I wonder if you've talked to people at GDC-

Fan-yun Sun33:34

Hmm

Host33:34

... and what are, what are they gonna do with it?

Fan-yun Sun33:37

Hmm. Yeah. No, the stance that we take on this front is like we're not gonna be more creative than our users.

Host33:42

To ship it out, yeah.

Fan-yun Sun33:43

Um, but we wanna make sure that we're building things in a way that really allows them to express their intent.

Host33:49

The thing that you said about here's the distribution that I want, I think text may be the, too low of a bandwidth to, to really demonstrate because I, I, you know, the, uh, I'm, I'm probably just gonna want to drop in a bunch of, uh, reference assets, and then you can figure it out from there.

Fan-yun Sun34:06

You probably wanna do a b- a mixture of both, right? Like you throw in a few images. "I wanted this style."

Host34:11

Yeah.

Fan-yun Sun34:11

"I want it to look like this." So it's, it's a mixture, right?

Chris Manning34:14

I, I think it's a mixture. I mean, yeah. I mean, there's clearly a visual component of this, and it's not that, you know, everything can be text 'cause of course you want to give a visual look. But there's also a massive amount of giving the overall picture of the look of the world and the behavior of things that you can express in a few words of text and it be very time-consuming and difficult to do via visual means.

So I think, yeah, you want a combination of both.

Host34:50

So one question I kind of have is how do we go about evaluating world models? So like there's many axes, right? One is like, okay, I have preferences. How well do we adhere to prompts? One is the simulation.

Evaluating Worlds34:50

Host35:00

One is like do things-- is there core logic that's broken? So coming from we know how to evaluate diffusion, there's fidelity, there's stuff like that, but what are some of the challenges that most people probably aren't thinking about?

Fan-yun Sun35:13

Yeah, I think this is like a great question and probably one of the hardest questions in world models because like I think it always comes back to what are you building this world model for? And depending on your end goal and purpose, the evaluation should differ.

So in the context of games, then the most direct way of measuring is how much time are people actually spending in this world that you create? And if your goal is to say, for example, in the context that we just talked about, like, hey, deploying, deploying action in body agent, then your, your end metric is then, okay, after training in these worlds that you generate, how robust it is to a- when you actually deploy to the target environment.

But then, you know, it's, it's hard to measure these end metrics. So today people have like these proxy metrics that I call that basically try to measure what we really care about, which is the end metrics. But then frankly, it's different for every use case.

Um, yeah.

Host36:07

Which seems like quite a challenge, right? Like in, in language models or video models, image models, your benchmarks are proxies, right? People aren't actually asking instruction following tool use questions. They're proxies of how well it will do downstream.

But for this, so like, you know, should, should teams, should companies have their own individual benchmarks outside of games if you think of stuff like, okay, video production, movies, stuff like that, that also wanna use world models. Should, should they sort of internalize like-

Host 236:36

Their own proxy? Is this something you guys do? Where, where does that kind of sit?

Chris Manning36:39

Yeah. I think this whole space is extremely difficult as things are emerging now. And I mean, it's not only for world models. I think it's for everything, including text-based models, right? 'Cause, you know, in the early days it seemed very easy to have good benchmarks 'cause we could do things like question answering benchmarks and could you answer the question based on these documents and the various other kinds of, you know, do pieces of logical reasoning or math.

But again, these are sort of-- and there are sort of visual equivalents of things like object recognition, right? They're these small component tasks. But, you know, these days so much of what people are wanting to do also with language models is nothing like that, right?

You're wanting to, um, have an interaction with the language model and get some recommendations about which backpack would be best for you for your trip in Europe next month, and it's not the same kind of thing, right? Um, and it's not so easy to come up with a benchmark as to does this large language model give you an effective interaction for guiding you in a good way for shopping, right?

So and it's the same problem with these world models. So if we take the game design case, well, success is that a game designer can produce what they are imagining in a reasonable amount of time, and that's really the kind of macro task.

But it, you know, that's a very hard thing to turn into a benchmark, and I think a lot of this is actually going to turn into people working, walking with their feet, right? I mean, I guess that's what's happening, you know, at the large language model level, right?

When people are choosing to use, you know, GPT-5 or Gemini or Claude. You know, individuals are trying out these different models and deciding, "Oh, I like the kind of answers that GPT-5 gives me," or, "No, I feel like I get more accurate detail from Claude," right?

Host 239:00

It's a lot of vibe check.

Chris Manning39:00

It's, it's a lot of, like-

Host 239:01

A lot of people just using it

Chris Manning39:01

... vibe checking. I realize that. But it's actually whether people feel it's giving them utility in what they want, right?

Host 239:09

And the, the interesting thing there is, like, a lot of people prefer the visual, right? This looks pretty, which is not the objective of what this is for, right? It's every... If a game designer is working on something, they care about the game engine-

Chris Manning39:21

Right

Host 239:21

... state. It's-- it can look whatever. You can fix that up later. Or you can have a really good game state and you can quickly edit it to 20, 20 different versions that keep state.

Chris Manning39:31

Right.

Host 239:31

But-

Chris Manning39:31

So that's a really important-

Host 239:32

Yeah

Chris Manning39:32

... distinction, um, for and for speaking to Moonlake's strength, right? So yeah. I mean, you know, great visuals are lovely to look at for a few seconds, but games are really all about the concept, the gameplay, and, you know, a lot of the time that doesn't actually even require great visuals.

I mean, there are just lots of very successful games which have relatively primitive visuals, and there are other games where people have spent millions producing photorealistic, um, visuals and the game sucks, right? Um, so, um, keeping those two axes apart is really important in thinking about what's important in a world model for different uses.

Host40:23

This conversation is reminding me of some game review and fiction discussions I've, um, had in my sort of non-AI related life. Uh, some, uh, some people might know Brandon Sanderson, who's a very famous, uh, fiction author, uh, ha- is, is a big, big game reviewer, and he, he's a big fan of video games where you change one thing about a normal what, what you might assume about, about the world.

For example, Baba Is You. I don't know if you might have come across that, where, like, the rules change as you play the game. And also, like, where, you know, you can do things like reverse time selectively or, like, change gravity selectively.

And I, I think this is also remind, reminds me of other kinds of world models that are created by authors, where Ted Chiang is, is my typical example, where he will take the world that you know today but change one thing about it and but then create a consistent world based on that.

Uh, which is a long-winded answer of me to, of, for me to say is it's easy to create alternative worlds that don't exist, but you change one thing, and then let's r- let's run a whole bunch of people through it to see if it works.

Chris Manning41:23

My first answer will be that seems a lot easier and more conceivable to do using techn- technology like Moonlake's- ... than with some of the other world models out there, um, where the sun can actually make it happen.

I'll let him give the second answer.

Host41:41

If I guess for you, you're constrained by the game engine tool, right? Like, at the end of the day, that's the, that's the thought, um, partner that you have. If I ask for something where like it never is allowed to reverse time or if gravity only ever works one way, then well, that's it.

But sometimes gravity might change.

Fan-yun Sun42:00

But it's a lot easier to change with code as opposed to-

Host42:05

Yeah

Fan-yun Sun42:05

... a model that is learned primarily on data of real world and virtual worlds that are... I guess, like, for example, Genie, where like there's actually trained on a lot of real world data and a lot of virtual gaming data, and it's hard to say...

Well, maybe it's easier to say, "Okay, I wanna change the visuals in, like, the time period of, of the world," but you can't change gravity, for example.

Host 242:27

I feel like you can to light bounds, right? Everything comes down to, like, code is a better way to execute it, but the models aren't that diverse and creative, right? You can say, "Okay, make gravity slower." It can do that, but it's limited to your representation of how you text it out, right?

Like, they're, they're only gonna do a few iterations, whereas programmatically, you know, if there's a game engine under the hood, you can, you can kinda go wild, right? So one of the- I don't know, one of the limitations of most models is that they're very over-trained to one style, right?

And extracting diversity is pretty difficult, at least. That's something that we've seen.

Fan-yun Sun43:03

I mean, are there other examples you have in mind where existing models, it would like-- it would be easier to do that's not using code? Like, uh, certain types of creative intent or, like, transition-

Host43:14

Uh, you know-

Fan-yun Sun43:14

State transitions

Host43:14

... clipping, uh... Other models, other world models are very good at clipping through things.

Fan-yun Sun43:19

Clipping? Like-

Host43:20

My, my, my legs clipping through a, a rock. Because, because it's, you know, it's just, it's just bad.

Like, uh, you would have to struggle very hard with your, your stuff to actually make that happen. Uh, which I think is m-maybe a topic that you actually prepared on, uh, uh, Gaussian splatting versus, uh, the other stuff.

Host 243:40

Yeah, yeah. It's just for those not super familiar, right? There's a-- There's Gaussian splatting. There is diffusion. Like, what works, what scales up. I feel like in February when Sora 1 came out, the, the blog post was literally titled, like-

Diffusion Limits43:40

Host43:52

We bring it up every now and then.

Host 243:53

You know, world, world, uh, video generation models or world simulators. Uh, it's super Bitter Lesson pilled. Yeah. Emer-

Fan-yun Sun44:02

It is.

Host 244:02

A, a lot of it is emergence, right? So, uh, not to go through their blog post. Basically, their whole thing was, as you scale up all this consistency, all this stuff just kinda solves. It's a very simple premise, right?

They just scaled up diffusion, and from there, you know, this is, this is Feb 2024. How much can we... I- it's already been two years, which is basically five years, you know. How much more in AI time do we need to just scale up?

Or, or do we hit a data cap? But I think we already talked about this a lot, right? Like, this is back to the beginning discussion of what's appropriate for the time, and that seems like your approach, right?

Fan-yun Sun44:36

Yeah. The point I'm trying to make is that there are very-- many, many different types of world simulators. And, like, having a world simulator that can produce pixel coherency is very, very useful for games and, you know, marketing and all these things.

But it's not as useful as people think when it comes to causal reasoning, when it comes to embodied AI. And yeah, like, it-- This, this title is true. Like, we're not saying that it's, it's like, you know, a not a great world simulator.

But actually, in the blog that we, we, we, we wrote, the bet is more so that there are gonna be disproportionately large share of value of real world task or in virtual tasks where high resolution pixel fidelity is not needed.

And yes, video models have their values. Yeah.

Host45:25

This is at the ex-absolute limit of my physics understanding, but one example that comes to mind is basically having to solve, like, bas- the equivalent of a three-body problem in a deterministic world, whereas the video models would just approximate it good enough.

Fan-yun Sun45:40

Yeah.

Host45:42

Right? Like, there's, there's some point at which your approach kind of runs into, like, the, "Well, you now have to simulate the world, please. Thank you very much." And, like, you're, you're trying to do that, but only to the extent that the game engine lets you, and, like, game engines cannot do some things.

Fan-yun Sun45:58

Yeah. No, I mean, I-I think the, the interesting or more technical question here actually is where do you draw the boundary between what's handled with, let's say, diffusion prior and where, when what's handled with symbolic priors?

Host46:13

Yes. Okay. Okay.

Fan-yun Sun46:14

And... Right?

Host46:14

Let's go there. Yeah.

Fan-yun Sun46:14

'Cause like this, this boundary can actually be fluid. Like, I think, like, maybe what you're trying to get at is like, okay, people are saying pixel prior everything. But what we're saying is, okay, there's a boundary that we draw where this is where we think provides the most economical value for the domains and things that we care about today.

And I actually do think, and this is something that we do internally all the time, which is like, okay, given new equations that we learn or new elements of the world and that we, we learn, or maybe some other knowledge that we acquire in the process of developing the models, should we still be k- maintaining this line exactly as it is today or should we move it a little bit left or a little bit right?

Right? Like sometimes that we realize that, oh, like, maybe customers or, or folks, like, want certain things that are better handled with pixel prior as opposed to, um, symbolic prior. Then we-

Host47:10

Yeah, your, your skin thing is a, is a example of moving it right. Yeah.

Fan-yun Sun47:13

Um-

Host47:14

Or left.

Fan-yun Sun47:14

Yeah, exactly.

Host47:14

I, I don't know what the difference, the left-right is.

Fan-yun Sun47:16

Yeah, yeah, yeah. No, the, the, the, the, the Revry model-

Host47:19

Yes

Fan-yun Sun47:20

... actually, we have a few iterations of them. They're actually a slightly different-

Host47:23

I know

Fan-yun Sun47:24

... uh, boundaries.

Host47:24

Yeah. You should, you should do that. That's a cool dimension to show.

Fan-yun Sun47:27

Yeah.

Host47:28

Is quantum mechanics the diffusion prior of our world?

Fan-yun Sun47:33

Uh.

Host47:34

Right? It's like th-that's the boundary of classical mechanics versus quantum, right? Like that, that's it, right? At one point, God plays dice, and the other point doesn't.

Fan-yun Sun47:42

I don't know, I don't know if Chris, you wanna say, but I think, I think generally I feel like physics is better with symbolic priors. Um-

Host 247:49

Even quantum physics.

Fan-yun Sun47:50

E-even quantum physics. Yeah.

Host47:52

This is starting to get to, um, MLST territory is, is what I call it, where, uh, he, he likes to get philosophical. Uh, we're, we're quite friendly.

Host 247:59

I mean, we need to get, we need to get singularity. I heard some of that.

Host48:04

No, no. I think that is actually really helpful, and, uh, man, I just want you to productize this. Like, as a product guy, I'm just like, "Well, okay, like-"

Product Vision48:14

Host 248:14

Also as a gamer, you know, I wanna-

Host48:15

"... like to a researcher," You know, like, it's cool, like, this, this is a theor-theoretical-- Like, you have a very good, I don't know, like, the way of thinking about these things, but I just wanna see you, like, you know, express it.

I do think, like, you're fundamentally things-- Wh-when you leave open new tools like, okay, use, use human intent to incorporate it into how you render. Well, artists are gonna have to take, like, two to three years to figure out what to do with this, and you just don't know.

Like...

Host 248:41

But I think, you know, this is, um, gives a much more approachable and controllable world for a designer-

Host48:49

Which is the beau-beauty of, uh, NLP

Host 248:51

... and that, that will

Chris Manning48:52

Enable it to be adopted and used, and we're very hopeful about that.

Host48:58

Yeah, yeah.

Fan-yun Sun48:58

Yeah. I mean, we are, we are very focused actually on commercialization in the sense that like we do, we do really believe in the data flywheel app approach-

Host49:05

Yeah

Fan-yun Sun49:05

... where, um, we put this in the hands of the creators and the users, and then they will teach us how and what capability our model should improve, and that's why we are, we are actually, you know, like products in b-beta.

Host49:17

Yeah. Focusing on gaming, what, what's like the adjacent thing to gaming?

Fan-yun Sun49:20

Embody AI basically. So maybe-

Host49:23

Yeah

Fan-yun Sun49:23

... we can, we can-- I'll, I'll maybe start with where we see the platform in three years.

Host49:27

Yeah.

Fan-yun Sun49:27

Which is like, okay, the users would tell us what they want to achieve. The end goal could be, "Hey, I just w-- I wanna make something to teach my kids the value of humility." Um, or it could be, "Hey, I wanna fine-tune my, um, drones to be really good at rescue situations."

It could be vacuum robots, "I wanna like train my manipulation or like vacuum robot to be very robust to my office." Right? But it's like whatever it is, the scenario-

Host49:53

Robust to my office.

Fan-yun Sun49:54

Or like navigate very-

Host49:55

Yeah, yeah

Fan-yun Sun49:56

... very robustly with- within my office. But then it's like whatever end goal that you want, our world model will say, "Okay, given what you want to achieve, let me generate a distribution of environments such that I can train and evaluate whatever it is you, you want."

Host50:11

Yeah.

Fan-yun Sun50:11

Right? Maybe for the purpose of games, it's just the end simulation, and that's the end product. For certain policies, it's like I can train it within these environments and then help you see where your policy's failing or not.

Host50:23

Yeah.

Fan-yun Sun50:23

And then... You know, so I think-

Host50:24

So in that case, much more of a training tool than in other applications.

Fan-yun Sun50:28

Training, evaluation, both, right?

Host50:30

Sure. Same, same thing.

Fan-yun Sun50:31

Yeah.

Host50:31

Yeah, same thing.

Fan-yun Sun50:31

I think it's just this world model that allows people to train any policy that can act in any multimodal environments.

Host50:38

Would it be harder to reward hack? Is there an angle here where it is harder to reward hack? Like just, I'll just put it generally. 'Cause I think that's a, that's obviously a key problem that a lot of people face when, when training agents in these environments and, I don't know, can you solve it?

Chris Manning50:54

I think not necessarily. I mean, to the extent that there's a misspecified reward that it seems like it could be hacked in a more symbolic world or in a more pixel-based world. Um, I don't know if Sun's got any thoughts, but I don't think that's really being solved.

Host51:14

The other thing that comes to mind is just you could just build a better Sora as a video generator model, right? Because then you'd, you would move the diffusion, uh, side a bit more further to the right, I think, if I got the directionality correct.

Um, and that's it. Maybe-

Fan-yun Sun51:29

It's better on domains, right? Like on consistency over an hour, for sure it exists versus something doesn't, right? So...

Host51:36

Yeah.

Fan-yun Sun51:38

Yeah. Is, is your question more like, like-

Host51:40

I'm just riffing on like how do you-- what can you build, you know-

Fan-yun Sun51:43

Oh

Host51:43

... with, with the stuff that you have. I do think that the mind of the academic does go immediately to training and in eva-evaluation, but like art tends to take un-unusual directions, like you might end up like-

Chris Manning51:55

Okay, yeah, but the question is, can you use this piece of software to develop compelling gameplay? And-

Fan-yun Sun52:03

Yes

Chris Manning52:03

... I don't think you can take Sora and produce compelling gameplay, right? If you want to have a world that you can wander around in a bit, you're good, but what are your abilities to have gameplay mechanics implemented the way you'd like them to be and to have things stay, you know, with the long-term history of your gameplay that influences future actions.

I think there's just nothing there for that.

Host52:28

Yeah. I do tend to agree. I, I'm just trying to sort of test the boundaries. I would also make the observation that as triple A games industry has developed, the line between what is a movie and what is a game has blurred.

Um, and you, you, you do end up basically producing a two-hour movie as part of your game.

Fan-yun Sun52:47

Um, no, honestly, there, there are so many actually applications in adjacent markets-

Host52:52

Mm

Fan-yun Sun52:52

... that our world model can go into.

Host52:55

Yeah.

Fan-yun Sun52:56

But yeah, it's, it's sort of fun to rif-riff on, although on the execution side, we sort of-- we, we need to stay focused with like, okay, what are the capabilities we wanna unlock over time, and there's a roadmap for that.

But yeah, if we're just riffing on sort of like the possibilities, I feel like whether it's endless, yeah, it's like-

Host53:09

Classic.

Fan-yun Sun53:10

And then-

Host53:10

The embedding for possibility and endless in my mind is very close.

Fan-yun Sun53:15

Yeah.

Host53:16

I do wanna ob-- uh, focus on one like weird choice. I, I don't know if it's weird, and maybe I'm-- I got something here. Uh, audio, right? You could have just said no audio. And audio in my mind has a lot of recursion, whereas in, in video, you can just do ray casting, and that's much computationally much simpler.

Audio & Roots53:30

Host53:33

Audio just seems way harder. I don't know if you wanna just comment on just the spatial 3D audio problem. Did you really have to do it? I, I guess you do to be immersive, but like a lot of people do treat it as like, well, we'll just stick a, a TTS model on top of...

Fan-yun Sun53:51

Well, there's a lot more to game audio than just speech, right? It's not just TTS like-

Host53:54

Yeah, TTS, SFX, BGM.

Fan-yun Sun53:56

Yeah, yeah, yeah.

Host53:57

Spatial, in my mind, echoes.

Fan-yun Sun53:59

Yeah.

Host54:00

And reflections, and I, I don't even know what's, what else. I don't, I don't know what, what are the problems in this space.

Fan-yun Sun54:07

Yeah. I think this point, like the-- it's sort of a more, more pointing to the benefits of using a game engine as a tool that's available to the model, right? 'Cause like part of the spatial audio is from the code that is underlying the simulation.

And while we do give our model access to other types of audio models as tools-

Host54:33

None of them would be spatial, I think.

Fan-yun Sun54:35

Right. But that's exactly sort of more point to we're giving our model an abstraction or a suite of tools such that it's able to achieve that. And you can argue that sort of spatial is like a, like a emergence out of the, the tools that we-- and abstraction that we provide to the agents.

And I think that's the beauty of this, this, this approach, is like there's a lot of things kinda like how humanity's built Technology and they're like Lego blocks that build on top of each other, and it's the same thing here.

Like there's gonna be things that so just sort of emerges from being able to put these things together in like a combinatorially interesting ways.

Chris Manning55:09

Right. So this integrated audio model exploits the understanding and semantics of the Moonlake world, right? And whereas in general for the Gen AI video models, there's no actual integration across to audio at all, right? That someone might stick some music or stick a soundscape or whatever else on top of their video so it's not a silent video, but they're in no way connected into a consistent world model and there's nothing that's okay, an action is happening in the video, therefore, there should be a sound that's coming from this part of the visual field.

Host55:58

Yeah.

Host 255:58

Is that different than Sora 2? Does it not have audio? Not to say it's not like amazing-

Host56:03

It doesn't have spatial audio.

Host 256:04

It doesn't?

Host56:05

No. I've, I've played around with it enough. It just sounds like someone put an ElevenLabs voice on top of it and just tried to s- do the lip sync.

Host 256:13

Yeah. I've seen, okay, generate a dog at the beach and reactions to big wave and move around.

Host56:18

Yeah, yeah.

Host 256:18

It's definitely like early, early-

Host56:19

So, so have the dog, have the dog move away from camera and see if the, the sound goes down, right? It doesn't, right? 'Cause they don't have spatial audio.

Fan-yun Sun56:28

We do want to basically like we-- our world model, like the one we're training, is basically towards the goal of having a combined latent representation across all these different modalities, right? Such that it can like reason across these different modalities.

Um, so for example, if I close my eyes and like you play a video-- uh, you play a sound of like a car skidding away from me, I almost can like visually extrapolate-

Host56:50

Yeah

Fan-yun Sun56:50

... that trajectory in my mind. And I think that, that type of capability we want our model to be able to reason, right? And that's the reason that we're sort of taking this multimodal reasoning approach. It's like we want this combined latent space that can-

Host57:01

Yeah. Oh, you said latent space and we like that here. We have, we have to play the, the bell every time there someone says latent space. Uh, no, you gotta train Daredevil one where you, you, you-- it's only audio, but you have to work out where everything is.

Cool. I, I think that was, that was, uh, that was about it for our Moonlake coverage. Uh, I do think, uh, we have like a couple of, uh, Chris Manning questions on, on IR and, uh, just any, any other sort of attention topics or NL- NLP topics.

Host 257:29

Okay.

Host57:29

Go ahead.

Host 257:30

Well, no, I mean, yeah, it's just fun. Uh, you know, we talked a bit about how you guys met, but you basically, you, you are like the godfather of NLP per se, right? You spent the whole career from early embeddings, early, early attention.

You did 2015 attention for machine translation, everything. Uh, you, you had information retrieval, so RAG before RAG. You know, we just wanna shout that out and admire a lot of that, right? So what prompted the switchover to world models?

How did, how did all that come about?

Chris Manning57:57

To some answer it is, um, the enthusiasms and creativity of students. But there's a bit of a history there, right? So yeah, so clearly most of my career has been doing stuff with language and, you know, how I got into research was thinking, oh, this is just so amazing how humans can produce speech and understand each other in real time, and somehow they manage to learn languages when they're kids.

How could this possibly happen? And so yeah, starting off, I was very focused on language, but you know, as it sort of got into the 2010s, I started, you know, going-- I'd been working on question answering, and then I started to get, um, interest in visual question answering, and that was an area where it was very noticeable that the visual understanding was bad, right?

You know, these were the days when like it sort of seemed like there was almost no visual understanding. You were just getting answers that came from priors. So you know, if you asked how many people are sitting at the table, it always answered two regardless of how many, how many people you could see in the picture.

And, you know, so it seemed like, oh, these models actually aren't able to get semantic information out of ima- images. And so I was interested in that problem and tried to work more on that. And so then that required knowing more about what's happening in vision and how you can represent visual information.

And then things start-- you know, there started to be this revolution of, um, doing generative AI images, and then I had students that started looking at that before the era of Moonlake. I was also working with Demi Goure, who founded Pika.

Um, and so, um-

Host59:54

And Ian obviously with GANs.

Chris Manning59:58

Yeah. Though Ian was never my student, but yeah. Ian-- I, I was very aware for the, the whole decade there of Ian with GANs, yeah. And I mean, Ian was a Stanford undergrad, but yeah.

Host 21:00:08

Richard does you.com, I believe he was your student. Um-

Chris Manning1:00:11

Um, yeah.

Host 21:00:12

Yeah.

Chris Manning1:00:12

And you know, there were, there were links across at that stage as well. So I mean, you know, there were several papers in that era of doing-- I mean, so Andrej Karpathy was a, um, PhD student at the same time as Richard, and so there was some joint language vision work in that era as well.

You know, it seems kind of ancient by modern standards, but yeah, we were trying to go from sort of textual dependency graphs to visual scenes.

Host 21:00:40

At a time the glove embeddings really took over a lot of TFIDF, like one hot encoding, all that. The early vision language models we saw were like lava style adapters, right? It's, it's technically still just embedding latent space.

Let's add image. So it's like mixed modality. So-- And that, that's one of the things you super put out there too, right?

Chris Manning1:00:59

Yeah.

Host 21:00:59

Yeah.

Host1:01:00

Yeah. Well, thank you for all of that. Thank you for advancing the world on, uh, world modeling. Uh, I honestly, uh, do think that if people deeply understand everything we just covered, they will see what's coming And I think you guys have, you know, made some, uh, s- really significant contribution here.

Hiring & Name1:01:00

Host1:01:15

What are you hiring for? You know, what, what is the-

Fan-yun Sun1:01:17

What do people find?

Host1:01:18

You know, we, we agreed that the CTA was a hiring call. Yeah, I mean, don't we have AGI? You don't need, you don't need engineers anymore, right?

Fan-yun Sun1:01:25

Yeah. On the model side, we, we are actually striving towards basically a self-improving system, but what that means is that we need people to set up the self-improving system. Um, so more, more specifically, people who have the intersection of knowledge within code generation and computer vision and graphics, right?

Host1:01:42

Yeah.

Fan-yun Sun1:01:43

That's, that's sort of the core research background that we look for within our team. And, and the majority of the team today do have, like, s- both backgrounds. Um-

Host1:01:51

When you say computer vision and graphics, are they the same thing, or is it computer vision one thing, graphics another thing? And how intertwined are they?

Chris Manning1:01:59

They're intertwined but different.

Host1:02:02

Yeah.

Chris Manning1:02:02

And I think, you know, this relates to some of the themes that we've been talking about, that the more explicit underlying world models that are being constructed inside Moonlake really draw on the computer graphics tradition, and so it's then combining that with the visual understanding of vision.

Host1:02:29

Got it.

Fan-yun Sun1:02:29

Yeah.

Host1:02:30

All right.

Fan-yun Sun1:02:30

So I think-

Host1:02:30

So if you've written a game engine, you're-- Come talk to us, right?

Fan-yun Sun1:02:34

Oh, yeah. Def- definitely. But I do think that the line is blurred, like increasingly blurred these days, where it's like if you have a general understanding of vision and graphics.

Host1:02:44

I think for your standards it is. Uh, for me, it feels like vision is, is... You know, I, I'll leave that to the big labs. Graphics, I, I, I can get that, you know, you would want to do that for more first principles.

But vision, there's so many vision models off the shelf that I can take, but m- probably not good enough for your-

Fan-yun Sun1:02:59

I see. I see. If, if you're sort of, like, making that distinction, then, then maybe we, we care a little bit more about having graphics knowledge.

Host1:03:07

Yeah, exactly. Exactly.

Fan-yun Sun1:03:07

Um-

Host1:03:07

Yeah. It could be like, uh, you know, sometimes a hiring call can be as simple as, like, if you know the answer to blah, you should talk to me. You know? Like the, the, the, the sort of core known hard problem in, uh, in your world.

Fan-yun Sun1:03:17

Ah, I see. Yeah. In that case, if you... Yeah, definitely if you've written a game engine before, if you've RL'd a variety of coding models on different objectives, like-

Host1:03:29

Easy

Fan-yun Sun1:03:29

... um-

Host1:03:30

Many of those, yeah

Fan-yun Sun1:03:31

... if you've done multimodal in space alignment. I, I intentionally included that space again.

Host1:03:38

Yeah. Our poor editor has to edit thing every time. Uh, yeah, lean space alignment. Hon- honestly, is it that hard?

Fan-yun Sun1:03:43

Well-

Host1:03:44

I, I-- There's some scripts out there that I've saved for the day I someday, someday have to do it, but I don't have to do it. But it's done.

Fan-yun Sun1:03:51

I, I think, yeah. There, there's, there's, uh versions of that that are done. Uh, but I, I think we are aligning audio, text, language, and video, right?

Host1:03:59

Yeah. Yeah, yeah, yeah.

Fan-yun Sun1:04:00

Like, and basically, we have these world models that are able to act as agents to, like, act in these worlds and extract long horizon videos-

Host1:04:08

Yeah

Fan-yun Sun1:04:08

... and encoding that back to the model to sort of self-improve. So it's an insanely exciting but also technically challenging problem.

Host1:04:15

Yeah.

Fan-yun Sun1:04:15

So people who wanna do their life's best work, you know, at Moonlake's the place.

Chris Manning1:04:19

How big are you guys? Where are you guys based?

Fan-yun Sun1:04:21

We're currently based in San Mateo, although we're moving up to SF. Um, we're about 18 folks right now.

Host1:04:27

My ending question was gonna be why M- what, what is the name and what's behind the name?

Chris Manning1:04:30

Yeah.

Fan-yun Sun1:04:31

Oh. Um-

Chris Manning1:04:33

Very cool graphics and design, by the way

Fan-yun Sun1:04:35

... actually, at the, at the time when the, when the, when we started the company, we were thinking a lot about how do we make a company name that gives people the vibe of, like, OpenAI but for, like, almost, like, industrial light and magic vibes.

Host1:04:48

Wow.

Fan-yun Sun1:04:48

'Cause it's like we care about creativity and using that as a funnel to solve AGI. So then we were-- we, we brainstorm a lot around, like, Dreamworks, right? Like industrial light and magic and, um... So there's a few, few basically, uh, space of things that we feel like are very, very semantically close to-

Host1:05:06

Yeah

Fan-yun Sun1:05:06

... the company's identity.

Host1:05:08

Yeah.

Fan-yun Sun1:05:08

And then it ended up being Moonlake partly because of the Dreamworks vibe, you know, the Dreamworks, uh-

Host1:05:15

Moon, lake.

Fan-yun Sun1:05:16

Exactly.

Host1:05:16

Yep.

Fan-yun Sun1:05:16

Um, so that was a little bit of that inspiration. And then the moon was sort of like a-- it basically was, like, about the reflection. The reflection part also implies the self-improvement-

Host1:05:28

Mm.

Fan-yun Sun1:05:29

... loop-

Host1:05:29

Wow

Fan-yun Sun1:05:29

... that we sort of like-

Chris Manning1:05:30

Wow.

Host1:05:30

That's very Moonlake

Fan-yun Sun1:05:30

... really believed in, and that's the path towards multimodal general intelligence. So that's, that's, that's that. I'll leave it at that.

Host1:05:36

I love a good name. I love a good name.

Fan-yun Sun1:05:37

This is great.

Chris Manning1:05:38

It's a very good name.

Host1:05:38

It's very good lore. I'm glad I asked the question. I, I will also say, you know, one of my favorite story, uh, books or biographies ever is, uh, Creativity Inc. with Ed Catmull's, uh, story about Pixar and how he w- you know, was rejected as a Disney animation artist.

Fan-yun Sun1:05:54

Mm.

Host1:05:54

So then he went into computing and brute forced his way into, back into-

Fan-yun Sun1:05:58

No, I love that story

Host1:05:59

... Disney.

Fan-yun Sun1:06:00

Yeah. And Walt Disney is also, like, one of my favorite founders. He's like, his, his story... Like, at the time you're like, "Okay, I'm gonna create this, like, immersive park." Like, people can't, can't-- don't even have that technology to create it virtually.

But they're like, "You know what? Let's just build it very physically such that people can-"

Host1:06:13

So he's the first world modeler?

Fan-yun Sun1:06:16

Um, no. I, I, yeah, I tell people that. Like, theme parks are world models too.

Host1:06:19

Mm. Yeah, yeah, yeah. I mean, uh, you know, it's a small world or it's, uh, like the Epcot Center with all the little, um, replicas of the countries. Yeah, those are very interesting. Um, okay. Well, thank you. Uh, we've covered, uh, you know, a, a huge amount.

Thank you for your time, and thank you for inspiring us.

Fan-yun Sun1:06:35

Thank you.

Chris Manning1:06:35

Thank you for having us.

Fan-yun Sun1:06:36

It's fun chatting.

Chris Manning1:06:37

Yeah, it's been a good time.