LALatent SpaceDec 6, 2025· 1:04:51

World Models & General Intuition: Khosla's largest bet since LLMs & OpenAI

Pim de Wit, founder of Medal and General Intuition (GI), turned down a reported $500M offer from OpenAI to spin out GI with a $134M seed from Khosla Ventures — Vinod Khosla's largest bet since OpenAI — arguing that world models trained on peak human gameplay are the next frontier after LLMs. Medal's 12M users generate 3.8B action-labeled clips via retroactive recording, creating a privacy-preserving dataset of 'episodic memory for simulation.' GI builds fully vision-based agents that see only frames and output actions in real-time, using pure imitation learning without RL, and can transfer from arcade games to realistic games to real-world video. Pim explains why world models need actions, memory, and partial observability (e.g., smoke, camera shake) compared to video generation, and how they distill giant policies into tiny real-time models that navigate and hide like humans. He recounts his path from running the largest RuneScape private server to reverse engineering, cold-emailing the Diamond (world model) paper authors to assemble a top research team, and advises data founders to train models themselves before selling. GI's near-term customers are game developers replacing…

  1. 0:00Intro
  2. 2:42Demos
  3. 14:17Pim & Medal
  4. 20:33Scaling Medal
  5. 23:57Team Assembly
  6. 33:29Khosla's Bet
  7. 34:58Data & Learning
  8. 41:03World Models 101
  9. 43:44Contrasting Approaches
  10. 50:30Go-to-Market
  11. 57:42The Road Ahead

Powered by PodHood

Transcript

Intro0:00

Pim de Wit0:00

You know, in a video model, you might predict the next likely sequence or the next most entertaining frame. What world models do is they actually have to understand the, the full range of possibilities and outcomes, um, from the current state, and based on the action that you take, generates the next state, right?

So the ne- the next frame. And so it is a, it is a much more sort of complex problem than, than traditional video models. So to me it is, it is a world that is accurately generated based on the actions that you take as a result of what's already been generated.

Host0:34

Hi, listeners. As you may know, I recently wrapped up the AIE Code Conference in New York, and while I'm traveling, I do like to visit top AI startups in person to bring you interviews that you don't find on any other podcast that just does a Zoom call.

General Intuition, or GI for short, is a spin-out of a 10-year-old game clipping company called Medal, which has 12 million users. But in comparison, Twitch only has seven million monthly active streamers. Medal collects this data by building the best retroactive clipping software in the world.

In other words, you don't need to be consciously recording, you actually just have Medal on in the background while you're playing, and you hit a button to clip the last 30 seconds after something interesting happens. It's very similar to how Tesla and self-driving does bug reporting, if you ever, ever done a self-driving bug report in Teslas.

The result is that Medal has accumulated 3.8 billion clips of the best moments and actions in games, resulting in one of the most unique and diverse data sets of peak human behavior actively mining for the interesting moments. They were also very prescient in navigating privacy and data collection concerns by mapping actions to these visual inputs and game outcomes.

As you saw on our Fei-Fei Li and Justin Johnson episode with World Labs, and with the recent departure of Yann LeCun from Meta, there's a lot of interest in world models as the next frontier after LLMs to improve on spatial intelligence and to work on embodied robotics use cases.

DeepMind has been working on this with Genie 1, 2, and 3, and SIMA 1 and 2. And this year, OpenAI seem to finally agree, because they have been pinning on LLMs a lot, and they made news by offering $500 million for Medal's video game clip data.

Our guest today, Pim, turned down that money and instead chose to build an independent world model lab instead. Khosla Ventures led the $134 million seed round, which is Vinod Khosla's largest single seed bet since OpenAI. We're able to get an exclusive preview of GI's models, which unfortunately we cannot show you directly.

But I can confirm they were incredibly human-like, and we chose to include the first 11 minutes of the demo discussion, even though I couldn't show it to you. It may be hard to follow, but I tried to call out what was noteworthy for you to know as your likely reaction if you were watching along with us.

Now enjoy the world's first look, and my first look, at General Intuition.

Demos2:42

Pim de Wit2:42

So what I'm about to show you is a completely vision-based agent that's just seeing pixels and predicting actions the exact same way a human would. Um, and so yeah. What I'll show you here is what this looked like, uh, four months ago.

So this was, uh... So again, this is just an agent that's seeing, that's receiving frames, um, and it's just predicting action. So you can see it has, like, a, a decent sense of, um, uh, of, of being able to, you know, navigate, uh, around.

Um, it tabs-

Host3:08

This-

Pim de Wit3:08

... a scoreboard, just like gamers always tab the scoreboard. So these are purely-- these are pure imitation learning.

Host3:13

I see.

Pim de Wit3:13

So this was-

Host3:14

At least slicing the knife.

Pim de Wit3:15

Yeah, exactly. So it's doing everything that, like, humans would. In this case, here's, here was the f-first interesting part that we saw, like, it gets stuck and then it has-- they have memory as well, so you see it can get unstuck.

Um-

Host3:24

How long is the memory?

Pim de Wit3:25

Uh, four seconds. Yeah, uh, four seconds from frame to frame. Okay, so this was four months ago. This was maybe a few weeks after that. So you can, you can see there is like-- it's still doing the scoreboard thing, but it's-- there's still, there's still, uh, uh, quite, quite like...

Y-you... And these are bots too, so you can see that-

Host3:40

It's very human, let's just say that.

Pim de Wit3:42

Yeah. Uh, and then, um, uh, right, so this was really, like, the early days of research where you can see, right, it does one thing and then goes for another. Um, and then we've been scaling, right, um, uh, on, on data and compute, and also we've just been making the models better.

Um, and this is where we are now. So what you're seeing is, um, pure, like I said, pure imitation learning. This is just the base model. There's no RL, no fine-tuning. Uh, this model sees no game states. It is purely capable of-

Host4:11

Like sequence, uh, existence.

Pim de Wit4:13

It's purely predicting the actions, um, from the frames. That's it. Um, and it-- this is playing against real humans, uh, just like, um, uh, like a human would play. And it's also-- it's running completely in real time. So there's absolutely-- Everything here plays exactly like a human.

Host4:34

Do you give it a goal?

Pim de Wit4:36

No.

Host4:36

It just figures out its own goal-

Pim de Wit4:37

No

Host4:37

... because obviously it's trained on the same tasks.

Pim de Wit4:38

Yes. Um, and I, I, I, I picked, right-- I picked a sequence where also it doesn't do well initially. So you can see, like, this is, this is just like a, a sequence, a random sequence.

Host4:47

But this is the-- the-- I mean, it's, it looks like it's doing well.

Pim de Wit4:51

So, um-

Host4:52

Oh, okay. Yeah, watch.

Yeah, this is pretty good. Maybe too good.

Pim de Wit5:04

Um, this is my, my favorite part. So you can see it does something that, like, um, here, like a human would never do this. Then gets unstuck, then has four, realizes, switch, and then in the distance.

Host5:23

So you're saying, one, it makes a mistake that a human would never make-

Pim de Wit5:26

Mm-hmm

Host5:26

... but it unsticks itself.

Pim de Wit5:27

Mm-hmm.

Host5:27

And two, what we just saw is it is doing superhuman things.

Pim de Wit5:32

Yeah.

Host5:32

Okay.

Pim de Wit5:33

Yeah. Um, I mean, there are things that, that, that humans said obviously. Um, but because it is trained on, on the highlights of things that all the exceptional things, it's inherited those behaviors.

Host5:42

Ah.

Pim de Wit5:43

Yeah. So it's not like Move 37 where we RL their way into something, but it's-

Host5:46

But just, yeah, we're replicating a superhuman-

Pim de Wit5:48

Yeah, exactly

Host5:49

... or like peak human.

Pim de Wit5:49

The baseline of our data set is peak human performance.

Host5:52

Yes.

Pim de Wit5:53

Yeah. Um, okay. So that w- that's the agent. Uh, so now what I'm gonna show you is we then are able to- Take those action predictions, and we're able to label any video on the internet using those actions.

So, um, um, and so this is, this is just frames in, actions out. Yellow is the, uh, model prediction-- or sorry, yellow is ground truth, purple is the model prediction, and then bottom left is, uh, compound error over the entire sequence, and then this is reset per prediction.

Host6:30

Reset meaning you ought-

Pim de Wit6:32

Yeah

Host6:32

... every-- now and then you reset?

Pim de Wit6:33

Yeah. So this just means it resets the baseline. Um, and so this basically, a, a single error in the entire sequence compounds here, but it doesn't compound here.

Host6:41

Yeah.

Pim de Wit6:41

If that makes sense.

Host6:42

Sure.

Pim de Wit6:43

Yeah. Um, so and again, this is just seeing frames, right? It's, it's not, it's not seeing any of the, any of the actions. Um, and so, you know, so what we did, right, is we, we, we trained it on less realistic games, and we transferred it over to a more realistic game.

And then, and this is where it gets really exciting, we transferred it over to a real-world video, which means that you can use any video on the internet as pre-training.

Host7:07

What was it predicting?

Pim de Wit7:08

Um, it's predicting it as if you were controlling it using keyboard and mouse. So if you were, if you were basically playing this sequence as a human.

Host7:16

There-- And is, is there some sense of error or?

Pim de Wit7:19

Uh, so that's why you, you transfer it to more realistic games first-

Host7:22

Yeah

Pim de Wit7:22

... and then you transfer it to real-world video because you can't get a sense from ground truth from, from the real-world video yet. Um, let's see. And then, um, so we don't-- they don't-- So I'll show you here.

Uh, this one is also, um, this is the same, uh, uh, agents that I just showed you.

Host7:43

This is playing against other AIs.

Pim de Wit7:45

Th-this one's playing against bots, yeah. Um, the previous one was against players. Uh, but with the sniper, it doesn't really matter that much, as you'll see. It's like, uh-- So one, one thing that's really interesting is you notice that it behaves differently as it has, like, different items, right?

Host8:02

That makes sense.

Pim de Wit8:02

Yeah.

Host8:03

Intuitively.

Pim de Wit8:04

Yeah.

Host8:05

I think there's, uh, also a question about egocentricity versus like, so the third person view.

Pim de Wit8:11

Yeah.

Host8:11

Does it matter?

Pim de Wit8:12

Um-

Host8:12

Maybe it doesn't matter

Pim de Wit8:13

... the third person I think will be very, very helpful if you're, for instance, trying to control multiple objects in an environment later on. Uh, right now I think having fully in, in perception, uh, first person is quite helpful.

Um, this one's also-- this is the policy itself.

Host8:28

What do you mean this is the policy?

Pim de Wit8:29

The agent. Yeah. Sa-same constraints that I just told you about.

Host8:32

Sure.

Pim de Wit8:32

Yeah.

Like this where, right, the, where it hides, that to me was just incredible. Like, just from, from knowing, being able to, to predict-

Host8:45

The difference is also high when you see it-

Pim de Wit8:46

Yeah, exactly. Yeah. Yeah

Host8:47

... through.

Pim de Wit8:48

Um-

Host8:49

And it needs the spatial intu-in-in intuition to go, "Well-

Pim de Wit8:52

Yeah

Host8:52

... this is hiding-

Pim de Wit8:53

Yeah

Host8:53

... and that's not hiding."

Pim de Wit8:54

Exactly.

Host8:55

Yeah.

Pim de Wit8:55

And, and, and right while it was reloading, yeah. Um, okay, so that, uh, so those are, uh, that, that's the policy, and this is a completely general recipe, meaning we can scale this to any environment. Uh-

Host9:07

Yeah. Is this work process-- Okay, no, uh, let, let, let's keep going on demos-

Pim de Wit9:11

Yeah

Host9:11

... until, um, I was gonna go into research.

Pim de Wit9:13

Yeah. Yeah, sounds good. Um, okay, so and then this is, this is, this-- So what I'm about to show you are the world models. Um, there's a few really, really interesting parts about our world models. So the first is, uh, we actually made the, uh, decision to, uh, transfer, um-- Sorry, we made the decision to, um, pre-train world models from scratch, but also we've actually been able to fine-tune open source video models to get a better sense of physical transfer.

Um, and so one of the things that you'll notice here is, like, our world models have mouse sensitivity, which is something that, like, gamers absolutely want, right? So you can have these, like, very rapid movements, which you couldn't do in any other world model.

Um, and so this is a hold-out set. So this clip was never seen before, um, at training time. Um, as you can see, it has, it has, uh, a spatial memory. This is, this is about a, a twenty-second-ish, uh, uh, generation.

And here's what's fascinating. This is an explosion that occurs, and you can see that in the, um, in the physical world, right, the camera would shake, and in the game, that would never happen. So you see, you see the, the world model, uh, inherits the, the physical world camera shake-

Host10:18

Yeah

Pim de Wit10:18

... but the, the, the actual, um, uh, uh, game never does that. Uh, which, which is, which is sort of-- That, that to us was quite fascinating, right? Also, the models that I just showed you that we used to transfer over from video, the two of those combined will allow us to, like, push way beyond games in terms of, in terms of training.

Um, this is another interesting-- So this is a world model. This is rapid camera motion. So like, again, this is stuff that we're ju-literally just taking one second from here in the context and the actions and replaying it here, right?

Um, and so you, you'll ne-- you, you never essentially have, um, uh-- Like what we're saying is the, the skill that you see in the clips, that, like, the speed and the movement, that also pays off at training time when you're doing world models.

Um, this is my favorite example. Uh, so this shows that the world model is capable, um, of performing, uh, with partial observability. So what you're gonna see is, um, uh, again, you're replaying the actions from here and here just using, uh, one second of video context.

Everything after that is completely generated. Um, so what you're gonna see is the model is going to encounter, in this case, smoke. Normally now, models break down. What you actually see is it comes out at the same place.

Um, and so it's capable of, um, of even with partial observability, still maintaining, um, uh, its position in the world. Um, and then here it is also interesting. So this is, uh, this is sniping. So this is gi-this gives you like a, um-

Host11:43

Reaction time?

Pim de Wit11:43

Uh, like the fact that it can do depth and like sequences in completely different views, right? So this is a completely different view than if you were to be outside of that view, right? And so it's, it's ga-it's able to maintain consistency, um-

Host11:56

While zooming in.

Pim de Wit11:57

Yeah, exactly. Um, uh, and so, um, yeah, so you can see. Uh, so even while th-this goes out of scope, right? Watch, and then it can-- and then it comes back, and you'll see it's still, still there.

Yeah. Um, and so, uh, yeah, this is the work that, that Anthony Hu has been working on.

Host 212:19

Um, just wondering how much game footage you have to watch in order to find these things.

Pim de Wit12:23

Um, we can ask Anthony. No, it's, it's, uh... I'm, I'm sure he's not gonna be too excited to play these games, uh, afterwards. Um-

Host 212:31

You're not playing, right? You're just watching.

Pim de Wit12:32

Yeah.

Host 212:33

Yeah, yeah.

Pim de Wit12:34

Um- Great. Okay, so those were the models. Um, let's see. These are interesting. So we also were able to distill into, like, really, really tiny models. Um, so this is for instance, a, um, a long sequence on a very, very tiny one.

You can see it makes, like, a bit more stupid mistakes. Uh, like, it do- it does things that are not as optimal. Um, but-

Host 212:57

I haven't seen anything yet that-

Pim de Wit12:58

Uh, at the beginning it was running into a wall for free. Exactly.

Host 213:01

Yes, yes.

Pim de Wit13:02

Um, uh-

Host 213:04

I mean, I do that too.

Pim de Wit13:05

Yeah. Yeah. Um-

Host 213:08

It looks-- I mean, it's doing pretty well.

Pim de Wit13:09

Yeah, and, and again, all these models are running completely in real time. There's so there's no, um-

Host 213:14

Okay. So I was thinking, your main model does real time anyway. What's the goal of distilling? Is it cost or...?

Pim de Wit13:19

Uh, yeah, parameters.

Host 213:21

Okay.

Pim de Wit13:21

Yeah.

Host 213:23

Yeah.

Pim de Wit13:24

Yeah. This is the interesting one, it peaks a corner. That's what we mean by, like, the spatial and the poor reasoning aspect, is, uh, humans actually they, they sort of simulate the optical dynamics of their eyes and how to actually spatially reason in the world.

Host 213:33

But you just have all the data-

Pim de Wit13:34

Yeah

Host 213:34

... right?

Pim de Wit13:34

Yeah.

Host 213:34

You've, you've seen all this.

Pim de Wit13:36

Yep. Um, exactly. And so, uh, like, even in, like, real wor- this is kind of interesting. Even in, like, the real world, um, with, uh, for instance, YouTube data, right? You have to s-first solve for pose estimation. Then once you have pose estimation, maybe you do something like inverse dynamics, right?

Where you basically are able to, like, somehow label some of the options that you're seeing. And then you still have to account for optical dynamics of, like, where are your eyes actually looking before the decision, 'cause there, there's just three levels of information loss.

Whereas when you're playing video games, you're actually simulating the optical dynamics with your hand, right? And I think that, like, that's why I think why games are a better representation of spatial reasoning initially than, um, uh, than YouTube videos, for instance.

Host 214:17

Okay. We're in the GI offices with the CEO, Pim de Wit. Welcome.

Pim & Medal14:17

Pim de Wit14:21

Thank you.

Host 214:21

Thanks for having us in your office.

Pim de Wit14:23

Yeah, excited to be here.

Host 214:24

If I'm in New York and you are one of the hottest raises of the year, I have to come and visit and, uh, thanks for taking some time on the weekends or-

Pim de Wit14:31

Yeah. Yeah.

Host 214:32

So you've raised 133 million seed, uh, for, for General Intuition. Most people don't hear about you. I, I guess 'cause GI is new, but mo- uh, more gamers would have heard of Medal.

Pim de Wit14:43

Mm-hmm.

Host 214:44

And before that you ran private RuneScape server.

Pim de Wit14:46

Yes.

Host 214:46

One of like the largest-

Pim de Wit14:47

Yeah

Host 214:48

... RuneScape server. Um, what's your reflection on just that, that journey of like now you're an AI founder-

Pim de Wit14:55

Yeah

Host 214:55

... and you started off playing RuneScape?

Pim de Wit14:56

Yeah. I think, um... So I grew up with Tourette's. Uh, I, uh, spent most of my time as a teenager coding and playing video games. Uh, so in that sense, it doesn't feel that much different. Um, but I think for, uh...

So yeah, so I started the largest private server RuneScape, worked at Dr. Sid Worthers for three years versus Ebola and then on like satellite, satellite based map generation for disaster response, um, uh, which was already like very AI related adjacent.

I, I built some models back then and then started, uh, Medal, which became one of the largest social networks in video games. I've always been kind of like AI, like adjacent. I, I, you know, I, I, I, I'm a self-taught engineer.

Uh, so for me, the modeling itself always felt a little foreign. I actually had to, uh, take a ton of, tons of classes over the summer and, and, and early this year to get better at it, uh, because I, it fe- it still felt like, like I was really, really good at, at the infrastructure side, and I had written like our, our, our transcoders for Medal myself.

So I was very, very familiar with CUDA and like the GPU side and all the video infrastructure that we were using, uh, for this stuff. But the, the modeling side itself was, was still quite foreign. Um, luckily, obviously we have, we have really, really good fo- co-founders, but they, they, they essentially put a bunch of coursework together for me to, to go complete to get really, really good at understanding the fundamentals better.

I think for me, I had seen inside of the labs that I had really, really good, uh, leadership with fundamentals on top and also the ones that didn't, and I think the ones that did were just like much better.

Um, and so for me, um, yeah, I wanted to be more like that. So in that sense it was a bit, it was first very foreign and then now I feel pretty comfortable with everything and, but yeah, like I think for, um, uh, there's a lot to be explored starting in video games and also reverse engine- like I think the interesting thing about reverse engineering is it kind of teaches you to look at problems, uh, very differently.

It's like the ultimate form of deductive reasoning in a way. Uh, and so um, uh, so this is just how I think, how I operate, and so for me it's, it's been a really, really interesting journey. Uh, you know, I don't claim to have any of the credentials or, or skills that some of the other guests you have had on, but hopefully it will make for a good time.

Host 217:03

Yeah. Well, your co-founders-

Pim de Wit17:04

Yeah

Host 217:04

... uh, definitely bring a lot of that-

Pim de Wit17:06

Definitely

Host 217:06

... capability and you bring a lot of the, I guess, gaming expertise or-

Pim de Wit17:10

We'll see, we'll see what expertise-

Host 217:12

Yeah.

Pim de Wit17:12

We'll, we'll see what I bring to the table.

Host 217:14

Just, just a little bit of history of Medal. Well, let's establish Medal-

Pim de Wit17:17

Yeah

Host 217:17

... for those who don't know. Uh, did nearly, uh, s- uh, Twitch-

Pim de Wit17:21

Yeah

Host 217:21

... like the year.

Pim de Wit17:22

Yeah.

Host 217:22

Um, that's you have more active users, concurrent users than Twitch, something like that.

Pim de Wit17:28

Yeah, on the creator side, I think. And the reason is because Medal is a lot more like Instagram than it is like Twitch. So people, um... So the way, the way to think about Medal is it's, it's a native video recorder.

Like unlike somebody like Twitch where you actually have to use other software to record and stream to Twitch, um, it's not a streaming software, it's actually a video recording software. And a lot of gamers love to put things like overlays on top of their videos.

Um, and as a result of that, we have sort of the largest data set of ground truth action labeled video footage on the internet by maybe one or two orders of magnitude. Yeah.

Host 217:59

What, what, what's an example of an overlay? Like the only overlay I usually think of is like face cam.

Pim de Wit18:04

Yeah, yeah. Also, um, controller overlays for instance, if you're playing, um, like let's say you're playing, uh-

Host 218:10

Console

Pim de Wit18:11

... yeah, like flight simulator, you get like, you know, the joystick and all, all the things.

Host 218:14

Yeah, yeah.

Pim de Wit18:14

So you get the actual actions that people take inside the games-

Host 218:16

Right

Pim de Wit18:16

... as well as the frames of the games themselves.

Host 218:19

Yeah.

Pim de Wit18:19

Which is a loop, right? Because it's essentially you perceive Then you act, then there's a state update, and then you perceive again, and you act, state update, which is, like, roughly precisely what you use in order to trace-- to train these agents.

Host 218:30

Yeah.

Pim de Wit18:30

Yeah.

Host 218:31

It's, it's, uh, almost perfect training data. The, uh, we, we were showing-- you were showing me in the demo when we showed some B-roll here on, uh, how you don't log key.

Pim de Wit18:39

Yeah.

Host 218:39

It is very important for you to log action.

Pim de Wit18:41

Yeah.

Host 218:41

When did you figure this out?

Pim de Wit18:43

Ooh, um, maybe starting a year and a half ago. Yeah, and, and we realized that, like, fig-figuring out the side of the research for us was we very much never wanted to be in a position where we eroded privacy or something like that, so we never wanted to actually log, like, a W or A or S and a D.

Which for researchers, the fact that we don't do that, like, often it sounds strange, like, why wouldn't you do that? But I think for us, the privacy-

Host 219:07

Well, we get the data.

Pim de Wit19:08

Yeah. I, I think, you know, a lotta, a lotta the, the, um, the researchers, they did-they hadn't quite understood yet that you can actually just get away with just doing the actions. Um, and the reason is, like, at training time, having the actual key is this noise any-anyways.

Like, if there is text in the screen and you would want to, in, in theory, uh, make that, um, part of the training, then, like, reading text from a frame is, like, really easy. And so for us, if we actually con-- So we convert-- Basically you hit, you hit the, uh, the input, we convert it to the actual action.

So we had thousands of humans label every single action you can take in every single video game over the past year and a half, uh, which is an enormous amount of action labels. Um, yeah, so when you act, you, we, we get the actual, um, action itself, and then it means that at training time, you, you can for, like, a, the general set of that, of that game, convert back into computer inputs if you want to, but you can never do it for any individual person.

And so that for us from, from, like, a design perspective was, was important. So we, we, we figured all that stuff out, then we actually started pushing, um... Like, we already had features as well with this. So for instance, like, gamers already love to be able to navigate their clips by, like, things that happened, so we have an events capture system, and then we also have the overlays where you actually just want to overlay and render the actions on top of your clip.

We developed kind of in tandem with the feature set itself, and then obviously when, when world models became a thing and it's very, very clear that all the, all the data for this was precisely, like, that sequence, yeah, we were able to sort of be first to market, recruit the best researchers, and start a lab.

Scaling Medal20:33

Host 220:33

Yeah. That's, uh, that's incredible.

Pim de Wit20:35

Yeah.

Host 220:35

Uh, one more question on Medal before I- we move focus on the AI. It's been 10 years.

Pim de Wit20:40

Yeah.

Host 220:40

What is the... I don't even know how you grow something like this, you know? I'm just kinda curious and, and, yeah, I'd like the opportunity to ask you what really worked-

Pim de Wit20:49

Yeah. Um-

Host 220:50

... that you became so, so huge? Because I'm-- you're not the only one.

Pim de Wit20:53

Yeah.

Host 220:53

But, uh, I'm sure it's performance and everything, but...

Pim de Wit20:56

A few things that really worked. I think the first was a lot of our competitors were focused on solving the social network and the recorder at the same time, and that never-- Like, our bet was really that we could get so many people to record with us that we could bootstrap the network on top of that, and that worked.

So while everyone was sort of distracted trying to bootstrap a social network, we were just focused on building a really, really good capture tool, and then we got tens of millions of people to use that, which then we were able to bootstrap a network on top of the share behaviors.

We already had, like, the profile behaviors and the share behaviors obviously, but the actual content consumption piece and, and the sharing piece really only came after we hit critical mass. It was actually early days during COVID when, like, the network really accelerated.

Fortnite happened, which was really important, and I think also the fact that Discord existed, um, uh, made it quite a different time than, uh, when other types of networks of these types had launched. 'Cause Discord essentially was, like, the connective tissue already between gamers that, like, never really existed before.

And so I think those combination of things really, really made it. I think we also built a product that, for instance, with, with most video recorders, you have to remember to start and stop the recorder. So you have to go into the application, then hit start, then start your game, and then, um, you know, maybe you'll play games for three hours.

Then you'll close the game, then you have to close your video application.

Host 222:09

Then you look for clip.

Pim de Wit22:10

Then you-- Well, then you have to process, like, a multi-gigabyte file. Uh, then you have to upload those somewhere, and so, like, this was a pain for people. And so what we did is we just ran this kind of recorder.

When you hit that button, it does a retroactive video record. So all the recording initially is in memory, and then when you hit that button, it exports only that sequence to disk and syncs it to your phone. And so that, that became super popular.

It also-- What, what was interesting about it, it also means that you're not sort of behaving or acting differently because it's always there and you can just export whatever happens, which is also very, very helpful for, for training obviously.

Host 222:41

Um, the thing-

Pim de Wit22:42

Yeah.

Host 222:42

You weren't the first to do that.

Pim de Wit22:44

Yeah. Yeah.

Host 222:44

The, the, the thing you were explaining just before this was, is similar to how Tesla does their bumper boats, right? You're driving-

Pim de Wit22:51

Yeah

Host 222:51

... from the time you disengage autopilot-

Pim de Wit22:52

Yeah

Host 222:53

... you're like-- they're like, "Well, tell us what happened."

Pim de Wit22:55

Exactly. Exactly. So you-- See, y-y-you're driving. Tesla doesn't wanna train on the, like, 10 hours of you driving through a desert where nothing interesting happens. You have the clip button on the steering wheel. Something interesting happens, either while FSD is engaged, and I'm not sure if you can use it without FSD as well.

But you hit the clip button, and it basically uses that precise sequence to mark, which is then more helpful for training because it's more unique as a training time.

Host 223:18

Yeah. Yeah. I mean, so one thing, and we're gonna get to this on the agent side. One thing that I-- that's-- that does pop up as well, a lot of life is boring. A lot of life is waiting for me.

A lot of life-- A lot of playing games is doing the boring stuff that is not clip-able.

Pim de Wit23:31

Yeah.

Host 223:32

Somehow you see the generalized fight.

Pim de Wit23:34

Yeah. Yeah. Yeah, it makes you think, right?

Host 223:37

It makes you think.

Pim de Wit23:38

Yeah. Yeah, it's also quite interesting. Like, I showed you the models, like what happens when you increase the size of the context window, um, and how behaviors actually are largely shaped by the size of the context window.

Host 223:49

Yeah.

Pim de Wit23:49

That, that to me was like one of the most interesting, uh, uh, parts about the research. Um, made me think about our own behaviors in a way.

Host 223:57

Yeah. Let's talk about also the, like, forming the team.

Team Assembly23:57

Pim de Wit23:59

Yeah.

Host 223:59

On your website, you have 12, and that's changed now.

Pim de Wit24:02

Yeah.

Host 224:02

Uh, before, there are three co-founders.

Pim de Wit24:04

Yeah.

Host 224:05

And just ta-- let's talk about how this team comes together, because you may not, because you're self-taught, you don't have that-

Pim de Wit24:11

At the beginning

Host 224:11

... network.

Pim de Wit24:12

Yeah.

Host 224:12

How you manage to get all these people.

Pim de Wit24:13

Yeah. I started reading all the research papers. By that time, I was already pretty deep into, like, having a con-- decent understanding of, of, of not world models in, in particularly-- in particular, LMs and, and transformer-based models. And so, um, there was Genie, there was SEMA.

Those two were really, really interesting. And SEMA in particular was interesting because what they do is they basically take 10 games, and then they, they have a graphic in SEMA, uh, where you can see kind of the precise actions that are inside of those games that they mapped, and I believe they found something like 100, um, uh, which are actually actions that also exist in the real world.

And, uh, what they did was they then, I believe it was specifically for navigation, they did a nine one holdout set, so they, they, they trained, um, an agent on the nine games, and then, um, they had it play the 10th game, the holdout game.

But then they also trained a specialized agent just on the 10th game, and they compared how good they did. And my-- and if, if I recall correctly, it did roughly as well, um, playing the 10th game on navigation specifically on the holdout on the nine game agent than it did on the one game agent, and that to me was really interesting because that's precisely the type of data that we had, right?

And so for us, the thinking was, okay, what if we did exactly what LMS did? What if we used, right, this, um, uh, right, so LMS were trained on predicting like text tokens on words on the internet. What if we predict action tokens on essentially what is the equivalent of the Common Crawl dataset, uh, but for interactivity?

Host 225:42

Vision input.

Pim de Wit25:43

Yeah.

Host 225:43

Action output.

Pim de Wit25:44

Correct. That's it.

Host 225:45

Well, I, I think, well, actually, I'm gonna double back a little bit to like a question I had, which is one of the, one of the reasons why I thought you would want to prefer keyboard and mouse over actions is the action space is potentially unbounded, right?

You can jump, walk left, walk right, but then also look up, look left, look right. It's, it's unbounded. So it's huge, isn't it?

Pim de Wit26:07

Yeah, I think-

Host 226:07

Problem?

Pim de Wit26:08

Yeah. There, there's benefits to the action space being small to start with. So I think we're g- we're gonna start with anything that you can control using a game controller.

Host 226:15

Yeah.

Pim de Wit26:15

But yeah, long term, we want to actually predict maybe like action embeddings and have models sit inside a general action space to be able to transfer out to other inputs as well.

Host 226:24

Got it.

Pim de Wit26:24

Yeah.

Host 226:24

Okay. And then let's, let's keep going on the, on the research side. So, uh, Genie, SEMA.

Pim de Wit26:28

Yeah.

Host 226:29

And then the co-founders.

Pim de Wit26:30

Yeah. So there was the Diamond paper, there was Genie, and then there was SEMA. The Diamond paper for me was really interesting because they had actually managed to get this world model, uh, called Diamond running on a consumer GPU, I believe it was a 4090, at 10 FPS, and you could play it, and they did that on like 90 hours of data, like 95 hours.

I think it was 87 hours and I think eight in the holdout set, or something like that. That was just incredible, right? That they had something playable on that little data. So I actually cold emailed the entire group of students, and I was-- and, and I, I told them, "Hey, I think we have this thing."

And then it was pretty interesting. So like right when that happened, a lot of the labs, uh, also started un- started understanding what we had, and so we started very aggressively. Multiple labs tried to bring us in in various ways and, and they were part of that.

Like, they basically were seeing that happen, and I think for them that also kind of like solidified how real it was. And then when we chose to do our own thing, you know, initially we thought that we were gonna have to just work on world models, right?

So, so we thought, okay, the main benefit of this data set is, is, is like Genie, is, is, is world models. What we didn't realize at the time is that we have so much of this data is that we can essentially do these world models in parallel and take the equivalent of like the LM bet mostly on imitation learning, and then use the world models after that to get into like RL stage, right?

And so for us-

Host 227:49

And eventually get rid of the world models. This is something that you can-

Pim de Wit27:52

I, I mean, ideally you get rid of the imitation, yeah. The, the imitation learning. But yeah. We essentially realized that, that we could get so far on just imitation learning. The way to look at it is we essentially, like let's, let's take the LM analogy.

We essentially have sort of the internet or like Common Crawl, if you will, and every single lab is trying to simulate that, right? In order to get similar data, in order to train their agents. And so for us, the reason why we stayed independent and we just did our own thing was we think we could essentially leap every single company that's forced to either be consumers of world models or, or build world models and, and, and take this, take this foundation model bet for spatial temporal, spatial temporal agents.

Uh, and be in a place where, you know, we have a lot of customers years before any of the labs even get there. And m- maybe the most similar, um, uh, comparison is like when Anthropic did with code, right?

Anthropic just focused really, really hard on nailing the code use case. Their models are incredible for it. A lot of their customers use it for it. So we just want to become incredible at this spatial temporal agent use case, and likely that starts in like game simulation and then using world models, we can then start expanding out to, to other, um, areas.

Host 228:56

Would you show me a little bit of how it does generalize out-

Pim de Wit28:59

Yeah

Host 228:59

... games?

Pim de Wit29:00

Yeah.

Host 229:01

Um, but although games is kind of the current player.

Pim de Wit29:03

Yeah. Games and simulation. Um, I would, I would, I would specify it as game engines in particular. So even if you're, for instance, uh, simulating human behavior in Omniverse because they're trying to create better training data for factory floors, um, you can use it.

Host 229:16

Yeah. Maybe Meta has a similar dataset because of the Quest.

Pim de Wit29:20

I never really asked them. I never really looked into the Meta Quest specifically. So you need a few things. You, you can't just... Like there's lots of companies that have like maybe recorders, but you also need the public graph, otherwise you can't train on the data, right?

You can't train on people's like private videos that they have saved somewhere, right? And so I think you, you, you need the social network graph components, um, because these videos need to be on the internet.

Host 229:42

To, to rank?

Pim de Wit29:44

No, to train on them. Yeah. I, I, I mean, I think, I think generally people don't, like people don't want to train on like... Like, 'cause these things, they live on your device usually, right?

Host 229:51

Yeah.

Pim de Wit29:51

Um, and you can't train on anything that lives on your device. Like you actually need to go and upload and do your thing, right? For Meta specifically, I think also VR, the scale of VR is still pretty small.

The amount of, um, environments in VR that are, that, that have like consumption at scale is probably in like the hundreds, um, whereas on PC it's probably in the tens of thousands, right? And so you get a lot less diversity.

Um, the three-dimensional input space of VR is pretty interesting. We see some of this too, obviously. And so yeah, I, I do suspect, you know, Meta, Meta starts using these types of things, but it's unclear to me whether they can get to like a, a similar scale of data or diversity on the environments as we can.

Host 230:31

Yeah. There are a lot of challenges there.

Pim de Wit30:32

Yeah.

Host 230:33

Um, okay. I wanna take this in, in, like, a few different ways, but I guess le- let's, let's fill out the, the papers. Uh, maybe one more to mention is TAIR.

Pim de Wit30:41

Yeah.

Host 230:41

Which, uh, I actually-- I interviewed the Gaia authors, but Gaia 2 seems like the particular, uh, insight that, that could have brought it overseas.

Pim de Wit30:50

Yeah. So, so Anthony Tu, who led the, um, research on Gaia 2, is, is also one of the engineers that joined our team. Uh, so it's all the Diamond, uh, the core contributors for Diamond, and then Anthony, um, and we just had three more researchers join this week.

It's been a good week. And yes, I, I think a, a lot of the approaches in Gaia 2 were heavily inspired by Diamond. And then Vincent, who was one of the, um, authors of Diamond also already was at Wave by the time that I emailed them.

Anthony also realized what this was and realized that, that, you know, you could scale world models to a much larger, like, scale and decided just to, to make the leap as well. So I think everybody that sees the dataset makes the leap because it's...

Uh, but it takes a while to wrap around-- wrap your head a-around it because it's like, oh, it's video games, right? Like, intuitively, it doesn't make sense. And then when you actually understand and you see, right, how we've been able to transfer it to physical world video and things like that, then it makes sense, and then everybody tends to jump at that.

Host 231:41

Don't call it video games. Call it RLMBs and then-

Pim de Wit31:43

Yeah. If I lived in San Francisco, maybe I would, yeah.

Host 231:48

Uh, just a quick note is that we actually cover all these papers in, in the Station Super Club.

Pim de Wit31:52

Yeah.

Host 231:52

Uh, SIMA 2 did not seem to have as much impact on SIMA 1, and I don't really know why. They did a lot more work. Gemini 3 had a ton of impact and but I, I also felt like because you can play with the model or people, it just seems-

Pim de Wit32:06

Yeah

Host 232:06

... an extension of all those things. But I guess, like, a-any quick takes on SIMA 2 and Gemini 3, which were both this year's like-

Pim de Wit32:12

Yeah. I'll, I'll talk about SIMA 2. The steerability of SIMA 2 was to me the most impressive part because lining up the action sequences and the, the text conditioning is, is quite hard to do, right? And so that-- And the fact that they were-- Like, it's also quite interesting that that means that they can sort of use Gemini as, as part of the flywheel, right?

Where, um, where you can sort of scale-

Host 232:34

Right

Pim de Wit32:34

... scale this orchestrator as, like, an independent, almost like a puppet master, if you will. And then, like, in theory, Gemini could orchestrate many instances of, of SIMA, right? That to me is the most, most interesting part is where I, I tend to agree with this, where, like, I think our models will initially be used as, like, uh, like you'll have, like, a-an orchestrator VLM of sorts that's kind of like managing instances and instructing them.

Um, and I think sort of SIMA showing that you can do this was, was fascinating. Also, the fact that you could, um, they didn't just have text conditioning, but they also were able to do, like, drawings and markings, uh, of where to go.

They really took an interesting end-to-end approach to me, uh, that I, I look forward to seeing a lot more of. Um-

Host 233:17

Are you talking to them? Like you said, is everyone collaborating or-

Pim de Wit33:19

Yeah, I, I think the, um... Yeah, we're very friendly with DeepMind. We like them a lot. I just saw the team not too long ago, and I think, um, you know, big fans of their work.

Host 233:29

The, the headline that I extracted from Alex Heath's coverage of you-

Khosla's Bet33:29

Pim de Wit33:32

Yeah

Host 233:33

... is, uh, you are the biggest bet that Vinod Khosla has made since OpenAI.

Pim de Wit33:36

Yeah.

Host 233:37

How did that conversation start?

Pim de Wit33:39

Okay. So what-- Vinod's style, and may-maybe I'll get slapped on the fingers for revealing this or whatever, but, uh-

Host 233:46

Forgive me if I, if this is bad

Pim de Wit33:48

... um, is he asks you to, like, draw a 2030 picture of your company, and I think he just picks N plus five years, but whatever. I don't know. Um-

Host 233:56

I know. He did the same to you.

Pim de Wit33:57

Yeah. Um, he asks you to, like, walk that back from first principles all the way from today. And, and, and he asks-- he expects you to do that flawlessly, where he can challenge any assumption, any part of the vision that, that...

And he asks you questions, right? He has a very technical background. He also has a bunch of technical people on his team. And he truly backs people that have these, like, very large visions on that vision and the ability to, to defend it alone.

Um, and that's what he did for us. Um, and I think that's why he made that bet. So I think also through this, uh, through, through this question, he, he, he gets to know a lot of things about how technical you are.

He gets to know how well you think from first principles because if that i-- if that vision is not connected to something real, it's very easy to suss it out by asking good questions. Um, and then, and then he just backs fully, I think.

Like, he, he really gets in your corner, um, if it's the right fit. And yeah, they've, they've been incredible partners. They, they, they've opened so many doors for us.

Host 234:58

I have to ask the question, I think just 'cause like it's, it's a, it's a very notable story. Uh, obviously-

Data & Learning34:58

Pim de Wit35:02

Mm-hmm

Host 235:02

... a lot of work went into it-

Pim de Wit35:03

Yeah

Host 235:04

... and but it's also worth it in, in the amount of-

Pim de Wit35:06

Yeah

Host 235:06

... upside.

Pim de Wit35:07

For sure.

Host 235:08

One of the things I also wanted to-- I, I think I kind of asked this question out of sequence, but, um, one of the things that excites me about talking to you is there are a lot of people like you who are founders of business and businesses that along the way have a ton of data, and yours happens to be highly valuable.

You pursued-- Before deciding to do an independent journey, you also talked to other companies about potential licensing or acquisition and stuff like that. What is your learnings from those periods? Also, like, one, one version of this is very simply how do you value data?

Pim de Wit35:43

Yeah. I don't think you can value it unless you actually model it yourself and see what the capabilities are. That's my, that's my real outcome.

Host 235:51

You say model, like train a model?

Pim de Wit35:52

Yeah. But that's obviously, like, not doable for, for everyone.

Host 235:56

Yeah.

Pim de Wit35:56

Um, and also I think my general advice would be as model capabilities increase, you... And models are also like, you know, these, these VLMs, they're, they're very, very good at labeling as well, generally, right? What I was afraid of when I was having some of these conversations was okay, like, you know, as, as the cap- the capabilities increase, you're just gonna need less, uh, ground truth data, and, like, you can do more model-based data generation or synthetic data generation.

I would recommend if you're gonna do large data deals, like, just try to get, like, a large chunk of equity in the company that you're doing it with, um, if you can. Now, a lot of them won't do this, but I think, uh, that to me would...

Or just go do the research, figure out what's actually possible. In our case, we were quite lucky in the sense that this is actually the foundation data.

Host 236:41

Data.

Pim de Wit36:42

Right? And I think- Right? Like, that's not true for, for every data set. I think, you know, we just happened to, to hit a particular gold mine.

Host 236:50

But you, you also did directly, probably you did the action thing, what, five years ago.

Pim de Wit36:55

Yeah.

Host 236:56

So, so you did work.

Pim de Wit36:57

Yeah, that's the thing. Like, you, you, you have to be grounded, right? And I think a lot of the, um... And, and I think that's the hard part, and I think a lot of what's interesting is you can also kind of look for if, like, scaling laws already exist on your data type, which, like, for video there were some, but for these, like input action labeled, uh, sets, there, there really wasn't any.

The other question is, like, does it go into LMs? Does it go into, uh, world models? Does it go into... Like, what type of model is it gonna be used for? And I think that's an important thing to know.

And so I just wanna... You know, if you're, if you're having these conversations with labs about data, just, like, make sure that you actually understand, like, what it's gonna be used for-

Host 237:34

Mm

Pim de Wit37:34

... 'cause that's a very, very good way for you to, like, make the decision yourself about what area you want to pursue that. Now, a lot of them won't tell you that.

Host 237:40

Mm.

Pim de Wit37:40

And I think, you know, I think in, in, in that case, you d- generally just don't want to do it because, like, I think, I think for our case, like we really cared that, like for instance, there weren't gonna be competing products with game developers built, right?

Because we didn't wanna h- like, bite the hand that feeds us, and I think we are a part of the games industry. So those questions I think are normal, and then we eventually decided, you know, you just have the data, we're just gonna go do it ourselves, and that's when the rest happened.

Yeah.

Host 238:03

And, and you assembled the team that can-

Pim de Wit38:05

Yeah

Host 238:05

... uh, take advantage of that. I, I feel like that's... You've aligned a lot of stars in order to make GI happen-

Pim de Wit38:12

Yeah

Host 238:12

... that other data founders, they are at the beginning of this journey.

Pim de Wit38:15

Yes.

Host 238:16

Or data founder. Founders who happen to have data. But they have a main business, right? Uh, I, I don't know if the other-

Pim de Wit38:23

There's two sides to this, right? There... It's really easy to be super naive about it, and like I had a lot of people tell me initially, "Oh, it's not that valuable. You're just, like, making this up." And, and, and so for me, like, doing the work and actually understanding it myself was a really, really big part of, of building that confidence to then go start a company.

Host 238:39

Mm-hmm.

Pim de Wit38:39

But a lot of times it is true that, like, model capabilities increase so quickly that, like, the-

Host 238:44

Mm

Pim de Wit38:44

... certain data you just don't need anymore.

Host 238:46

Yeah.

Pim de Wit38:46

Um, and so I think it is, it's really important to, like, get people to do the work such that you can make these types of distinctions.

Host 238:54

Yeah.

Pim de Wit38:54

And, and, and so, so my recommendation would be go build models with your data. See if you can create any sort of capabilities that, that aren't clearly already there-

Host 239:01

Yeah

Pim de Wit39:02

... um, or on path to being there, and then figure out, um, where you go.

Host 239:05

Yeah. I did wanna ask this earlier, but you gave me the opportunity to. Uh, when you say do the learning, do coursework and all that, and your co-founders gave you some homework.

Pim de Wit39:12

Yeah.

Host 239:13

Uh, is this like some books? I mean, Coursera?

Pim de Wit39:16

No, this was, um, Francois, uh, Fleuris, uh, Fleuris. So he has a little book of deep learning, and then he also has a full course that he's published-

Host 239:23

Okay

Pim de Wit39:23

... um, uh, on his website.

Host 239:25

I, I-

Pim de Wit39:25

I went through the entire course, uh, over the summer. I believe it's like something like 30 or 40 lectures, which also take home projects and things like that.

Host 239:33

Okay, amazing.

Pim de Wit39:33

Um, and I would recommend anybody, uh, uh, uh, does this. It, it goes through, right, history of deep learning, like the, the topology. It takes you through, um, the linear algebra, the, the calculus. You eventually end up with, like, chain rule, and by this time you've, you've done, like, all the, the, the more important concepts.

It takes you through how do you create neural networks using, uh, using these concepts that you've learned.

Host 239:56

Wow, this is super first principles.

Pim de Wit39:57

This guy, and I've, I've, I've had the, the f- the, uh, opportunity to spend some time with him as well. He is one of the most first principles people I've met in my entire life. I'm convinced... Like, I actually asked him, "Why'd you do this course?"

He's like, "Oh, uh, 'cause I thought all the other courses weren't right."

Host 240:10

Yeah.

Pim de Wit40:10

And because, because he's so first principles and he can only explain things from... Like, everything you see and how he explains this thing, it's everything is from first principles, including, like, the history of deep learning itself was part of, of the course.

And, um, yes, he goes, uh, um... So all- so he goes through everything and then, uh, and by the end of it, I think you're... Like, I now have, like, a pretty good intuitive understanding of how everything works, but obviously still, right?

Like, I, I like to describe it as, um, I'm like the, the guy who just got his driver's license. I can drive the car, and, like, my co-founders are like the F1 drivers that, like, have done this for years.

They know where all the, um, uh, where all the, the gaps are. And, and so I, I enjoy getting to learn from them. The cool thing is also that world models is just, like, a very, very new space.

Host 240:53

Yeah.

Pim de Wit40:53

And so, you know, I, I get to bring ideas to the table that, like, no one's thought of, and not because I'm great at this. It's because it's such a new space that, like, people just haven't tried it yet.

Host 241:01

Mm.

Pim de Wit41:01

Um, so.

Host 241:03

Let's get a hit on definition.

World Models 10141:03

Pim de Wit41:04

Yeah.

Host 241:05

What are world models to you?

Pim de Wit41:07

You know, in a video model, you might predict the next likely sequence or the next most entertaining e- the next most entertaining frame. Um, what world models do is they actually have to understand the, the full range of possibilities and outcomes, um, from the current state, and based on the action that you take, generates the next state, right?

So the ne- the next frame. And so it is a, it is a much more sort of complex problem than, than traditional video models. So to me, it is, it is a world that is accurately generated based on the actions that you take as a result of what's already been generated.

Host 241:42

And just a fact-check. Uh, that is, it needs to understand physics. It, it needs to understand if I'm building a type of material, it need... how it interacts with some type of material.

Pim de Wit41:52

Yeah. I think the interactions is the most important part. I think the reasons why world models are so fascinating, one of the things that I did when I, I was studying over the summer was I tried to actually build a super rudimentary, um, PyTorch-based physics engine, which I would not recommend writing a physics engine in PyTorch for obvious reasons, but I wanted to be able to, um, 'cause it's, it's differential, so you can, uh, you, you can generate the-

Host 242:13

It's very fatal, but-

Pim de Wit42:14

Yeah, exactly. You can... And then you can, um, uh, uh, train. And so I wanted to... You know, I got so many people ask me about, you know, why aren't you just using, uh, uh, why aren't you just simulating or generating this data?

Um, and I really wanted to understand from first principles why. And I think the most important thing that I figured out was the compute complexity of simulation goes up really, really rapidly with three variables. First, the numbers of agents in an environment Second, uh, their dof, so, uh, their individual-

Host 242:42

Degrees of freedom

Pim de Wit42:43

... yeah. And then third, the information that each action reveals. Um, so like, um, for instance, if you, if you have a te- if you have a, a text action or a speech action, the environment can change so much based on whether you say, right, water or fire, that the outcomes are going to be completely different of like how a human would behave in that type of situation.

And so it goes up so quickly with those three variables, that at some point you just hit a point where you just want to maximally bet on either video transfer or generation of these environments using world models, because that type of stochasticity is just incredibly difficult, but it's already very, very present in a lot of the video pre-training, uh, that goes into, into these world models, right?

And so I think for, for us, it, it is more so about making a maximal bet on video transfer and interacting with things that are difficult to simulate, and the steerability is also really interesting with text, uh, than it is on betting against simulation or something like that.

And so I think there's still a large market for, for traditional simulation engines, specifically in areas where video is really hard to get.

Host 243:44

Is this exactly what the big labs are also saying when they're talking to them?

Contrasting Approaches43:44

Pim de Wit43:48

I honestly haven't talked about the big ... to the big labs. Like, since, since we started working on them ourselves, I think people are more reserved with what they share with us. Yeah.

Host 243:56

Of course. It makes sense. It's a funny question.

Pim de Wit43:59

Yeah.

Host 243:59

How would you contrast your version of world models with Fei-Fei Li-

Pim de Wit44:03

Yeah

Host 244:03

... Yann LeCun?

Pim de Wit44:04

Yeah. So I don't know exactly what Yann LeCun is doing today. My understanding it's based on the Feijia Ba, like, LeJiaBa approach, which is... So I'll start with Fei-Fei Li. I think what's really interesting about Fei-Fei Li's approach is that you in some way are able to reuse the, the, um, the spots, right, in game engines and in things that let you stay in verifiable domain, um, which I think is a really interesting approach.

Um, however, my understanding is they're currently not interactive, which in my opinion is, like, the whole point of, of world models, right? It's, it's environments, they're great environments, and I think from a business perspective, I think they, they, they picked a really important part of the tool chain.

But to me, that's not really a, a, a world model. But I... My, my guess is they'll get there, right? They'll, they'll start generating-

Host 244:46

Yeah, they just haven't released it.

Pim de Wit44:47

Yeah, exactly. Exactly. And I think, right, Fei-Fei is one of the, like, founders of the entire space. Um, uh, so I think it's gonna be really interesting to me on, on, on what maybe that interactive piece looks like for me to really judge their approach.

I, I, I think-

Host 245:04

We interviewed, just before we moved to Yann, uh, we interviewed her with Justin Johnson, uh, her co-founder. He was, he was more focused on the physics side of things-

Pim de Wit45:14

Yeah

Host 245:14

... and the interactivity, uh, and each type of thing they still get. I, I, I do think that basically the, the splats, if you just add more dimensions on, I guess, the forces acting on them, then, then you get interactivity out of the box.

Pim de Wit45:28

Yeah.

Host 245:28

Because you are basically... These are virtual atoms that then has all the normal physics applied to them.

Pim de Wit45:34

Yeah. I'm, uh, I'm excited to see what that looks like when they actually release it. It's really hard to, really hard for me to comment on anything beforehand.

Host 245:43

Anything, yeah.

Pim de Wit45:43

I really like the, um, uh, the, the f- the frame-based approach, um, because all of our video or all of our training data is in this format.

Host 245:51

Yeah. Yes.

Pim de Wit45:51

Um-

Host 245:51

Yeah, so they were... We actually asked them about this, and they were like, "Yeah, it's possible, but we're choosing the splat approach."

Pim de Wit45:56

Yeah. Yeah, and you can also go from splat to frames, right? I'm sure you can write, like, at some... It's, it wouldn't be easy. Like, you'd have to actually render out the environment, do the... Sure. It, it's not, it's not gonna be a simple problem, but, like, in theory, it has to be something that you can do if you really wanted to.

So like, 'cause it's almost like having a more sort of ground truth through three-dimensional representation of the underlying world.

Host 246:15

Yeah.

Pim de Wit46:15

Right? So I think it's an interesting approach. Um, it might be overkill, right? Uh, uh, you're also dealing with, like, a much larger, like, degrees of freedom on the output space, right? So, so who knows how well it scales.

I like the fact that, like, I think these video models also use things like autoencoders, right? You can actually have the world models predict, like, much smaller, um, uh, maybe like a-

Host 246:37

Resolution or size.

Pim de Wit46:39

Yeah, exactly. And then you can use, like, diffusion upscaling or, or methods like this to actually, um, uh, enrich. And so I think that world models just allow a much more... or world models in my sense, for a much more, like, controlled space that, that, that we know really well.

Host 246:53

Yeah.

Pim de Wit46:53

Um, I'm not suggesting their approach is wrong. I'm just, you know, like, this is I think what we really like about it.

Host 246:58

Good.

Pim de Wit46:58

Honestly, Yann's podcast that he did, I don't remember which one it was, but a long time ago where he, where he basically proclaimed LLMs to be a dead end, um, uh, was one of the things that inspired me to do this.

Host 247:11

I think this is very consensus among world models people. Basically, everyone who has this insight stops with the LLMs and just goes straight to world models. I would say that the main pushback, I asked this exact question to Noam Brown from OpenAI, and he was like, "Well, they're learning physical models," right?

So there's basically the, the difference between implicit and explicit. Um, but what do you want to build on here, Rune?

Pim de Wit47:32

Yes. So, yeah, I, I, I'm not one to proclaim LLMs are a dead end, personally. I think, um, I think they're actually quite useful, in particularly as orchestrators. Like, the way I think about it is, as humans, right, we had sort of a three-dimensional world, then we invented text as like a, i- in a way, an compression method, right?

So you had... We invented text in order to communicate with each other in a, in a, in a common way, uh, in a, in a way that actually compresses all this information that we are perceiving in three-dimensionals space into just, like, a, a single sequence.

And I think that allowed sciences to emerge, right? It allowed so many literature, like, so many, uh, parts of the world that we, that we cherish. So I think it's a critical part of, uh, of the whole picture.

I also agree that, that, uh, it's very, very clear that they do build sort of the internal implicit world models, uh, inside LLMs. Um, and so I think they'll be very helpful as things like orchestrators. Um, the, the problem is when it comes to the generalization, I think, uh, text as a generalization backbone- When most of the, the, uh, um, when most of the, the pre-training is, is, is text, right?

Or, or, or largely te- uh, text sequences, then I think you want that backbone to be kind of more spatiotemporal in nature, and then also just have text, like, as one of the, as, as, as part of that.

And I think the actual argument of, um, of LLMs is also, for instance, the autoregressive nature of the prediction itself. So the, um, the fact that it's running the entire output, right, through the transformer, and then in order to predict the next token, which doesn't ...

Like, the environment in the real world is continuous, right? It's always, it's always changing, and LLMs kind of just forget about that, right? I think a lot of the, the, the argument is in, first, right, is so I think the, the fact, the fact that, like, text doesn't necessarily generalize well to spatiotemporal, um, context, and then the autoregressive nature of the prediction and using text for that, right?

So I think those are, those are the two main arguments. Um, I think, I think text prediction is just one of the actions that is gonna come out of, of these, you know-

Host 249:33

Yeah

Pim de Wit49:33

... these, these policies and world models. I think speech and text generation will just be-

Host 249:38

Mm

Pim de Wit49:38

... one of the actions that, that can, that can be a part of that. I think that there will just be labs coming at this problem from both sides. Um, and everyone ends up in roughly the same place, and the same place will be whatever people think is cool, uh, right?

Like, whatever the consumer use case is.

Host 249:56

Whatever is closest to AGI.

Pim de Wit49:57

Yeah, yeah. And so I don't think there's, like, a clear answer. I think it's really interesting to come, to come at it from the world modeling side, but it's also because we have to, right? 'Cause, like, text has largely commoditized.

We can import all the text.

Host 250:11

I- I think it's, um, interesting attempting that. Well, I'm attempting it makes sense that you can probably recover. It's sort of like you're, you're taking a step back. You're studying your branch of the ML research field, but you might actually just end up recovering all the other tech stuff emerging.

Pim de Wit50:26

Yeah, yeah. We can import a lot of that research, right? It's, it's-

Host 250:28

Yeah

Pim de Wit50:28

... a lot of that is, um-

Host 250:30

That's really cool on the, on the research side. Let's talk about the stuff that GI is producing more, like the, the, I guess, the sort of research and product output. You mentioned the word customers. What are your target customers?

Go-to-Market50:30

Pim de Wit50:43

Yeah. So we're, we're already working with some of the largest game developers in the world.

Host 250:46

Yeah.

Pim de Wit50:46

Uh, we're also working with game engines directly. And so really what we're doing at the moment is replacing essentially the player controller inside of an, a game engine. So anything that you're currently, that maybe, like, behavior trees or things that you're deterministically coding, we hope to replace with a single API, which is just you stream us frames and we predict actions, and that can be inside an engine or it can be, um, ev- eventually even inside the real world.

Hopefully those are then also steerable. So the models that you saw weren't text steerable yet, but I think we want to get to a point where they're fully text steerable.

Host 251:19

Well, does is steerable means, like, well, I want you to go to share-

Pim de Wit51:23

Mm-hmm

Host 251:23

... figure anything else out in the pre-

Pim de Wit51:25

Yeah, I think it, it's, it's text conditioning on the generation. So yeah, the ability to, to ... You're right. We want to get to a point where you can generally ... And that's why it's called general intuition, where we can sort of can mimic the intuition of all these gamers into human-like behaviors in any situation.

Um, as I mentioned also, the lab is named after the-

Host 251:44

Demis Hassabis quote

Pim de Wit51:45

... Demis Hassabis quote from AlphaFold, which is wouldn't it be amazing if we could mimic the intuition of these gamers, who are, by the way, only amateur biologists, um, on his path to, um, he tried to get an AI to train Foldit to generate a lot of data for, for AlphaFold.

And so for us, really, the, the North Star, right, what we hope to get to one day is being able to represent scientific problems in three-dimensional space and then have a spatiotemporal agent capable of perceiving that space and using hopefully also the, the, right, the, the text reasoning capabilities that LLMs have today in addition to the spatiotemporal capabilities to be able to work on the other side of that problem.

So that for us is, is sort of the North Star. That's why, you know, we're, we're sort of trying to be hyper-focused spatiotemporal workloads the same way that Anthropic was hyper-focused code, and use that to then get into organizations and expand from there.

Host 252:30

Yeah. Just as a side note since you mentioned Anthropic, um, any idea what they did on this, uh, to, to solve folding? Yeah.

Pim de Wit52:37

No. Out of any lab, I probably know Anthropic the least to be honest.

Host 252:41

Yeah.

Pim de Wit52:41

Yeah. I admire them though.

Host 252:43

Yeah. Well, the, the, the current working theory is that they had a super lucky, um, roll of the dice. But well, I ... And then, and then it compounds from there.

Pim de Wit52:53

That sounds like a nice story. I'm sure it's not that.

Host 252:55

Yeah. Okay. So, um, wh- why do the game developers want this?

Pim de Wit53:00

So if you're a game developer, how well you're actually retaining players is, like, um, if you have a game that's already at scale, is, like, decently dependent on how good your bots are. So if you're logging in at an obscure time, let's say 3:00 AM in America, and your player liquidity is low, then you need really, really good bots to keep those players engaged.

Host 253:19

Is this known? Is this a thing?

Pim de Wit53:20

Yeah. For sure.

Host 253:22

Especially for, like, Fortnite and whatever.

Pim de Wit53:23

A lot of gamers it is.

Host 253:24

Yeah.

Pim de Wit53:25

Yeah. Um, and so if, if you're-

Host 253:27

Even like as a human, do I want to play against bots?

Pim de Wit53:29

Usually it's not just bots. It's, like, players mixed in with bots because you don't want to play just against bots, but it's better to have a full game than to have, like, an empty game.

Host 253:36

Yeah.

Pim de Wit53:36

Um, and so I think as long as it's part of the environment, I think it's okay.

Host 253:40

That means you also have to sort of grade the skill level.

Pim de Wit53:43

Yeah, yeah. Which we can do, um, 'cause we have ... We know exactly how good people are at these games. Yeah. Yeah, I, I think for us, um, bots is kind of like step one, uh, right? So what I, what I was showing you is we're building a general agent that can sort of play any game, um, in real time, but really that extends into all of simulation, right?

Like in GTA V, for instance, people are genuinely role-playing real life.

Host 254:03

Mm-hmm.

Pim de Wit54:03

Right? And so they're actually behaving in quite aligned ways with, with the goals they set for themselves. So you have all these examples represented in video games, right? You have Truck Simulator, PowerWash Simulator.

Host 254:14

PowerWash Simulator?

Pim de Wit54:15

There's PowerWash Simulator where, like, actually the behaviors that you'd want, uh, an agent to be able to perceive, they're all there.

Host 254:21

Okay.

Pim de Wit54:21

Yeah.

Host 254:22

Most definitely. Yeah. It's really, it's really funny how seriously some gamers take Truck Simulator. Um, if you haven't seen these clips, you should watch it.

Pim de Wit54:30

Yeah.

Host 254:30

They buy the whole, like, truck driving set And they're doing the job-

Pim de Wit54:35

Exactly

Host 254:35

... of a truck driver.

Pim de Wit54:36

Yeah. W- what I mentioned to you, we have more people at any given time on Medal playing with steering wheels in, like, Truck Simulator and these types of games than Waymo has cars on the road. Um-

Host 254:44

Yeah

Pim de Wit54:45

... it's a ridiculous stat, but it's true.

Host 254:46

Yeah, yeah. I mean, it's, it... So, you know, I, I, I used to think that in order to self-driving, you kind of just need to play a lot of GTA V. Um, y- yeah, I mean, it's not, it's not bad, but it...

Yeah.

Pim de Wit54:57

Yeah. Our, our bet is not that we can zero shot any of these things. It's just that, like, the next self-driving company can maybe have- collect 1% of the data.

Host 255:03

Mm.

Pim de Wit55:04

Because, right, also, for instance, clips already self-select into negative events and, and adversity, right? And so, like, a lot of our dataset gives already highlights. Is, is really, um, precisely what a lot of these companies spend, like, their last 20% doing.

Host 255:17

Yeah.

Pim de Wit55:18

Right? And I think that's the main argument if you're, if you're, if, if you're another company that's looking at what we're doing. I think the thing that people are not... that people won't understand is that anything that you- that you're currently doing in, in pre-training, as long as your robot can be controlled using a game controller, we hope that we can move that to post-training for you.

Host 255:34

Mm.

Pim de Wit55:34

So our bet is not that we can create the next self-driving car company. It's just that the next self-driving car company hopefully only needs 1% of the data or maybe 10% of the data, I don't know, right, to be able to deliver a really good product.

Host 255:44

Yeah, yeah. It's also the, the term that comes to mind a lot is active learning. I don't know if you've, uh, used to identify with that. It's start... It got less cool for a bit, and now it- it seems like they're on the uptrend, uh, a bit.

Which, which obviously you have the best dataset for the sort of high-intensity or... You said negative-

Pim de Wit56:01

Yeah

Host 256:02

... but I feel like it's not negative. We- it could be negative or positive.

Pim de Wit56:04

Yeah, for sure. I think negative events is just because it's the most common term that people use for, like, if your, if your Tesla, you want the crashes, you want, like-

Host 256:12

Right. Um, yeah. Right, right, right. But, but so in gaming-

Pim de Wit56:14

It's both. Yeah, yeah. So, you know, the, the model that you saw obviously had really, really incredible moments, and, and that was largely-

Host 256:20

It was stupid human beings

Pim de Wit56:21

... because of the fact that it... Yeah, yeah. That, um, uh, that it had the large representation of people at their best.

Host 256:26

Yes. Yeah.

Pim de Wit56:26

And worst. Yeah.

Host 256:27

Yeah, yeah. Amazing. Okay, cool. Uh, anything else on the customer development side that you wanna sort of flesh out?

Pim de Wit56:33

Yeah. Um, uh, we're also already working with robotics companies, but again, the... and manufacturing, but the key is that the robot has to have gaming inputs. So we're... Like, our bet is not that we can transfer over to, like, higher-doff robots than the keyboard and mouse is.

It's really just that we can move the hard work of, of, of pre-training hopefully to post-training.

Host 256:52

Yeah. It, it's, like, kind of like the foundation model that is a very good basis to start.

Pim de Wit56:55

Yeah. You're gonna straight- you're gonna give us frames and, and likely some text.

Host 256:59

Or you license the model to ... Because they're gonna want to post-train them.

Pim de Wit57:02

Yeah. Our, our business model is initially gonna be an API.

Host 257:05

Yeah.

Pim de Wit57:05

Again, like the Anthropic API. Um, but you also saw, for instance, some of the video labeling models that we've been able to develop.

Host 257:11

Yeah.

Pim de Wit57:11

So, um, the goal is for any company to be able to take in their, uh, their video data as well, and we can create, first, obviously, custom versions of the policy for you, the agent. Um, if that doesn't work, then, um, we, we've already working with a customer that, that is doing...

We distill a model and, and they, uh, turn that into a product for themselves.

Host 257:30

So people can engage with you on the agent level, the API level.

Pim de Wit57:33

Mm-hmm.

Host 257:33

People can engage with you on the sort of model level. Can you also buy data or?

Pim de Wit57:38

No.

Host 257:38

No. Right. Yeah.

Pim de Wit57:39

We don't sell data.

Host 257:40

Okay, cool. So that's the, that's the business.

Pim de Wit57:42

Yeah.

Host 257:42

Um, and is there a world in which... I, I mean, I, I, I think this is on your landing page. If you are f- you know, Frontier Labs for, for world models, is there a world in which there is a more sort of application layer thing that you...

The Road Ahead57:42

Host 257:56

that comes out, like a ChatGPT for whatever?

Pim de Wit57:58

Yeah. You're gonna see us launch a few things on, on Medal itself that are gonna blow your mind-

Host 258:03

Mm

Pim de Wit58:03

... uh, as a result of this, this, um, this agent. I'll, I'll leave it to the imagination for now.

Host 258:08

If people to figure it out, you know, and-

Pim de Wit58:10

Yeah. On the world modeling side, like, I think one people underestimate is that Medal is already one of the largest, you know, video consumption platforms as well. People watch millions and millions of videos a day. Um-

Host 258:21

Whizzle.

Pim de Wit58:21

So, um, world model-based entertainment and things like that, while it's not, like, a focus for us right now, I think we'll be... Like, on the consumer side, we have the ability to move very, very quickly here, um, and, and get it integrated in a way that I don't, I don't think anyone else can.

Host 258:34

Yeah. You could theoretically do video gen, like the Sora, like, uh, what is, what is that Instagram one? What's, what's the Meta one? MetaMUSE? Not, not Muse. Um-

Pim de Wit58:45

Oh, Wise

Host 258:46

... Wise.

Pim de Wit58:46

Yeah.

Host 258:47

You could theoretically generate clips that nobody play, but you know it's gonna be viral.

Pim de Wit58:52

Yeah. I, I th- I think for us, the games being so human-centric is, like, a really big part of what makes it special.

Host 258:58

Yeah.

Pim de Wit58:58

Like, I, I actually, I actually just don't think that would work. Like, one thing that we are really excited about though, I'll, I'll give you one sneak peek of what we're thinking about, is what if you could literally replay any of the clips that you have inside a world model or your friends can play them?

Like, I showed you a model that already took part of your clip as a context-

Host 259:13

That seems to replay internet worlds

Pim de Wit59:15

... but it's also how we go from imitation learning to RL, right? 'Cause, like, it's part of our research roadmap anyways to make every single, every single clip on Medal playable. Um, so, uh, yeah, who's, who is to say that that doesn't apply to just the actual clips that you take?

Host 259:28

Yeah, yeah.

Pim de Wit59:29

Yeah.

Host 259:29

Interesting. Can you say more about the RL potential?

Pim de Wit59:31

We describe Medal as, as the episodic memory of humanity and simulation. So when you take a clip, really the way to think about it is you get the highlight of what is maybe three hours of playtime, right? You maybe get, like, two to three minutes of the things that were the most out of distribution, right?

It is genuinely your episodic memory, um, of that playtime and simulation, the things that you most want to remember and share. We want to be able to load, uh, and this is the work that Anthony Hu is doing.

The reason why we build world models is every crash that you run into in Euro Truck Simulator or American Truck Simulator or a driving game, we want to be able, right... And again, these are ground truth labels, so we know precisely the actions that lead up to the negative events.

Um, they're also title labeled. When people upload it onto our platform, they say, "Oh, good, it's a crash," right? And so we can select all these events, and if we can put them inside a world model, we can go into, right, we can, um, uh, we can train reward models to then, uh, reward based on how you perform in clips that actually contain negative events, for example.

And so for us, it's, it's very much about, um, uh, right, we can, we can create this, this, this, like, LM moment on, of, you know, imitation learning, but actually making every single clip on the platform playable, um, at billions of clips scale is how we go from imitation learning to RL.

Host 21:00:43

Cool. Uh, we covered a lot of it. Uh, is there anything else that you wanna do before we sort of wrap up the, the, the long-term vision stuff?

Pim de Wit1:00:49

Yeah, yeah. I think, I think for us, um, this is a very, very ambitious long-term bet.

Host 21:00:56

Yeah.

Pim de Wit1:00:56

We need the best researchers in the world-

Host 21:00:58

Yeah

Pim de Wit1:00:58

... that, that, that, that want to work on this stuff. Um, it's really exciting not being extremely data constrained. Um, like we really get to-

Host 21:01:05

Yeah

Pim de Wit1:01:05

... like we get so many learnings every week that we didn't think were possible, and it makes it for, for a joy working here. Also, the other thing is because we have such a large data moat, we don't have to be as concerned as the LLM companies about publishing because we don't-

Host 21:01:17

No one's gonna be able to-

Pim de Wit1:01:18

Exactly. No one can replicate the models, right? And so for us, um, we really want to bring back like the original culture of, of open research, which is why we did the partnership with Quitai in France. Um-

Host 21:01:29

Can you say it again? I, I actually didn't, didn't-

Pim de Wit1:01:30

Yeah. We just did a, um, we just announced our partnership with Quitai in France, which is, which is an, an open science lab in Paris, one of the best re- research labs in the world. Um, Eric Schmidt, I, I believe, funded it in addition to some, some French people.

They are essentially acting as the partner that's currently doing a lot of open research on the data. We also want to partner with universities who, um... because, like, we do believe this is the frontier, but it's so data constrained that really no...

everyone has their hands tied behind their back right now, and so we want to help fix that. So for instance, um, we want to work with universities to build like negative event prediction models for maybe like trucks in India on all the truck data where all these crashes occur.

We have all these things that we know we can do that we just haven't had the time to do. Um, and so if, if you're listening to this and, and you're, uh, maybe an academic institution or something and you want access to some of this data in a research, um, in an educational research fashion, I think we're, we're quite open to doing that 'cause we want to educate people.

And, uh, yeah, and other than that, we just want to work with the best infrastructure and, uh, research engineers on the planet as we're going into scaling, you know, runs that have thousands, tens of thousands, eventually hundreds of thousands of GPUs.

Host 21:02:32

Yeah.

Pim de Wit1:02:32

Yeah.

Host 21:02:32

Amazing. Uh, I primed you this as like the closing question-

Pim de Wit1:02:35

Yeah

Host 21:02:36

... of like at the... it's a little bit the Vinod Khosla 2030 vision.

Pim de Wit1:02:39

I didn't know.

Host 21:02:40

Yeah. So what does GI become in five years?

Pim de Wit1:02:43

Yeah. In 2030, we want to be the gold standard, um, of intelligence. Uh, and any sequence, uh, long enough is fundamentally a spatial temporal, right? Which I think is, um... so by nailing sta- spatial temporal reasoning, you go after the root cause problem of intelligence itself.

What the world looks like is we want to have 8... So I sort of group, um, the sequences of AI in three stages, and I credit Andrej Karpathy for, for teaching this. Um, bits to bits, atoms to bits and bits to atoms, and then atoms to atoms.

In the atoms to atoms stage, I want, like I want GI models to be responsible for 80% of all the atoms to atoms interactions driven by AI models. Uh, uh, and, and, and the reason, the, the reason for that is because we were able to unblock intelligence so quickly in robotics, like intelligence is the bottleneck, that supply chains actually converged on gaming inputs as their, as their primary input methods.

And, and they converged on essentially simpler systems that let us do a lot more a lot quicker. So we are essentially the, the 80% market approach, and then you have lots of companies that have kind of like specialized maybe humanoid robot OS stacks that are, that are the other 20.

And then so I, I wanna be responsible for 80% of all the atoms to atoms interactions driven by, uh, by these models and be the gold center for intelligence and maybe 100X more in simulation, 'cause I think simulation will actually be the larger market initially.

So I think in simulation, um, because you have very little constraints, uh, also from a safety perspective, simulation is much easier. So I think, uh, a lot of the takeoff initially assists in simulation, so a lot of the simulation use cases like, uh, what I mentioned, scientific use cases I'm really, really excited about.

And so, um, yeah, 80% of atoms to atoms interactions, uh, coming downstream from these types of spatial temporal foundation models, and then 100X more in simulation.

Host 21:04:22

Yeah. Yeah. It reminds me a lot of the, uh, what Mark and Priscilla from the Chan Zuckerberg Institute are doing with virtual biology, 'cause you can do a lot pretty simulation than you can do-

Pim de Wit1:04:33

Yeah

Host 21:04:33

... or you can do it a lot faster, uh, with interest. Um, amazing. Thank you for inviting us to your office.

Pim de Wit1:04:38

Yeah.

Host 21:04:38

And thank you for sharing a little bit about your journey.

Pim de Wit1:04:39

Thank you. Yeah.