LALatent SpaceJan 23, 2026· 1:32:05

Captaining IMO Gold, Deep Think, On-Policy RL, Feeling the AGI in Singapore — Yi Tay

Yi Tay, who leads Google DeepMind's Reasoning and AGI team in Singapore, explains how Gemini Deep Think achieved IMO Gold by abandoning symbolic AlphaProof for an end-to-end RL-trained model. He details the on-policy RL philosophy—models learn from their own generated outputs rather than imitating others—and the critical role of self-consistency through parallel sampling and internal verification. Tay describes the IMO effort: four co-captains in different time zones, a one-week training sprint, a live competition in Australia where researchers punched in problems as they were released, and the tension of waiting for human scores to determine the gold threshold. He discusses why the team believes one model must subsume everything for AGI, the data efficiency gap compared to humans, and his hiring focus on raw talent and research taste. Tay also shares his personal fitness transformation—losing 23 kilos and improving HRV—as integral to research productivity.

  1. 0:00Return to GDM
  2. 4:49On-Policy RL
  3. 12:40IMO Gold
  4. 26:37Pokemon
  5. 32:24Reasoning Defined
  6. 37:00AI Coding
  7. 45:09Architecture & Scaling
  8. 55:34Data Efficiency
  9. 1:07:50DSI & Rexus
  10. 1:18:23GDM Singapore
  11. 1:29:06Health & Wrap

Powered by PodHood

Transcript

Return to GDM0:00

Yi Tay0:00

The thing that I find the most useful about, like, these models in general is, like, when I have these big spreadsheets of a lot of results and I just need plots of it, I think models can quite good at just screenshotting a plot of this.

I hate making these method live stuff about. It's so annoying. There were so many moments this year where AI s- suddenly crossed that, like that emergent thing. So I think AI coding is one of them that we just discussed.

I think, like, Nano Banana also got to the point where I usually, like, we make these images, it's just like doing it for fun and just tr- troll your friend or something like that. But, like, Nano Banana actually really got so good.

Host0:37

Welcome back.

Yi Tay0:38

Yeah.

Host0:38

How are you?

Yi Tay0:39

Yeah. I'm good. I'm good. Great to be back.

Host0:41

It's been one, one and a half years.

Yi Tay0:43

Yeah, it's been one and a half.

Host0:43

Feels like a long time. So last time we talked, you were at Reka.

Yi Tay0:47

Yeah.

Host0:47

And then you joined GDM again, working for Kwok again.

Yi Tay0:51

Yeah.

Host0:52

And more recently you've started GDM Singapore.

Yi Tay0:55

Yeah.

Host0:56

Is it GDM Singapore or Gemini Singapore? I don't know if you've named the team.

Yi Tay0:59

Oh, I, I, I think we have a Gemini team in Singapore, yeah.

Host1:01

Gemini team in Singapore. It's called Reasoning and AGI.

Yi Tay1:04

Yeah. Reasoning and AGI.

Host1:05

Is it important to have AGI in the name?

Yi Tay1:08

I think it was, like, a vibes thing that we put AGI in. Yeah, I think that, like, one reason why we work on these models is that we want to get to AGI and I just, it was a vibes thing that we added the AGI to the job posting, yeah.

There is no, like, formal name of the team yet.

Host1:23

Team yet.

Yi Tay1:23

But it's basically the Gemini team Singapore, yeah.

Host1:25

I mean, I think people are, like, trying to triangulate. Amazon has an AGI team, you guys have an AGI team, and then let's say Meta now has a super intelligence team. What are people signaling when they choose these names for their teams?

Do they have, "Oh, we have a plan," or is it just vibes?

Yi Tay1:42

Are you trying to fish some hot takes from me? No. Uh-

Host1:45

You have officially AGI in your job title.

Yi Tay1:47

No, it's not a team name. It's, it's not a team name we have. Yeah, it's just, you know, we just want to signal the north star of we are building these models to get to AGI.

Host1:54

Yeah. Yeah, no, I, I wasn't really fishing for hot takes. Okay, so you rejoined GDM.

Yi Tay2:00

Yeah.

Host2:00

And I think last time when you talked about, I listened back to the whole thing, it was an amazing episode last time. You were talking about how it's like externally you were in Brain and came out and now you're back in GDM.

Yi Tay2:10

Yeah.

Host2:11

I wonder what's your just, your general reflections just plugging back into the Google infrastructure.

Yi Tay2:15

Oh, yeah. So I guess coming back is very interesting because it felt, and return to Google, like, everything, including your LDAP, your username is all the same. It's like you play Pokemon, you, you leave it aside and then you go back and you click continue game.

Host2:29

You save game.

Yi Tay2:29

Yeah, you save game and then continue game. It's like that. Obviously the last one point five years while I was away, many things have changed. Brain is now part of GDM and stuff, so I think that obviously a lot of things have changed but I think overall the coming back has been pretty seamless.

Obviously, I, I love Google infrastructure and I think debuts are great and stuff like that. Yeah, and I'm very glad to be back to-

Host2:51

Yeah

Yi Tay2:51

... to Google Infra, yeah.

Host2:53

And was the intention always that you were going to work on Deep Think?

Yi Tay2:57

No, not really. I think I missed research a lot, like, doing, like, research. Not like super fundamental research but, like, close to model research, right? But I really miss being at the frontier and trying to s- go be- beyond that, right?

So I really miss that a lot, and I think when I came back, Deep Think wasn't a thing and I don't think there was any plans actually. It was just like, I'm just going to work on research and see what happens.

Yeah.

Host3:20

I'm sure, I guess, there was some inclination that reasoning is the next frontier and that's, like, obviously the most rewarding research path, especially this year.

Yi Tay3:30

Yeah, I think reasoning, these days reasoning and RL is, like, what would be quite-- is RL reasoning comes up. I spent a lot of my past life, I call it the past arc, working on, like, architectures and pre-training but I think now I more, I would, like, transition more into RL research.

I'm not like old school RL, but the games RL and the old school RL and, and to be honest, I had almost no RL background coming back. But I think, like, like, RL is the main means of modeling these days and, yeah, so I think it was pretty easy to jump back in.

And I think a lot of fundamental skills in research is for general purpose and universal and, and it's, it's quite easy to innovate even in a tool set that you're not super used to. And yeah, so I think RL is basically the main modeling tool set that we play around with-

Host4:15

Yeah

Yi Tay4:15

... these days.

Host4:16

Superficially, I see some, you know, in your UL2 and Fan-- uh, T5 work, um, so some overlap of like, you know, the, the focus on objectives and the focus on the stuff that you're trying to incentivize. So I would have maybe guessed there was more overlap than you are saying right now, which is interesting.

But I know-- I understand it's very super-

Yi Tay4:38

Technically-

Host4:38

Superficial

Yi Tay4:38

... shape is objective and they have some, like, overlap, right? Yeah. I think it's just mainly, like, on policy and off policy-ness of designing these things that change how, like, also the learning algorithm itself, right?

On-Policy RL4:49

Host4:49

Let's just introduce this kind of terminology to people if they're not that familiar with the sort of RL policy. I do think that a lot of people are, like, trying to understand what is working about this generation of RL research.

Anyway, so Jason had this interesting post which I think you were co-signing, which is basically you always want to be on policy. Instead of mimicking other people's successful trajectories, take your own actions and learn from the reward given by the environment.

Basically correct your own path instead of trying to imitate other people's path.

Yi Tay5:17

Yeah, yeah.

Host5:18

And first of all, he writes really well, and I wish that more people wrote like him. But I don't know, what's your reflection on that or your addition on top of that?

Yi Tay5:26

Yeah, so I think, like, the biggest analogy of on policy and off policy is basically that on policy is basically like when you SFT something, it's on policy. Basically, you take some other model, larger model stuff, and then it's basically like this off policy.

Somebody else's generated outputs, trajectories, and whatever. I think on policy is mainly like the core idea of like modern LM RL, where you, like, generate and then you reward the model based on its own generations, and then the model trains on its own generations.

Host5:51

Yeah.

Yi Tay5:52

So it's more-- It's a bit like- Self dis- distillation to some extent, you, the model generates its own output, and then you reward it, and then trains on its own output. So I think on-policy nurse is basically this idea of, like, model training on its own outputs and letting the model, like, generate its own trajectories, and then let- letting some reward verify it, and then the model train its own outputs.

I think this is more generalizable in general. I think there's still a lot of, like, science out and still to be done about the gap between SFT and, and RL itself, but I think basically on-policy and off-policy, right?

And I think bringing this analogy back to real, like, life, and I mean, we-- this on-policy nurse is more like, like, humans, we are more on-policy because we go around the world, we make mistakes, and then we, ah, okay, this is...

But, like, imitation learning is mostly somebody else-

Host6:35

Not first principle. They just copy

Yi Tay6:37

... they just tells you what to do and then you just copy.

Host6:39

Yeah.

Yi Tay6:39

So I think, yeah, this philosophy, bringing back this philosophy to life is quite, like, powerful. Like, when I-- like, n- now I have a kid and everything, like, want my kid to try stuff and then you tell them like, "Okay, this is, like, where this went wrong, where this went right," and stuff, rather than, "Okay, you just copy everything somebody else does."

Host6:54

Yeah. There's a Montessori schooling is mostly that, right? Like, very unstructured learning. Like, you discover your own path, and we just give you a safe environment to do it.

Yi Tay7:02

Yeah, yeah, yeah, yeah.

Host7:03

What is the point at which you should transition from imitation to on-policy? I do bounce back and forth.

Yi Tay7:09

For humans, right, not models, right?

Host7:10

I, I would say in models it seems like there mostly has been a very concrete, like, first you imitate, and that's pre-training, and then you RL at the end.

Yi Tay7:18

I mean, technically SFT is still imitation, but I think for humans as well, a little bit of this, right? Because if you basically, like, it's like sports, right? When you play sports, you start off by imitating, like, hardcore imitating, but then you cannot imitate forever because you need to, like, imitation-- I, I don't know whether this is a good analogy, but watching a lot of tutorials and stuff is more like imitating.

We learn, try to learn certain movements and stuff like that. But then, like, on-policy nurse is like going to the game itself and trying to get a reward signal from that, right? But so I think that humans do need some form of imitation learning, but, like, I think everybody starts off by imitating.

But then again, the human and model kind of is not-- it, it's just fun to have analogies, but we shouldn't, like, take things, like, super literally and stuff like that.

Host7:58

I actually am a quite a serious taker of machine learning insights into human learning.

Yi Tay8:06

That's how we learn from our models now?

Host8:07

Yeah.

Yi Tay8:08

Hmm.

Host8:08

Because I think, like, machine learning is the most scientific way we have ever studied learning, just in general.

Yi Tay8:15

And that's true. That's true.

Host8:16

Where we had to invent curriculum from, like, scratch.

Yi Tay8:18

Yeah, that's true.

Host8:19

And things like learning rate. If your learning rate's too high, blah, blah, blah. Learning rate's too low, blah, blah, blah.

Yi Tay8:23

Like, wait, do humans even have a learning rate?

Host8:25

So I do tell people to, to keep an idea of their own learning rate and to be wary of it being too low. So for example, if you've been wrong once, you should ask, "Where else have I been wrong?"

And typically, usually, l- let's say learning-

Yi Tay8:41

Oh, okay

Host8:42

... you know what I mean? People usually update slower than they should when they have been wrong.

Yi Tay8:47

Is it stubbornness?

Host8:48

It could be stubbornness. Um, I don't know, is that, is that the right word for it? It could be, like, they are, they are too Bayesian when actually it sh- like, their prior assumptions are wrong, and they need to completely throw out their previous assumptions because one counterexample invalidates all prior experience.

Your entire world model is wrong. Throw it away. So Bayesian are actually wrong. Let's say you've lived for 10 years under some assumptions, and you have one example that breaks your narrative.

Yi Tay9:16

Okay.

Host9:16

You shouldn't be like, "Okay, now I have 2% update." No, actually, you should be like, "Oh," like, "something's really freaking changed. Everything I've assumed for the last 10 years is probably wrong. What else am I wrong in?" And update 20%, update 50%, not 2%.

You know what I mean? That's your-- that's a learning rate thing for me. So my, my direct example is the whole getting into AI stuff. I was watching GANs for 10 years.

Yi Tay9:40

Yeah.

Host9:41

Has it been 10 years? Uh, 2012, 2013.

Yi Tay9:43

Time flies, yeah.

Host9:45

I was watching GANs and I was like, "Okay, this is cool. It's getting more detail. Not that impressive." Then all of a sudden, stable diffusion came out, and you can run it on your laptop, and that was my learning rate.

Okay, like, fuck, like, my mental model of generative image- images did not include this. And so I was like, "Okay," like, "I am very wrong and I need to pivot everything," and that's how I started latent space.

Yi Tay10:06

So would this mean that, like, your, your learning rate is high?

Host10:08

Yes. I, yeah, I will nudge it up. I, I schedule my learning phase- ... because a world model has been violated.

Yi Tay10:15

Okay. I think it's a good, it's a good strategy. I think also this brings a little bit to, like, when new paradigms happen, like how fast people are to adopt it or, like, to invalidate their understanding of things.

I think as scientists, we def- definitely, a lot of times we do have to keep, as the field progress, we do have to keep, like, invalidating our own world model. It could be, like, in certain ways, the way to do s- like, something all along and suddenly something comes along and invalidates it.

Host10:37

Yeah.

Yi Tay10:38

Yeah.

Host10:38

Yeah. You can be very proud of your priors until it's, like, becomes your prison.

Yi Tay10:41

Yeah. No, yeah, yeah.

Host10:43

That is, like, actually very dangerous.

Yi Tay10:44

Yes. Yes.

Host10:44

Yeah. Okay. That was a bit of a tangent. I don't know how we got there. You did highlight Danny's LLM reasoning lectures where he kind of traced the intellectual history of reasoning in LLMs.

Yi Tay10:53

Yeah, yeah.

Host10:53

From chain-of-thought to, to now RLFT. And then the one part that I was gonna prompt you a little bit was also self-consistency, right? Which-

Yi Tay10:59

Yeah

Host10:59

... I think people roughly know. I think it's more crudely implemented with OpenAI than with you guys, where it is straight up they have eight inferences and-

Yi Tay11:09

Yeah

Host11:09

... they judge or whatever. But I do think that also is relevant to on-policy distillation, where it's like, literally you have eight different paths and y- they're all from the same model. Is it-

Yi Tay11:20

Yeah

Host11:20

... so checking my intuition there. Basically, the stuff that you were saying about why on-policy is important and using, let's say, an external verifier to, to improve your reasoning, you can also do that with parallel reasoning.

Yi Tay11:33

Oh, yeah, yeah. I mean, like, when we train our models, they sample multiple times. So yeah, to some extent there's some form of-

Host11:39

Yeah

Yi Tay11:39

... self-consistency.

Host11:40

Is, is that directly-- That's self-consistency, right?

Yi Tay11:42

Yeah. Self-consistency is a little bit more-

Host11:43

Budging

Yi Tay11:43

... it's a more nuanced version of-- If you talk to Danny, he'll tell you it's not majority voting for sure, but it's, it's more an-

Host11:48

I agree. I agree.

Yi Tay11:49

Yeah. It's a more nuanced version of that, but I think parallel thinking definitely is related to self-consistency. Yeah.

Host11:55

Yeah. I think for those people, OpenAI also actually put out some interesting papers on majority voting versus other forms of, like- Multiple output consensus. Then basically like you, like the highest level is actual, an actual LLM judge that decides like, you know, this is actually a worthwhile trajectory that is more valid based on some internal consistency or just like inspecting the chain of thought.

Which is, again, very cool that we can train models to do that.

Yi Tay12:19

Yeah, for sure.

Host12:20

Yeah.

Yi Tay12:20

Yeah. I-- Self-consistency is a big, like, a big fundamental idea in-- I mean, chain of thought itself was also a big idea, and then self-consistency was also like a big fundamental idea in, in, in, uh, in modern, like, LLM-

Host12:32

Mm.

Yi Tay12:32

-uh, literature.

Host12:33

Yeah. Amazing. Okay. So let's bring it to, I guess, the-- one of the headlines of this podcast is going to be about diving into the IMO work.

IMO Gold12:40

Yi Tay12:40

Yeah.

Host12:41

So this was around about May, March? March? In around July?

Yi Tay12:45

July.

Host12:46

You guys announced-- Oh, this very nice photo here. This is the photo I was looking at. This is in London, I believe, where you had Dishoom.

Yi Tay12:52

Yeah, Dishoom. Yeah. Oh, you've got to be at a, like you've got to be at a photo taking to, to get the credit.

Host12:59

That's bullshit, right? What the fuck?

Yi Tay13:01

No, no, no, no. No, I'm just kidding.

Host13:02

The contributor list is bigger than this.

Yi Tay13:03

Yeah, yeah. But, like, they were like saying that, "Oh, okay, you should go to the photo."

Host13:06

To-- In order to get a literal gold medal?

Yi Tay13:08

No, no, no, no.

Host13:09

Oh, okay.

Yi Tay13:09

Like, like to get the, the credit for being in the IMO effort is to be in the photo. So it's, it's just a joke. It's just a joke. Yeah.

Host13:15

But anyway, okay, could you tell the story of starting this IMO thing? Apparently, it was done in one week.

Yi Tay13:20

So let me like be a bit more clarified a lot, a lot of things, right? So the IMO effort has been like very long standing. So Tang and basically and Kwok has been really working on this even last year, right?

Last year they got-- So I was not back at Google at that time. So like they had the AlphaGeometry stuff and then there were like AlphaProves and stuff. So it's, it's a very long standing effort. But I think this year was the-- we wanted to try to like use-- actually use Gemini as a end-to-end model.

Basically, no-

Host13:45

No second system with AlphaProof.

Yi Tay13:48

No second system. Like text in, text out-

Host13:48

Yes

Yi Tay13:49

... like a model. And-

Host13:50

Even that was a not intuitive thing. I covered the silver result from last year.

Yi Tay13:55

Yeah.

Host13:55

And I was like, okay, it's pretty close. Like it's one point off from the silver.

Yi Tay13:58

Yeah.

Host13:58

Just try a bit harder, you'll get gold. The decision to abandon it, I think, was pretty bold. I don't know.

Yi Tay14:04

I personally was believe, always believe in if we are not-- Like in retrospect, it's easy to say this, but it's a bit like if the model can't get to IMO gold, then can we get to AGI? Basically, so it's basically at some point we have to use these models to, to try these Olympic com-competitions.

And I think that one of the goals this year was like, okay, we're going to do a end-to-end, like text in, text out model. That's where like my involvement came in. So basically, I was not like involved in the IMO effort only until the model training part.

So I have to say that, that Tang did, did most of the IMO thing. I just trained the model with a bunch of other-

Host14:41

Okay. What, what does that work involve? What are some things?

Yi Tay14:43

So, so basically we, we just prepared the model checkpoint for the actual IMO itself, right? So there, there's also something that's easily overlooked about the IMO thing was that many times you, you want to chase benchmarks or stuff like that.

It's always like a thing that you can kind of keep running and running and hill climbing until you get there and then you-- But like the IMO was a live competition. Like some members of the team were in Australia for the thing, and there was like this happening thing was happening live, was unfolding live.

Host15:08

Oh, it's a very AlphaGo. You receive the thing, you like punch it into your system, and then you like-

Yi Tay15:12

Yeah, yeah, yeah. So, so like some of the professors from Tang's team were like in-- when they went to the IMO itself and stuff like that, the conference, I don't even know whether IMO is a conference, but it feels like there were people there.

Like in Australia. And then, so it was a live thing, and there were people who actually the job was to run inference on, on, on this IMO P one to P six that came out, and they also came out on like different days.

So it's like different sets, like one day one, day two, something like that. So the fun part is that I knew nothing about IMO like at all. I'm not like a kid that took part in IMO. Like I stood down for that.

Host15:48

You're a piano player.

Yi Tay15:48

And yeah, I play the piano. But what I only knew was that, okay, we delivered the checkpoint, and that checkpoint was used to do the IMO Gold. But then like there was somehow a week in London where everybody gathered there.

So everybody was flying to London, and then this photo was taken there. And then you get to see how all the different parts like come together and like also being in the other rooms, in the rooms with the other like co-captains and then it felt a little bit like a hackathon thing.

So yeah, I think this was like the training process of this IMO model itself was like maybe a week or so. Not the actual, like the whole like basically like-

Host16:26

Yeah.

Yi Tay16:27

Yeah.

Host16:27

I think the, the question is I'm still not over the decision to throw away AlphaProof.

Yi Tay16:33

Okay. Yeah.

Host16:33

Basically. I think it's very major. A- a- and I understand that you have this goal of AGI, obviously, like at some point one model should do it, to do all of it, right?

Yi Tay16:41

Yeah.

Host16:42

But I think if you pointed a gun at me and said in twenty twenty-four, "What do you need is to do IMO and IOI and, and CPC," and all the other stuff that you guys did, was you need an LLM reasoning system that knows how to operate a computer and knows how to write Lean and run-

Yi Tay16:59

Yeah

Host16:59

... Lean verifier and all this. But basically you RLFT the Lean verifier into the chain of thought. Is that the-

Yi Tay17:06

Wait, I-

Host17:07

Like, so basically, like it is not obvious that you can do that-

Yi Tay17:10

Okay

Host17:11

... at all. Because I think-

Yi Tay17:13

So, okay, so I suppose what you mean is that like some, in some way-

Host17:16

You need a system one system two-

Yi Tay17:18

... this is encoded in parameters of the model somehow.

Host17:20

Yes.

Yi Tay17:21

Hmm. Yeah. I mean, it's just whether at the end of the day you, you just believe in like this like connection is one, one model, lots of parameters. I mean, there's also two use, right? There's also two use which, you know, and stuff like that.

But I think to some extent the model we-- Like I think we should be able to get to a point where to-- Like in the past when the LLM first started, the model could not even be a calculator.

Now it can somewhat be a calculator. So technically, like a tool like a calculator is somewhat encoded in the parameters of the model. So I think we, we, we'll eventually get a point where whether there's things that cannot be expressed in the parameters of the model is like an open question.

We, we don't know where is the limit, but I think we will keep pushing and pushing this limit. So whether like a Some- something like a lean system or like some other things to solve other, like maybe a physics engine or something where could it still-- we still continue to push that, that, that boundary, yeah.

But I actually don't know, like, whether there were a lot of debates about s-symbolic system versus end-to-

Host18:22

Yes, that's the word I was trying to-

Yi Tay18:24

I actually don't really know whether there was... Like, to me, I was just like, "Oh, let's train the model." And then, then some-so-someone told me to train the model, and then I trained the model. Basically, there was, like, overarching, like I-- people at, like, the IMO effort that, that decided this.

And I also think that because the-- basically these specialized systems are very, like, one-off systems that are like you could create, like, a chemistry engine, you could create a math engine, you could create the thing, right? But at the end of the day, you want one model for everything.

So yeah, so I think this kind of fits that direction a little bit more where you have one model for... And then this model was also, like, launched as Gemini Deep Think, as a general purpose Gemini Deep Think.

So-

Host19:00

It's basically unchanged but with maybe some config toned down a bit.

Yi Tay19:03

Yeah, so the, the, the inference time config was, like, the one served to most people is different, but-- and the full IMO co- like inference config was prese- get, like, shipped to some mathe-mathematicians just because of the inference cost, right?

But that was good enough to be a general purpose model. I think my take is that this i-i-- intuition was what led to the trying to go towards one model instead of... 'Cause this specialized systems, there's no end, right?

You can create many specialized systems.

Host19:31

Yes.

Yi Tay19:31

The, the most I can see in the future is there'll be a model then c- that, that is-- if there's something that really cannot be s-subsumed by a model, then you just use a tool or something like that, right?

Host19:40

Yeah.

Yi Tay19:40

But my prediction is that, like, I think most things can be, can be subsumed by the model. I think, yeah. I mean, AI researchers are quite good at hill climbing.

Host19:50

History would say that you're-- you have a lot of evidence backing you up. Is this the model output? This is it, right? This is what-

Yi Tay19:56

Yeah, I think this is the model output, yeah.

Host19:58

A-and what do you see when you look at this? You just see obviously it looks like a well-written prompt. It looks like something a real human mathematician would do. I-- people did compare yours versus the OpenAI one, where OpenAI is a lot more raw or had to clean up their versions.

We don't have to talk about O-OpenAI, but I just-- I think what is interesting to you when you saw this kind of output?

Yi Tay20:17

I want to give a little bit, like, a special disclaimer is that I know nothing about math. Right. So I think the wonderful thing about this era of LM is that, like, you can be a, like, a AI researcher engineer and you don't have any domain knowledge, and you can still, yeah, solve a-- get a gold medal in a-

Host20:34

It's a universal tool

Yi Tay20:35

... that you don't know anything about it. I can't parse this at all.

Host20:38

Okay.

Yi Tay20:38

Like, this is foreign to me.

Host20:39

But, like, maybe a proof is a particular kind of chain of thought.

Yi Tay20:43

Yeah.

Host20:43

But I would say that the other interesting thing that some of the-- some of your co-collaborators were talking about, he was, he was like, "Oh, this is the first example of reasoning in a non-verifiable domain." Which to me, isn't proofs by definition verifiable?

I just want to give you things to riff on or debates that are-- might be worth digging into.

Yi Tay21:03

So I think there's g- there's a lot of... aside from proofs, there's a lot of domains that are, like, non-verifiable, and I think not, not easy to verify. So it's like when people mean non-verifiable, it's like non-trivial to verify or, like, just not as, not as easy as, like, the solution of, like, a math problem because proofs are long-form and it's also-- that's why it's non-trivial to, to verify unless you convert it to Lin and then you do all the kinds of things, right?

So I think there's a lot of work to be done in, like, these type non-verifiable domains, yeah. I'm going into this territory where I'm not sure whether what I can say or what I cannot say.

Host21:33

Okay. So we, yeah. Sure. No, say it. I think another thing that is an open topic of debate was how much domain specific work or post-training was done because you then went on to do the IOI and CPC stuff as well, right?

The same model.

Yi Tay21:48

I was not directly involved in the ICPC, but I was related to some extent. That's all I can say. Yeah.

Host21:56

Yeah. Any other interesting call-outs maybe just on-

Yi Tay22:00

I won't-

Host22:00

Just on the team. You called out Jonathan as someone who's co-captain on, on this effort, and yeah, basically how does the effort like this come together?

Yi Tay22:08

So I think there were four captains for the IMO. Two from London, Jonathan was from Mountain View, I was from Singapore. So I think four of us basically trained this model together, and I think o-o-one-- I was also trying to receive what Tang was saying.

But I think one, one interesting thing was that, like, we're all in different time zones and we're all f-- and there's something also very interesting about, like, passing on the job. There's no, like, really, like, fixed workflow how to work together between captains.

So it's more like, "Oh, I'm going to board the plane now. I'll be like AFK for 12 hours, so someone take-"

Host22:37

So just babysitting the run.

Yi Tay22:38

Sometimes there are bugs and stuff, the job comes down sometimes. So basically, it's very ad hoc and it's very, very-- it's really between the captains how we decide to, like, work together. And yeah, but I think it was a kind of interesting time also because we were all flying.

Like, I think the London folks were not having to fly, but I had to fly and Jonathan had to fly. And then, like, when you visit another country, you have another-- like, if you visit an office, you have many meetings.

So I was in and out of meetings, and it was a pretty, like, interesting... And I also think that we-- nobody really knew whether we would get gold at that time because IMO actually has-- had- hasn't happened. Yeah, it, it was interesting and exciting.

And then I think, like, the whole process of this getting verified by the IMO committee and, you know, like, you know, there was like, "Okay, but we're not going there." But I had to learn a lot about how the IMO works, right?

Apparently, the gold score is not even, is not even a fixed number. It's like a bell-

Host23:29

Yeah

Yi Tay23:29

... a bell curve, right?

Host23:30

Yeah.

Yi Tay23:30

So it was like a time where you just look at the score, you're like... Like, I was even, like, looking at the-- watching the human partic-participants and then seeing, like, what their scores were because-

Host23:39

Mm

Yi Tay23:40

... whether Gemini will get gold depends on, like, how the humans do. So you, like, looking at, oh, if a certain percentage, you're like, "What's the-- Do we get gold?"

Host23:47

Yeah. To some extent, you don't have any control over that, so it's-

Yi Tay23:50

Yeah, but you're just curious, right? Because-

Host23:52

Yeah.

Yi Tay23:52

So I would say that it, it is definitely more int-- like exciting. Like, there's more adrenaline than, like, just r-running on a benchmark and getting a number. It was like a process that, that, that took some time. Yeah, yeah.

But I think overall, if you have specific questions, you can ask also, but I think this whole thing has been a highlight for me. This IMO effort has been like-

Host24:11

It-- Yeah, I, I would say, uh, most people, if you ask them maybe two years ago whether a model could get an IMO Gold, they would have said it like impossible

Yi Tay24:20

Yeah.

Host24:20

Then the, the silver helped, right, from last year. But, like, the fact that you can throw that system completely away and then just take ex- existing Gemini and scale up Deep Think and then just run it for IMO Gold, I think is also, like, very non-consensus compared to last year.

Yi Tay24:35

Yeah, definitely to some extent I think researchers were also surprised. I wo- I won't say, like, surprised, but, like, it was more like a pat on, on the back kind of surprise how we actually made a lot of progress in-

Host24:46

Mm.

Yi Tay24:46

We as in collectively, the, the, all the engineers and researchers working on Gemini. There's a lot of progress being made. Just look at how much we went in one year. Yeah. And I also think that-

Host24:54

Yeah

Yi Tay24:54

... it just five years ago, like, the not two years, like five year, you just, you just imagine the outcome. Like, you just look at the state of AI now, like just generally, the, uh, IMO and the ICBC Gold and, like, also even, like, things like NanoBanana.

If you just look at the AI progress now and five years ago, I think people would think that we already reached, like, AGI.

Host25:10

Some form of AGI.

Yi Tay25:11

To some form of AGI.

Host25:12

Yeah.

Yi Tay25:12

We're just moving ... If you just traveled, like you take these checkpoints and you traveled back into five years ago, uh, someone should make a drama about this. But I think it's really quite impressive, like, how the field has moved so quickly.

Host25:23

Yeah.

Yi Tay25:23

Yeah, yeah, yeah.

Host25:24

The hard parts you would say were scaling inference.

Yi Tay25:27

In what aspect? Like hard in terms of like expensive?

Host25:31

Hard as in maybe the most amount of brain power expended on the team. I saw some comments where they were like, "Actually, the hardest part was the inference optimization," or like the very, very long horizon inference that Deep Think needed compared to a normal Gemini, stuff like that.

Yi Tay25:47

I didn't work on the inference time scalings. I-

Host25:49

Yeah.

Yi Tay25:49

Yeah, I wouldn't know. Yeah.

Host25:50

That is mostly that. Oh, and then there was this, the codename was apparently IMO Cat, which you named after your desk.

Yi Tay25:58

Okay. That's not really... Like, it was in the... So I- I think I tweeted about it at some point, right?

Host26:04

Yeah.

Yi Tay26:04

Like bringing up the tweet. So the IMO Cat was basically, like, okay, it's not like a official codename or something. It's just like the name that the config of the job was, like, IMO Cat. That's, that's, that's the-

Host26:16

You just need some kind of name.

Yi Tay26:17

Y- yeah, I mean, I just like, you know, I just, I like cats and yeah.

Host26:20

Yeah, fair enough. That is mostly it on IMO unless you want to bring up any- anything else. We have other sort of research-y topics, but beyond, before I go into sort of research-y topics, I did wanna maybe leave the floor to cover what else should people know about the reasoning effort that's going on at GDM?

Yi Tay26:37

Let, let me think of where to start-

Pokemon26:37

Host26:38

Yeah

Yi Tay26:39

... with this. What do people need to know?

Host26:41

So-

Yi Tay26:41

You know the model's really good, yeah, but that's what people need-

Host26:43

Maybe an easy one to start with would be a lot of people were focusing on maybe academic benchmarks two years ago, last year maybe LM Arena, this year Pokemon. Pokemon's a very interesting reasoning, visual reasoning and just general long horizon agent planning-

Yi Tay26:59

Yeah

Host26:59

... benchmark, and I, I don't know, you seem to focus on it a lot, and I think Gemini did very well, so obviously I think it's something that is easy to talk about.

Yi Tay27:08

I think I probably should rephrase this. There's actually nothing specifically done for Pokemon.

Host27:12

Of course.

Yi Tay27:13

Yeah, of course. Like, there's nothing specifically done for Pokemon, and I think that, I think Logan had this tweet recently about the recent Gemini 3 on Pokemon Crystal.

Host27:22

Mm.

Yi Tay27:22

And Pokemon Crystal is like-

Host27:23

There's so much more to punch

Yi Tay27:24

... Pokemon special. Yeah. I think Pokemon is like... So I used to play a lot of Pokemon, and I'm a big Pokemon fan in general, and I think it's a, it's a great, like you said, it's a great long ho- horizon benchmark and stuff like that.

And I think it's good to check in once in a while on these benchmarks that, like, almost never get contaminated or let people actually, like, don't spend time to-

Host27:45

Benchmarks

Yi Tay27:45

... like, Hugh Clive benchmarks. It's, like, kind of silly to, like, people are like, "Okay," like, "what are you working on?" If some people, "Oh, I'm working on Amy, I'm working on HLE, I'm working on Pokemon maxing," or something like that.

That's kind of, like, funny.

Host27:56

We did interview the Cloud-based Pokemon, uh-

Yi Tay27:58

Yeah

Host27:59

... AI. I think his name is David. And it showed serious flaws in Anthropic's screen understanding vision capabilities.

Yi Tay28:08

Yeah.

Host28:08

Couldn't, literally couldn't tell I'm trying to, like, get past this wall, but you keep, just keep running into it-

Yi Tay28:14

Oh

Host28:14

... 'cause it doesn't know that the wall is there. And so it doesn't have any spatial reasoning at all.

Yi Tay28:19

I mean, some of it could be like a harness, like the harness thing or, like, also whether the model has access to, like, game state information or is it complete visual.

Host28:28

Yeah, Cloud's implementation is very game state heavy. They dumped effectively, like, all the, the memory of what's going on in the emulator.

Yi Tay28:36

Yeah, yeah, I see.

Host28:37

Yeah.

Yi Tay28:37

I think for... But I don't know whether I'm jumping off tangent or something like that.

Host28:40

No, it's fine.

Yi Tay28:41

I, I think solving Pokemon is gonna be more of, like, how fast you solve it and then, like, the thing that I have not really seen so far is, like, whether the model can com- complete the Pokedex.

Host28:49

Why is that?

Yi Tay28:50

I think you need-

Host28:51

Even more challenging.

Yi Tay28:51

No, completing the Pokedex is so hard, right? You need to plan, you need to, like, you need to search up, like, informa- like, there's some things if you don't, like, go online, basically you need to, like, have a little bit of deep research in-

Host29:01

Yeah, yeah

Yi Tay29:01

... in this.

Host29:01

Yeah.

Yi Tay29:02

There's, the model will just never know, like, that it needs to trade, okay? If it's able to go online, post on forums, and then find someone to like, "Hey, Claude, can I trade with you? This Pokemon to evolve."

Some Pokemon needs to be traded to evolve.

Host29:13

Yes, or mated. I think it's-

Yi Tay29:15

Yeah, but-

Host29:16

Yeah

Yi Tay29:16

... yeah, but anyway, I, I don't know. I have not seen the model be able to complete a Pokedex. But completing the Pokedex is really hard actually for models. So I was, I think that's actually an interesting, int- in, like, like, an interesting one.

Yeah.

Host29:28

Yeah. I wonder what the real world analogy would be once, let's, or if, let's say we, we have a model that is capable of doing that, what can we make it do that we cannot do today?

Yi Tay29:38

Well, there's a lot of planning involved, like-

Host29:39

Just real deep research.

Yi Tay29:40

Like planning, there's also a lot of planning involved and it's more, I think completing Pokemon is p- p- but the Pokemon game, the l- is very linear, right? The, the Pokedex is, involve a lot of, like-

Host29:50

Backtracking research.

Yi Tay29:51

Yeah, a lot of research and a lot of, yeah. So it's a probably a different nature-

Host29:55

Yeah

Yi Tay29:56

... itself.

Host29:56

Is that as interesting to you as, for example, a lot of other people in the AI for science world are trying to discover things that you cannot look up, right?

Yi Tay30:04

Novel knowledge.

Host30:05

Yeah, novel knowledge. Because basically what you're saying is we're not even there yet. We're at the place where models cannot consistently apply knowledge that they look up, right? Like you give Gemini access to a web search and you say, "Okay, go and try to collect all the Pokemon in a Pokedex."

You don't have high confidence that it'll do it. I, I don't know if someone's actually tried. Probably not, right?

Yi Tay30:26

No, I, I think the hard part is actually like- Trying to use the s- like synthesize the web knowledge and then apply it in, in, in the game itself with all that visual state going on and stuff like that.

It probably will be solved in one or two-

Host30:37

Yeah, yeah

Yi Tay30:37

... no, it's not-

Host30:38

It is challenging

Yi Tay30:38

... it's not challenging, but not like super-

Host30:40

But like is it interesting? Like you're basically just... Y- the task really is can you look up the guide to do it, and then can you apply the guide? That's it. You know what's even more s- intelligent than that?

Creating the guide. Like being the first to figure out how to create the guide. Which is what in-

Yi Tay30:58

Oh, yeah, yeah. But then when it comes to this one, it's mostly there's a exhaustive search thing. The model will try and try like humans guide for that. Yeah, yeah.

Host31:04

Okay, so that's actually less interesting to you. Interesting.

Yi Tay31:06

Okay, actually to be-- okay, when you think about it, it's like not super, super interesting, but it's okay. It's just that I have not seen the model to try to do this, so it's-

Host31:12

Yeah

Yi Tay31:12

... just like-

Host31:12

Uh, yeah, I think like efficient search of novel idea space is interesting. Obviously, you can brute force anything, but we're not talking about brute forcing. We're talking about trying to create an AI scientist.

Yi Tay31:23

But novel knowledge is actually an interesting thing that, that I think is gonna be quite a big thing, being able to generate-

Host31:29

Yeah

Yi Tay31:29

... novel know-

Host31:30

Google has done stuff there, which I don't-- uh, you're probably not that close to those teams that has done AI scientist work.

Yi Tay31:36

Just some things that have been... It's like, for example, if you freeze the model weights at like 2015, you freeze time at 2015.

Host31:44

Mm.

Yi Tay31:44

And then you, even with the current model, let's say you have... Okay, let's assume there's no l- l- leaking of information somehow. Can-- if you ask the model what's the best ML, okay, not 2015, like 2012 or something.

It will just tell you that SVMs are like the best, right? This is the way that machine learning works in general, right? Then the question is can you invent the transformer, right? It might not be able to... Like even today models, they b- might not even be able to invent the transformer.

Like if you freeze the time at a certain time, and even you bring the te- I mean, the model is a transformer, so I just say there's no-- assuming there's no leakage, there's like, like no-

Host32:14

Oh, yeah. Yeah, no, totally possible. Yeah.

Yi Tay32:15

Right. So, so I think there's still a lot of open questions about like whether the model can really, you know, innovate and generate like, ooh, really novel-

Host32:22

Yeah

Yi Tay32:22

... uh, uh, knowledge. Yeah.

Host32:24

Yeah. One related question on that, which I think is related to the, to the Denny paper, which is I think people have this sort of mythicism on what reasoning is.

Reasoning Defined32:24

Yi Tay32:32

Mm.

Host32:32

And reasoning, if you really demystify it a lot, it's whatever happens inside the chain of thought tags.

Yi Tay32:39

Mm.

Host32:39

Right? And you post-- you, you... You're eliciting that reasoning behavior from some stuff that is already latent inside of the pre-trained corpus. Is that-- that's one version of this interpretation.

Yi Tay32:53

I think these days, like reasoning itself is very vague and it's very open. So it's like most- mostly different people will have different definitions of what reasoning is, right? So I agree that like this chain of thought is like basically when people think of reasoning, they, they associate it with chain of thought, and obviously it's what happens in, in, in the thinking and the, okay, reasoning, right?

But I think these days it's more like, as said in the earlier part, it's like reasoning in RL is almost like the s- like basically it's anything that is like post-training to elicit capabilities, basically. It's like, like RL and post-training to elicit capabilit- capabilities.

So I think the actual like technical definition of reasoning is making models better with thinking and post-training. Okay? Yeah. So basically like RL-ing the model to think better, right? And thinking is more like thinking tracers and tr- thought trajectories and stuff like that, right?

There's also this line of work for la- like latent thinking and, and stuff like that. Like whether latent thinking and discrete token thinking is like going to be the same thing or like something like that. It's like a open question.

Host33:49

Meaning adding extra tokens to your vocab that represent or-

Yi Tay33:52

There's all this, I forgot what's the name, like these academic papers that do these loopy things or-

Host33:58

Track tokens

Yi Tay33:58

... or like they basically, instead of decoding discrete tokens, you actually simulate this by doing this in latent space, right?

Host34:05

Uh-huh.

Yi Tay34:05

So when you do chain of thought thinking and like re- reasoning, you basically decode extra tokens, hide it in the thinking pack, and then you do decode stuff. But like latent thinking is basically you, you just don't decode-

Host34:16

Yeah

Yi Tay34:16

... tokens. You just don't bother by it.

Host34:18

It might start speaking-- the native language of thinking is numbers, not passing it through some filter of English. And sometimes it might start thinking Chinese or something else.

Yi Tay34:27

Yeah, yeah. Generally, I'm not a-- I'm not really-- I don't really believe that model thoughts have to be the same with human thoughts. I'm actually like-

Host34:33

Yeah

Yi Tay34:33

... generally in ML, I'm more of the school of thought of let the model do whatever it wants in general.

Host34:38

There was a discussion. There, there's a latent representation hypothesis paper that I think you are maybe sympathetic to if you haven't already read it. To me, sounds obvious that basically image models will have the same idea what a laptop is versus a text model will have-- they converge on like the same latent-

Yi Tay34:53

Yeah

Host34:53

... and obviously you can align them and you can do all those stuff with them. And so it, it totally makes sense that their concept would be just a vector of numbers that represents laptop. That's the concept.

Yi Tay35:02

Yeah, yeah.

Host35:02

And okay, maybe you have some numerical differences between one model's idea of what a laptop is versus another, but it mostly would be the same. Yeah, very interesting. The question I was kinda leading into was that because there's now we're in this age where LLM text is in the corpus of like stuff that we train on, where it's a little bit of a recursive loop, right?

Like the reasoning tokens are out there now, and so pre-trained models themselves, pre- pre-trained based models-

Yi Tay35:32

Yeah

Host35:32

... are also capable of reasoning, and they're increasingly so as more and more reasoning text goes into the corpus. Isn't that interesting or is that worrying?

Yi Tay35:43

Do you actually see much reasoning trace on the internet?

Host35:46

So-

Yi Tay35:46

I've never seen those though

Host35:47

... I would say that-

Yi Tay35:49

Like on Hugging Face?

Host35:50

Yeah. People are publishing that specifically. As, as to whether or not people are-- researchers are actually including that in their training corpuses, who knows, right? Like this, that's their, their choice. But I would say that percentage on Common Crawl that has CoT tokens in there went from zero to 0.001% and it will just go up over time because like people are publishing it.

Yi Tay36:14

Yeah, but I think if the sources are like quite clear, you can actually like filter away those because usually people put it on GitHub or put it on-

Host36:19

Do you want to filter it? Maybe you don't.

Yi Tay36:21

There's, there's a choice for the, for the-

Host36:23

Yeah. Quite re- quite literally the whole reason why, I don't think we covered this in our previous part, but- Two years ago, a lot of people were like, "Oh, you just include more coding tokens in your pre-trained corpus.

It will be-"

Yi Tay36:35

Wait, but the coding tokens are different from like coding token-

Host36:38

No, it generalizes outside of code for reasoning.

Yi Tay36:41

Oh, that was like-

Host36:42

You don't believe that?

Yi Tay36:43

No, no, no, as in like, I don't know if it's still true today, but I see.

Host36:46

Yeah. That, that was just our general coverage of reasoning. I would say that there's a lot of interesting work here and more to do. Maybe I'll cover one thing which I know that you have personal inputs on, which is that you have started using AI coding.

Yi Tay37:00

Oh, yeah. So I actually don't really u-use much AI coding in the past, but I think we've reached a point where AI coding has started to become really useful. Like, okay, so before AI coding, the most-- the, the thing that I find the most useful about, like, these models in general is, like, when I have these big spreadsheets of a lot of results and I just need plots of it.

AI Coding37:00

Yi Tay37:20

I think models can quite really just screenshot making a plot of this. I hate making this method like stuff about it's so annoying. Okay, but that's basically, like, one thing that I can remember about, like, how I used AI in the past.

But I think AI coding has started to become the point where I run a job, I get a bug. I almost don't look at the bug. I paste it into, like, Antigravity and, like, I told you that would fix the bug for me and then I relaunch the job at that like...

Beyond, like, vibe coding, it's more like vibe-

Host37:47

Training

Yi Tay37:49

... V-Vibe ML or something like that.

Host37:50

ML.

Yi Tay37:51

I would say it, it does pretty well most of the time, and it's actually there's, there are classes of problems that it's just generally I know this is actually really good for and in fact, maybe probably better than, like, I would have to spend twenty minutes to find, like, figure out the issue and then-

Host38:06

Okay

Yi Tay38:06

... fix the thing and then relaunch.

Host38:07

To be-- So yeah, that's very interesting because I would say, like, level one vibe coding is you actually know what to do, you're just too lazy.

Yi Tay38:14

Yeah, it's just, "Ah, just do it for me." Like, I've done this a thousand times, like, "Just go fix it, like I know exactly what to do."

Host38:20

Here you're saying it's like on, like, the next level where you actually don't even know. It's, it's investigating it for you. As long as, like, the answer looks right, you'll just ship it.

Yi Tay38:28

At the start, I was a bit like I did check it, look at the thing, and then at some point I'm like, "Okay, maybe the model knows better than me," so I'm just going to let it do its stuff, and then I relaunch the job based on the fix that the model gave me.

And I think the models will just keep, keep getting better and better.

Host38:43

Yeah.

Yi Tay38:44

So yeah, it's something that, that-

Host38:45

Yeah.

Yi Tay38:45

But I also think that I mean recently there's this, there's Antigravity. I think, I think also because these tools were not like that. In Google infrastructure it's not that easy to-- You, you don't-- I'm not that familiar with what is available outside, and when I was at the start, I didn't really-- I think the models were not, like, so good like one and a half years ago.

So it's also like a forcing function that, like, it's also people are like, "Oh, try Antigravity, it's a game changer and stuff." Like, okay, so I just started using and yeah.

Host39:10

Yeah. You spent some time with Varun recently. What did, what did-

Yi Tay39:13

Oh, no, I really just said hi and greeted him. Okay.

Host39:15

I guess you were telling me you're an AI researcher that doesn't even have to use much AI, and like now you're actually like AI pilled as a user.

Yi Tay39:23

There were so many moments this year where AI suddenly crossed that, like, the immersion thing. So I think AI coding is one of them, like we just discussed. I think, like, Nano Banana also got to the point where I usually, like, we make these images, it's just like a little for fun and just thr-throw your friend or something like that.

But like Nano Banana actually really got so good that-

Host39:39

You can use it for charts.

Yi Tay39:41

That, that-

Host39:41

Yeah

Yi Tay39:41

... that you can use it for like basically, yeah. So it's getting really good and I think, yeah, th-this year the stuff like-- and even things like the past, many of these LMs will like halluc-hallucinate things a lot, but now I just trust it like automatically.

I think we just-- I think people are just enjoying the utility brought by-

Host39:58

Yeah

Yi Tay39:58

... by, by, by these models. So now I'm like-- I was always AI pilled. Like AI is a good thing, like I, I don't see how anybody can disagree with that, but-

Host40:06

Yeah, but you are actually using it for things that you are high expertise in, which is your own ML work.

Yi Tay40:11

Yeah, yeah, yeah.

Host40:12

And just to, just to confirm, you, you are-- you-- do you have a special version of Gemini that you use internally that we don't have access to or you're-- it's like public Gemini? Or do you have-

Yi Tay40:21

I, I think it's the public Gemini.

Host40:23

Okay.

Yi Tay40:23

Yeah.

Host40:23

The, the-- I was just saying, like, it would be entirely reasonable to train Gemini internal for only your code base and your work.

Yi Tay40:31

Oh, actually, I'm not sure though. You see, these things are just like really-

Host40:34

Yeah, yeah

Yi Tay40:34

... like strategy-

Host40:35

It's further away for you

Yi Tay40:35

... further away for me. Yeah.

Host40:36

Yeah. But obviously, if it obviously improves, improves your pro-productivity by, I don't know, ten percent-

Yi Tay40:40

Yeah

Host40:41

... worth it, right?

Yi Tay40:41

Yeah, yeah.

Host40:42

So I, I think that's interesting. A-and there's the interesting thing, le-levels of how much do you trust it, how much of your jobs do you automate away-

Yi Tay40:50

Yeah

Host40:50

... that you no longer need. There's also the question of, I guess, about how people come up and train in the field if they-- you no longer need juniors, because the Gemini is your junior ML, ML researcher. So I think these are, these are all interesting questions.

Yi Tay41:03

I want to say one quick thing first, right? So I think that when it comes to, like, whether a model can be like a junior SWE or like something like that, right? So I think if you think of it this way, of if a job from a one, 1x SWE, one time SWE can be replaced by a model itself.

But let's say you are a manager, right? The objective that-- the metric you track is, like, your time and then if you can have a model that saves you, like, the same amount of time as the work that your reports do, but you don't actually like replace one person per se, but you, yeah.

Host41:31

A little bit from everybody.

Yi Tay41:32

Yeah, right. Then you can-- I, I definitely agree that, okay, like when you count the net time saved, there are times where the model can fix bugs that, like, would have cost me like one day, right? One day is huge.

Host41:42

Right.

Yi Tay41:42

And these things are definitely like if you-- I don't know whether anybody has done any like real, like, metric evaluation on these type of things, but if you use time as a me-real metric and then not as a number of, okay, maybe three hours, like, kind of me-me-metric, right?

But these things are not like, like going to replace one person as, as it is, but more like a passive aura that buffs everybody. The in-game terms, right?

Host42:03

Yeah.

Yi Tay42:03

Yeah.

Host42:03

I often think of myself as a bard because I tell stories and I, I plus everybody around me, you know? Like that's, that's a, that's an ideal situation for me in a D&D group.

Yi Tay42:12

Oh, okay. Yeah, no, I don't play D&D, but okay. Yeah, I, I get it. I agree.

Host42:15

I said they're the kings of Metamora.

Yi Tay42:17

Support, support hero.

Host42:18

Support, support. Yes. Yeah, okay. AI support I think is very encouraging. I think, like, where is it still not working for you? That you've tried and you're like, "Oh man, I expected it to be better."

Yi Tay42:29

Oh, there are times when models get-- try to get lazy and try to fix something in a, like, they are still-

Host42:36

We will have-

Yi Tay42:37

They get lazy, and then they try to, like, gaslight me into thinking that, like, the bug is fixed.

Host42:41

Ah.

Yi Tay42:42

So there, there are still classes of-- there are classes of problems that are, like, very easy for the model and very hard for humans. There are some things that are very easy for humans, very hard for model, reverse paradox and stuff.

So it's still very hard to characterize these, these things into these proper quadrants and stuff like that. So I would say that the capabilities on models these days are good enough to be r-like, really helpful, but, like, it's still-- you-- it's a bit like it still has some...

But yeah, but I think this will-- like this-- I don't think there's anything that to, to be done to specifically, like, focus via these things. It's more like general capability improvements. The models just get better over time, and then these things will just, like, go away.

Host43:15

You say that, okay, so yes, I think obviously in the grand scheme of things, just trust the process, keep scaling in every dimension, and, uh, things will just fall away, things will emerge. But you've also said in the past, I can't remember the exact tweet, where you were like, "Each additional data set compounds over time."

They're just small additions. And I would say that when you s-say things like focus fire on things that you would think humans-- it's easy for humans, hard for machines, those are easy wins where you can just add a data set that would focus fire on that.

Yi Tay43:47

Mm.

Host43:47

And isn't hill climbing just a sequence of doing that until you reach AGI?

Yi Tay43:53

Okay, so I, I get your point. I think that it's true that sometimes a lot of progress on the whole is just a series of small incremental changes that-

Host44:01

Yeah

Yi Tay44:01

... that push. And that, that-- I think that's accurate. That's true. There's also-- it also feels that there's also a lot of, like, small mi-- like, seemingly m-m-minor, for the lack of a better word, like, that push AI to the state where it is today.

So I definitely agree. So nothing aga-- like, against people who are, like, focused fire, but it's just that when I mean that, like, it might be not easy to focus fire on things that are, like, not very easy to characterize.

So it's just that, like, when it has-- there's something targeted, right? You know, like, okay, I want to improve this capability, add some data. So I think, like, defining, like, the evaluations, defining the problems and stuff like that is, like, characterizing it.

And if it can be characterized, and then, okay, then fine. But I think it's, like, what I was trying to say with the coding is that these things are not even some of the class of problems are, like, I don't work on coding, but, like, people who maybe work on coding, they know, like, they have, like, terminology for different types of failures, but so maybe somewhere somebody is focusing fire on this and then make the model better.

That's great for everybody.

Host44:54

Yeah. I mean, that's why it takes 1,000 people to, like, get all these things together.

Yi Tay44:58

AI is definitely, like, a big collective-

Host45:00

Yeah

Yi Tay45:00

... effort these days. It's a big machine.

Host45:02

Yeah. It's, it's really crazy. Okay. So I just wanted to broad-broaden out to general things people are talking about in the community on research.

Architecture & Scaling45:09

Yi Tay45:09

Yeah.

Host45:09

Which again, I know that you are very locked in, so you don't, not necessarily have read all the papers or anything, but we can just riff on ideas.

Yi Tay45:18

Okay. Yeah, yeah.

Host45:18

You can obviously ask me what I think as well.

Yi Tay45:20

Yeah.

Host45:21

Is attention all you need?

Yi Tay45:22

So attention and transformer has been-- is, like, a core idea in the recent times. Like, pre-training and scale is the thing that made attention and transformers, like, actually shine, right? Because without-- I think the first transformer paper was this, like, a machine translation thing, and then basically, like, GPT and BERT were the ones that, like, actually showed the full, like, big potential of this idea.

So in terms of, is attention, like, really, really all we need? Like, uh, probably no. But I think it's, like, one of the-- from the architectural point of view, also may-maybe no. But it's not all we need. It's-- but we, we need it definitely.

I, I-

Host46:00

So what else are you thinking about on that same level? Are you talking about MOE stuff? What do you mean when you say it's not all you need? Like-

Yi Tay46:05

No, you definitely need, like, you definitely need, like, the scale of pre-training you need-

Host46:08

Okay

Yi Tay46:08

... or the tokens you need, like RL.

Host46:10

I think when I say, when people say-- I say it-it's attention all you need, it's, it's mostly-

Yi Tay46:14

From architectural point of view?

Host46:15

Well, will transformers get us all the way to AGI, right? I guess it's the-

Yi Tay46:18

Oh, so, so basically when you get to AGI, it's the problem-

Host46:21

Will it still be, will it still be a known architecture or, like, meaningfully different?

Yi Tay46:26

It will be a transformer, I think.

Host46:27

Really?

Yi Tay46:28

Like, like it will-- like, people-- it depends on what you call it. But I think unless the paradigm shifts completely, which is-- I mean, as a scientist, you cannot, like, like, completely say no to, like, like, that th-this would never happen.

But my feeling is that it's been, like, what? Like almost 10 years since transformer.

Host46:44

2017.

Yi Tay46:45

Nine.

Host46:45

2017.

Yi Tay46:46

Nine years, like, since, since the transformer. I think we have not r-replaced self-attention. Like, it's, it's some form of it. Like, you could rename it, you could name it something else, you could sometimes-

Host46:57

You can do local, global, style window.

Yi Tay46:58

Like, yeah, it's still a transformer in the end, and I think, like, that's not going, like, anywhere unless the whole thing with, like, backprop, like, everything, like, goes-- like, the whole thing just changes completely. Like, then there's a different story, there's a different conversation to have.

But if it's still within the same scope and bounds-- So I spent a lot of time thinking of-- about architectures and, like, whether there's alternate architectures and stuff like that.

Host47:20

Yeah.

Yi Tay47:21

Okay, at the sequence processing level, like there's-

Host47:24

There is the ultimate-

Yi Tay47:25

There is-

Host47:25

... sequence-to-sequence transformer

Yi Tay47:26

... it's probably the self-attention is there was this whole big era, which I was also involved in this era, where people tried to, like, undermine the attention as much as possible. Like, they tried to remove it, simplify it, make it efficient, like this whole, like, like, efficient attention era.

At the end of the day, the outcome was always like, "Oh, remove all attention, but we have one layer of self-attention there, and it still works." Like, that's at the end, like, of the s-- like, always the story.

Host47:50

Which even Noam at Character, he published some stuff about how he has some ratio of mixing of local and global attention, right? Like, basically still attention, but modifying it quite a lot for-

Yi Tay48:01

I, I would consider local and global attention to be, like, still a- ... still attention.

Host48:07

Just like how much you're skipping it.

Yi Tay48:08

Yeah, yeah. The only, the only question is that if the formulation changes too much, your QKV becomes like A, B, C, D, E, F, G or something like that or some-

Host48:15

Okay. Maybe I'll give you some motivating constraints-

Yi Tay48:17

Yeah

Host48:17

... in order to do this. You guys are still charging 2X for over 180K token context or 240, something like that. And the max is, theoretical max is two million tokens, right? What if we need 200 million? Is there some point at which where even this concept of Input token context is irrelevant because you are doing continual learning, uh, that, that kind of stuff, where you're modeling it as, okay, the AGI will be achieved through a sequence to sequence transformation, therefore an attention is the best sequence to sequence model or architecture, therefore attention is all you need.

But I think other people are like, sequence to sequence doesn't accurately capture intelligence.

Yi Tay48:59

Well, but that's not really about sequence to sequence. That's more about like the whole b- gradient design and backprop thing, right?

Host49:06

Mm-hmm.

Yi Tay49:07

It's not architecture itself. That's, that, that's a problem that we-- that is more of like the learning paradigm itself rather than the architecture itself. I think the architecture is just basically like the interface between the learning algorithm and the tokens.

I think it's more about the learning algorithm itself. And this cont- like continual learning this like there's many ways to think about processing many like insanely large like context, right? Like 200 million, 1 billion tokens or something like that, right?

Like whether it's gonna be like you have a new le- learning algorithm that every time you run inference, you, you learn on it, right? Then you can technically have some kind of memory, like human being is learning as I'm talking to you, right?

So that, that's also like one way. The other way is like whether, okay, maybe somebody will say that, "Okay, the attention is like just too expensive for 200 million, 1 billion context, so we need new architecture." Or some people will say that, "Oh, we just improve the chips like accelerators."

So I think many ways to interpret it, but I think it's like if it's about-- There, there's a lot of fundamental things that if it's about continual learning and stuff like this, a lot of fundamental things about the learning algorithm and, and stuff like that as well.

Yeah.

Host50:09

That will have to change?

Yi Tay50:11

I think like the learning paradigm and architecture and stuff like that goes like hand in hand. And I think as the field progress and ideas just stack on top of one another, right? So there's also this thing about a idea that was proposed has to be compatible with all the work that have been done before to, to shine, right?

It's a bit of variant of this hardware lottery like by, by Sarah. It's not the hardware lottery that I wrote about the TP- GPUs failing, but it's the, it's the original hardware lottery, but it's more like a bit of lottery of like the things proposed have to play well with the things that were proposed before.

So it's a bit like going down this local optima minima to some extent. So now we are like in this local minima of like transformers everything, everything, right? Maybe it's n- not easy to like get totally out of this because al- also a lot of people's investment-

Host50:53

Optimisation

Yi Tay50:53

... optimizer have been, have been done. So the things that play well need to play well with the ideas before. And the, the way I see it now is it's very difficult to like-

Host51:01

Yeah

Yi Tay51:02

... c-come out of it.

Host51:03

Okay, I'm not entirely convinced. I, I see what you're saying, but let's call it Gen AI. I fucking hate that term.

Yi Tay51:10

Yeah.

Host51:10

It's still a very young field, and so yes, there's been eight years of work on the transformer, but what's that in the grand scheme of things? Maybe we're in a local minima and we've got to nudge ourselves out of it.

I do wanna leave that open-ended. I don't have an idea. I do, I do think that people are in what's called-- what Ilya Sutskever has been calling the age of research, right?

Yi Tay51:29

Mm-hmm.

Host51:30

Like we're, we're like, okay, we scaled up what we can scale up. There's-- we know what the next maybe one, two orders of magnitude look like in scaling-

Yi Tay51:37

Mm-hmm

Host51:37

... on every dimension that we know about. But what is the next dimension to scale?

Yi Tay51:42

There's this like m- mis- misunderstanding a little bit about like, oh, the last five years has just been like scaling things, scales.

Host51:50

Okay. Please tell me more.

Yi Tay51:51

Right.

Host51:51

Yes.

Yi Tay51:52

I think-

Host51:52

You, you made that joke about sca-- now we scale res- uh, researchers' salaries.

Yi Tay51:55

Okay, let's not go there. But, but I think that ideas like matter, and I think that there have been a lot of good ideas in the last five years. It's just that maybe it's just not-- So it's not been like blindly-- Like if you took a MLP, right?

Just like without a self-attention, and you just, "Okay, I'm going to throw like $100 trillion on this and scale up that thing," and the thing-

Host52:15

It's never gonna work

Yi Tay52:16

... it's never gonna work.

Host52:17

Right.

Yi Tay52:17

Yeah, yeah. So there's no like-- There's part of it there's also like I think the bitter lesson gets used too much in like too conveniently used around. But actually, there's, there's also a little bit of, uh, not a bit, there's also a sweet lesson where it's like ideas matter.

And I think even to, to today, right, like people downplay ideas and stuff like that. Yeah.

Host52:36

Do you think the rate of new i-- without being specific about what ideas, 'cause obviously you can't share, but like do you think the rate of ideas has increased or decreased because there's like kind of a law of diminishing returns?

Are they smaller-

Yi Tay52:46

I think the number of ideas is always proportional to the number of researchers working on a certain problem. So by definition, by definition it should increase, but I think the number of ideas that actually work is not decreasing compared to the last-- Like it's, we are not in the era of diminishing returns yet.

Host53:00

Well-

Yi Tay53:01

So I think ideas are still very important and there's still very good ideas that are game changers that are being invented.

Host53:07

Yeah.

Yi Tay53:08

Yeah, yeah.

Host53:08

And I think I know the answer to this, but is the closed lab advantage increasing versus open source or decreasing?

Like the Chinese labs, they say they keep publishing open source models and some of the American labs as well publish open source models. Would you say that the ideas that I see there, NVIDIA has Nemotron, OpenAI has GPT-oSS.

These are all basically checkpoints on what is publicly known about training models as of this year. You know, the-

Yi Tay53:39

Oh, okay, okay

Host53:40

... they're like it's declassified information 'cause everyone-

Yi Tay53:42

Yeah. Like-

Host53:43

Okay, yeah, everyone does this.

Yi Tay53:44

I think that the gap is increasing.

Host53:47

I don't think it's, it's completely predictable from stuff that you said before.

Yi Tay53:49

I think the gap is definitely increasing. Yeah.

Host53:52

I think that would make-- That justifies researchers. Other- otherwise, what's the point of having researchers if not finding new tricks that compound over time? Yeah. Um-

Yi Tay54:02

Yeah, yeah. But, but definitely I think it's, it's, uh, it's increasing. Yeah.

Host54:06

Okay. I'll do a side tangent. I don't know if you have any comments on this.

Yi Tay54:09

Okay.

Host54:09

So then this is very related to NVIDIA's recent purchase of Groq, which I don't know if you have views because, you know, you're very TPU centric, but are we memory or compute bound? And this is re- re- relevant to the transformers discussion of like-

Yi Tay54:23

In terms of what, like serving?

Host54:25

E-e-exactly. I think the c- the classic view is that we are compute bound 'cause we just need more compute for pre-train and RL and then inference.

Yi Tay54:33

Yeah.

Host54:33

But actually the Counterargument I would make a- against this is I actually have these charts of Moore's laws. I wish I could just pull it up easily. Moore's laws of the scaling of compute versus scaling of memory versus scaling of network and bandwidth, and compute has a much higher slope for scaling than the other two.

Yi Tay54:53

Memory, what, the cheap memory?

Host54:55

Yeah.

Yi Tay54:55

Like, honestly, I don't think about this, this memory bound that much, so maybe it doesn't... Yeah, so I would disagree-

Host55:02

Yeah

Yi Tay55:02

... with it, but I don't have high confidence in, in-

Host55:04

And because you're mostly on the research side, less on the inference side.

Yi Tay55:07

Yeah. There's, yeah-

Host55:08

Maybe the inference guys would be like, "Oh."

Yi Tay55:09

Yeah, yeah, yeah. I don't think about... I don't wake up and think about serving, so yeah, maybe I don't think about the inference-

Host55:15

Yeah

Yi Tay55:15

... that much. Yeah.

Host55:16

My, my previous line of discussion here was like, NVIDIA is very fo- foresighted by Mellanox because it, it actually is the real bottleneck in scaling-

Yi Tay55:24

Mm-hmm

Host55:24

... uh, 'cause it has the lowest Moore's law, and then the second one now is memory-

Yi Tay55:28

Mm-hmm

Host55:29

... which is very interesting.

Yi Tay55:30

Okay, okay. But honestly, I don't think about, uh, doesn't think about this, this, that, that, that much. Yeah, yeah.

Host55:34

Understand. Okay. Data efficiency. So this is a joke, but implicit in this is that there's some kind of maximum data exposure, right? And, and so I-- so previously I would say that a lot of the training paradigms is like one epoch is all you need.

Data Efficiency55:34

Host55:49

That's the-

Yi Tay55:49

Yeah, yeah

Host55:50

... meme-y title of this idea. I would say that the real number maybe is between three to four epochs, and I do wonder what the theoretical limit of data efficiency of a model in terms of training and compression should be.

I, I don't know if that means-

Yi Tay56:08

Data efficiency and, basically by you asking the question in a way, asking like how much repeats is tolerable. Is that-

Host56:14

One is we've-- in, in tolerable is contingent on does it actually improve in meaningful ways.

Yi Tay56:20

Oh.

Host56:20

It's not about like you actually want to do it for its, for its own sake. But I do think there's that, and then there's also just the sheer amount of stuff that we can learn with limited data. So you say, as they say, like we are-- you are not compute bound, you're not memory bound, but let's say you are data bound, right?

Last time that we were on the podcast, we talked about Chinchilla versus inference optimal training. But now actually I think a lot of people are even talking about like data optimal training. Like a given limited data set, how well can you learn from it?

I think that's an interesting research direction that not enough people are talking about. Maybe it is something that is commonplace in the labs, but it seems very clear that we are very unoptimized with regards to how much we, we learn from our data.

I'll just put it there.

Yi Tay57:07

I think in general, the, like, learning m- more, like extracting more from every data point is definitely valuable, but I think that's also related to the fact that we're, like, running out of tokens in the world. So I don't work on data for pre-training, and I think things that I say would-

Host57:27

General state of industry, not nothing-

Yi Tay57:28

The general state of industry, right. So I think that the-- I don't even know whether data has, like, diverged, like the way that these things are done. It has diverged too much across these labs and no open-

Host57:40

There's a lot of cross, cross-pollination for sure.

Yi Tay57:42

Cross-pollination? Okay, okay. Yeah. But I don't think about data that much, like the pre-training data that much this-

Host57:47

Yeah. Maybe earlier this, the first half of this year-

Yi Tay57:50

Yeah

Host57:50

... I would have said that kind of pre-training is dead and that everyone's like just funneling all their work towards RL, and you-- we had this like Groq chart, which was very interesting, where we're sending the same amount of compute on-

Yi Tay58:02

It's like Groq

Host58:03

... as to-- You think it's a psyop?

Yi Tay58:04

No, I don't know why. I have no idea. Yeah.

Host58:07

I think that people are taking it seriously. They are like, yeah, okay, whatever. Es- especially in the agent labs like Cognition and Course, they're taking the open source models from, from whoever and then adding, let's call it pre-train scale RL on top of it-

Yi Tay58:23

Mm-hmm

Host58:23

... if they have that level of info which, uh, the data which they do.

Yi Tay58:26

Mm-hmm.

Host58:27

Which is very interesting. I would say, yeah, this data efficiency argument, yeah, I- I think it's, to me, is also more trying to discover new paradigms of learning in order to get where we wanna all go, which is, yeah.

Yi Tay58:40

Mm-hmm.

Host58:40

And the existence proof is humans, right? Your two-year-old daughter can-- is much more capable than, than an LLM in some things, having seen way, like eight orders of magnitude less data. That's very interesting.

Yi Tay58:55

Yeah, comparing human learning to machine learning is definitely like, like-

Host58:57

Purely as an exis- existence proof that we could probably do better. Three examples of dog.

Yi Tay59:02

Yeah.

Host59:03

Fourth example of unidentified animal, I can probably tell it's a dog as a human, but machines, classically you take 20.

Yi Tay59:09

The data efficiency of humans is definitely way higher than, than, than models. Yeah, the only question is that where does this thing come from? Is it actually like putting more flops on every token or like maybe it's like back to the question about whether the transformer is the optimal architecture.

Maybe it's backprop, it's the problem-

Host59:25

Yeah

Yi Tay59:25

... maybe it's the off-policiness, maybe it's the-

Host59:27

Yeah.

Yi Tay59:28

So what is the-- like, where is the bug, right? Where is the bug, right?

Host59:30

Exactly. Exactly.

Yi Tay59:31

Uh, but maybe it's a feature, not a bug. I don't know. But, uh-

Host59:34

So now, okay, we've identified probably, it took me a while to get this across. This is the kind of data efficiency I'm, I'm talking about. I think it's emerging. Basically, at the end of every year I like try to take bets as to, okay, what will be the big theme, so next year.

Yi Tay59:45

Yeah.

Host59:46

I think this is one of them that people are really trying to focus on. Be- because you're feeling this data crunch, even though everyone's like still investing in data. I forgot to mention that I would say that I've been wrong on pre-training being dead.

Yes, I've now met pre-training leads from both Anthropic and OpenAI, and I've also seen the, the talk from the DeepMind guy recently, and so everyone's investing in pre-training still, which is like nice to see.

Yi Tay1:00:11

Nobody said pre-training was dead.

Host1:00:12

I know. No, it's a, it's a theory that we're trying to disprove-

Yi Tay1:00:16

Okay, okay

Host1:00:16

... or, or, or prove.

Yi Tay1:00:18

Yeah, yeah.

Host1:00:18

Anyway, so I think, okay, let me wind back to my general idea, right? So yeah, data efficiency seems worthwhile. You would treat it as like, okay, well, show me where the bug is, and I'll go fix it.

Yi Tay1:00:29

Yeah.

Host1:00:29

We don't know where the bug is.

Yi Tay1:00:30

Yeah, yeah.

Host1:00:30

We just have existence proof that it could be better. And then I think the final logical chain in this for me is that everyone is focusing on some idea of world model. As a version of this form of more efficient learning, which potentially might not take the form of a sequence to sequence transformer.

I don't know how that works. Like, definitely a little bit out of my depth here. To me, that is more efficient because every world must be infinitely consistent, and if the next piece of evidence come in and invalidates those worlds, then you no longer need to pursue those paths e- ever, and you can just narrow in on the world that you've identified.

And so to me, that is learning, where you're learning to fit world models.

Yi Tay1:01:13

Or to the actual data.

Host1:01:14

Yes. Oh, so yes.

Yi Tay1:01:16

Mm.

Host1:01:16

Maybe you can treat the learning process as curve fitting.

Yi Tay1:01:19

Yeah. So you're learning the world instead of learning the... So the, the world model, yeah. So you're learning the world model, right? Okay. By sampling multiple world models and then finding out which one fits the data the best.

Host1:01:29

So I guess my query, is this what people talk about? Is this... If you-- I mean, obviously feel free to attack it because I'm just spitballing it. But this is what I pick up from talking with multiple people about, okay, what are you talking about with wor- world models?

What are you talking about data efficiency and learning efficiency, and, like, how do you gel it all together in a cohesive sense of the future where we can actually-

Yi Tay1:01:49

Can, can, can we def... What's the definition of wor- world model at the start, from the start?

Host1:01:52

Yeah. There are two kinds.

Yi Tay1:01:54

Okay, okay, go on. Yeah.

Host1:01:55

First kind is the VO kind.

Yi Tay1:01:56

VO kind.

Host1:01:57

Or the, what's the other one? Genie.

Yi Tay1:01:59

Yeah.

Host1:01:59

That, that DeepMind has, which is the sort of video world model.

Yi Tay1:02:01

Yeah.

Host1:02:01

You model everything with some kind of Gaussian splats or whatever, and you're like, you inhabit those, that 3D space.

Yi Tay1:02:07

Yeah.

Host1:02:08

Second, let's call it the Yann LeCun/Meta school of thought, which I don't know if you're that familiar with it. He has published the JePa architecture, and then separately, FAIR has also published the code world models, where you're basically, specifically for code, very interesting, you are executing code and modeling the internal state of the execution environment as you go line by line.

Yi Tay1:02:30

Oh, okay.

Host1:02:31

So the, the LM actually, like, learns to predict th- those things.

Yi Tay1:02:33

Yeah.

Host1:02:33

And actually it seems a lot more efficient at the scale that they've tested it out, which is very cool.

Yi Tay1:02:38

Which definition are, are you anchoring on?

Host1:02:40

The third one.

Yi Tay1:02:41

The code one, the code world model?

Host1:02:43

That's the second one.

Yi Tay1:02:43

Oh, second one, okay.

Host1:02:44

Those two are bundled together.

Yi Tay1:02:45

Yeah, yeah.

Host1:02:46

The JePa fans are probably hating me right now because I'm lumping all Meta's work under one school of thought-

Yi Tay1:02:51

Yeah

Host1:02:52

... but whatever.

Yi Tay1:02:52

Okay.

Host1:02:53

Okay. The s- the first one is VO, Genie, I feel it's like super spatial intelligence, those kind of video based world models. Second one is some execution or some sort of explicit modeling in the, uh, as you, as you sort of run through the, the corpus.

Yi Tay1:03:08

Yeah.

Host1:03:08

And then I think the, the third one is this amorphous thing which I think people are trying to get to, where they are doing what I said about the resolution of possible worlds and you're curve fitting as you learn, as you inference.

Yi Tay1:03:22

Yeah. But what is the world model itself, itself? Is it like-

Host1:03:25

It is a mental model of where everything is and how you think the world works. How, what I think you think, that, everything.

Yi Tay1:03:34

Okay. But the, technically it's like a-

Host1:03:36

It is something in the latent space.

Yi Tay1:03:37

Okay, okay. So you can, for simplicity it could be just be like a transformer model pre-trained and-

Host1:03:42

Yes

Yi Tay1:03:42

... and, and-

Host1:03:42

Yes.

Yi Tay1:03:43

So like a more-

Host1:03:43

So to, to me, that is the most coherent thing to the current paradigm, which is you cannot actually do this in current transformers. I think the way that you train it will probably have to be different.

Yi Tay1:03:52

Okay. I see, I see.

Host1:03:54

I, I don't have any conclusion here. I'm just throwing it out as something where I know you're interested in this kind of stuff, and I don't have that many knowledgeable people to talk to about it.

Yi Tay1:04:03

No, I don't think about world models that, that often. I think because world models are just not really well defined in the first place.

Host1:04:08

Yeah.

Yi Tay1:04:08

But I think when-

Host1:04:09

So don't necessarily say world models, but the problems is-

Yi Tay1:04:11

Data efficiency

Host1:04:11

... learning efficiency and maybe, I guess, accuracy or, like, ca- AGI capability that is not easily unlocked right now on our current path of scaling.

Yi Tay1:04:22

Yeah, yeah. So I, I think, I think when it comes to like, like, data efficiency, I think it's more like a unbelievable of finding ways to spend more flops per the token, right? Because you actually-- basically you're-- if you are data bound, you want higher data efficiency because you can learn more from every data point.

You squeeze, you squeeze out more points, right? So things like that can extract more, can use more flops on every token is definitely like a form of data efficiency. Then there's the learning algorithm, right? Because I think there's this-- there's a different scaling law for like humans is this, machine is this, uh, like dogs are this, cats are this.

There's this different, like, exponent, right?

Host1:04:56

That's the, that's the Ilya chart.

Yi Tay1:04:57

Right. Yeah, yeah. There's this famous, famous, famous chart. Uh, and point one and point two are just, like, not entirely, like, d- different things because it could be that better architecture is actually just spending more flops per token.

So if you are-- you come to a point where you are very data bound but not compute bound at all, you just find algorithms that spend a lot of compute on every token, on every token. So I think the, the overarching point is just that, okay, that, that, that it's a learning algorithm thing for data efficiency and then if whether the correct way is actually just to apply more flops per token just to squeeze out more from every, every data, every data point.

Also because humans actually don't like-- like, they are exposed less. When you say less or more data, it's also very ambiguous because they are technically like on like twenty- four/seven, and then you have-

Host1:05:41

And it's mostly visual. Yeah

Yi Tay1:05:42

... have a lot of like different types of inputs, right? And whether they actually spend more flops on everything they listen is also a question because maybe they are just data efficient just because-- I mean, I think somebody needs to count like how much flops the brain use to process, like how much-- maybe they're just spending more compute on every token, and also maybe the learning algorithm is different.

So, but I agree that data efficiency is very important given that, that I think we're gonna, like, there's limited amount of data in the world.

Host1:06:07

One more thing before we go into DSI. You know how like we're talking about RL and like you're, you're working on RL stuff. Why are people paying so much for RL environments?

Yi Tay1:06:17

Wait, so who is paying for RL environments?

Host1:06:19

OpenAI and Anthopic at least. I do-- nobody say anything about DeepMind. So a lot of the model labs that are not you are well known for paying at least seven figures for external startups to create RL environments for them to train in.

Yi Tay1:06:35

Okay.

Host1:06:35

And I think the question is, if you are, your models are so good at coding, why don't you do it yourself? And so I think there's some amount of expertise that's being distilled from human experts into an RL environment that you can then let your agents run wild in.

But I'm curious if there's any other deeper insight than that because I'm not satisfied with my own explanation.

Yi Tay1:06:56

RL environments that are like, they have a lot of domain expertise are probably very valuable. And actually, I, I don't know specifically about what RL environments people are actually explicitly buying. But what, what's the thing that you're not satisfied by?

Like the-

Host1:07:11

It was just so valuable and a lot of people are saying like, "Look, it's a Next.js app inside of a Docker container that logs stuff out when you send inputs in." Then you could probably do it yourself internally, right?

Why you pay so much for some startup that you don't know to do it for you?

Yi Tay1:07:27

I actually have no clue about what-

Host1:07:28

Yeah

Yi Tay1:07:29

... like, why this is happening. Yeah.

Host1:07:30

Yeah.

Yi Tay1:07:30

I have no clue.

Host1:07:31

And a classic example would be like, if you wanna build computer use agents for buying things on a, in e-commerce, you would want RL environments that perfectly replicate maybe the top 1,000 e-commerce websites.

Yi Tay1:07:44

Yeah.

Host1:07:44

And then you just parallel roll out on all of them. Does that seem valuable?

Yi Tay1:07:49

I don't know.

DSI & Rexus1:07:50

Host1:07:50

All right, cool. DSI and LLM Rexus. A big bet for me this year for my conference was we actually, like, started focusing on LLM Rexus. The other-

Yi Tay1:08:00

Actually, what's the motivation behind starting LLM Rexus?

Host1:08:02

Track?

Yi Tay1:08:03

Yeah.

Host1:08:04

I think Rexus is the king AI problem in consumer, is the single most valuable thing. All your feeds, any s- even search is Rexus.

Yi Tay1:08:12

Yeah. Basically, it's search, basically. Like, it's retrieval.

Host1:08:15

It's the God problem, right? 'Cause Rexus is ranking, but then also filtering, also personalization-

Yi Tay1:08:20

Yeah

Host1:08:20

... also re-indexing and-

Yi Tay1:08:23

Yeah

Host1:08:23

... and, and, like, performance. It is the God problem, and you get paid a lot for it. Engineers are not that excited by it, which is very weird because they don't see-- A lot of them don't work on Rexus, and they probably never will.

Yi Tay1:08:35

Yeah.

Host1:08:35

But they don't see the monetary value that can come out of a good Rexus. The other two pieces of updates for me, which I actually didn't even know that DSI, like, directly tied into this-

Yi Tay1:08:45

Okay

Host1:08:45

... was, one, Twitter publicly adopted their feed algorithm as an LLM Rexus.

Yi Tay1:08:54

LLMs are just used everywhere now. But whether it's actually like a gen-

Host1:08:57

Like a big LLM.

Yi Tay1:08:58

Like whether it's, like, a generative retrieval type of-

Host1:09:00

Yes

Yi Tay1:09:00

... models, like, it's, like, another question.

Host1:09:02

Correct.

Yi Tay1:09:03

Is, is it?

Host1:09:04

We don't know. All we know is that they have said that they have swapped out their current Rexus for an LLM-based Rexus. That's all they've-

Yi Tay1:09:10

Okay. Okay

Host1:09:11

... but what has, what is published is YouTube, where they actually adopted semantic IDs for YouTube's Rexus.

Yi Tay1:09:16

Yeah.

Host1:09:17

And YouTube is obviously a big deal.

Yi Tay1:09:18

Is it, like, public information?

Host1:09:19

Yes.

Yi Tay1:09:19

Okay. Okay.

Host1:09:20

They came and did a talk about it with us.

Yi Tay1:09:22

Oh, s- okay.

Host1:09:22

And then they published a V2 this year as well.

Yi Tay1:09:25

Okay. I see.

Host1:09:25

It's more info. And I just thought, so basically, the last time we, you were on the podcast, we didn't talk about DSI that much, but you have actually some background in IR. You care about IR.

Yi Tay1:09:34

No, I don't care about IR, but, but I think DSI, okay, like DSI or, or generative retrieval was like, I think one of my favorite works in the old of like, uh, I, I, I have some IR background in like when I was doing a PhD, I did some Rec- Rexus work.

I did some retrieval work with Rexus and stuff like that. So I have some IR and Rexus background. So I think generative retrieval and generative Rexus is re- all very conflated. DSI started as a retrieval thing, so we did like natural questions like ranking of like documents.

Host1:10:04

Learning, yeah

Yi Tay1:10:04

Like doc- documents that re- everything. The, it started off as, I think that's actually we did an interview with Yannic, like me and Don, we did interview when, when the paper came out, like long time ago. So at that time we wanted to like reimagine retrieval and search.

So we wanted, at that time, LLM, we were still using T5 models at that time. It was like, not, we're not in the LLM era yet. It was pre-LLM era. It was like, okay, pre-training works kind of thing.

And then there were like some pre-trained models around. So we wanted to re- reimagine retrieval, right? But retrieval Rexus, they're all the same formulation, ranking, retrieval problem, right? And then that's where we started to imagine retrieval as one giant RAM that, that encodes everything in the memory.

But we tried so many different, like semantic ideas. Actually, my collaborator, Vin, was the one that came up with ideas, semantic ides, uh, ideas that basically, at the start of this whole generative retrieval was actually basically literally just trying to give a document like identifier and just predicting like raw brute force predicting like this.

It actually works.

Host1:10:58

Real ID.

Yi Tay1:10:58

Like, it actually works because the models can memorize some something. If you look at the literature from all the way to things like Doc2vec, it's very of like model-- The words have no meaning. They're just ID in, in a vocab, but it's another's number, right?

And technically, the models have enough capacity to pr- predict. But I think semantic IDs was an idea that, that basically you have some like semantic association, and then you actually try to break down the search space hierarchy, right?

So how this would evolve in the Rexus was at that time after DSI came out, right? So at Chi's group and, uh, Mahesh, the guy who, who, they did some exploration of applying DSI to, to Rexus, and that's how that generative, generative Rexus recommended system paper came out.

Host1:11:42

Yeah.

Yi Tay1:11:42

Uh, then-

Host1:11:42

I didn't even know he was involved. That's crazy

Yi Tay1:11:44

That was like basically us like transferring this like basically, okay, DSI works. We, we try to try it on, on Rexus. And then I think the recommended system people have a slightly different way of doing semantic IDs, but it's basically just because the pro- the domain is slightly different.

But after that, I think we were done with the invention part, and we just like

Host1:12:05

The rest is details.

Yi Tay1:12:06

The, the, the rest are, are details. So over time, I also left Google and stuff like, so over time, these things evolve a little bit here every day. I think I also saw like something like Spotify is also using something like YouTube, Spotify, like they, they use this type of like semantic IDs, this type of DSI like models.

I think from the research community point of view, like the DSI work was the first one that like decode semantic tokens. But then when we went to, I don't know, academic community is like strange in a way that like they will do things like, oh, this is generative retrieval.

It's not generative Rexus. It's like they'll do this type of like random things that is like a, a bit strange. But yeah, I mean, this was like the whole history of this generative retrieval. Apparently, there's also a lot in, like of people working on, I don't follow-- Actually, I don't follow at this at all now.

It's just not even in my mind.

Host1:12:49

Yeah.

Yi Tay1:12:49

But there's, there was once I went to, even in the S- Singapore office, there are Swis actually like working on- Generative retrieval. They don't work on-- I don't know whether they're still working on it, but I met a person that tried to explain generative retrieval to me.

It was quite funny that, you know, I kind of like co-invented generative retrieval. But yeah, I think this is-- this whole IR thing has been-- it's just an interesting phase, and I, I definitely think DSI is one of my more creative works that I've done that is not like really LM, but it's like-

Host1:13:14

Under the general principle of apply ML to everything, if a Googler is working on generative retrieval, would that be like AI overviews? Is that s-something similar?

Yi Tay1:13:23

I have no idea.

Host1:13:24

Okay. For the people listening, I, I did have a track there. I think you just type in AI.engineer and you'll get it where the Gemini guy was talking-- Sorry, the YouTube guy was talking about how to use Gemini for their REX-IS.

I don't know what size of Gemini 'cause he, he didn't talk about it, but this is public work now, and basically every YouTube video uploaded gets encoded into some kind of codebook, and they, they retrain this every day, every-- on some kind of batch job.

Yi Tay1:13:53

Ah.

Host1:13:53

Yeah, just interesting. So yeah, I, I don't know if you even know what AI-- what Gemini is being used for.

Yi Tay1:13:58

I, I don't follow like this-

Host1:14:00

Yeah, yeah

Yi Tay1:14:00

... this day, these days.

Host1:14:01

I do think like in the sense of like for people who are not-- still not getting it, applying intelligence and the, the general intelligence of an LM to the retrieval, to the recommendation task means you can accommodate such weird recommendations, like such weird queries as well, that normally, like no classical system can ever handle.

And I think like it's also somewhat emergent in the sense that when you were using T5, you just couldn't actually add that much value on top of a normal BM25 retrieval technique. Would, would you say that's accurate? It is not, it's not just about paraphrasing, it is about understanding query intent.

Yi Tay1:14:39

I-- BM25 is a really strong baseline, actually.

Host1:14:41

At least.

Yi Tay1:14:42

Uh, BM25 is a really strong baseline, yeah. Uh-

Host1:14:46

Is that like, unlike I, I-- Sorry, uh, so I don't, I don't know the, the, the, the comparative delta v-versus T5 for you guys versus BM25, but I don't expect it to be very high, and I expect it to be a lot higher for, for a true LM-based REX-IS, depending obviously on the query set.

Yi Tay1:15:02

I didn't really think about it this way before, but because I've done modeling in many different domains, including like search and, you know, in the search commun- AI community, there's also set of benchmarks and stuff like that for like, that people will climb on and some like there's Amy of, or like a, of...

I don't know what is it c-called these days anymore, but generally the modeling dynamics of IR tasks are-- is very different from like in REX-IS tasks. It's very different from standard language tasks or like vision tasks and something like the way we kill climb LM, where you train models into like the way that modeling things interact with this environment is very different.

So I think that I honestly, I hated working on like REX-IS and retrieval stuff. Okay, I'm just talking about old days. When you work on like T5, you work on-- You change, change architecture, you try to improve perplexity, you use SuperGLUE, like this is the olden days.

You are on like even now, when you train LMs, you just do zero-shot, few-shot, stuff like that. Your things that you-- 'Cause as a researcher engineer, you just interact with the environment a lot by this. You're just like, okay, RL by this environment.

But REX-IS and IR has a very strange feeling to it. Strange feeling in the sense that it feels like you're the, like whatever works-- It's like you are in this, in a world with the, the gravity is different or like you are in a world where the modeling things that feel intuitive are not intuitive.

So, so it feels like a very strange space to... So I wrote some papers back in my days on like REX-IS and stuff like that. Well, every time I ran some modeling experiments for, for REX-IS and stuff like that, I, I didn't enjoy the...

It, it feels like the environment was rude, and it feels like the, the vibes are just like, like-

Host1:16:32

What makes it rude?

Yi Tay1:16:33

It just feels-

Host1:16:34

Transactional?

Yi Tay1:16:35

No, no, not, not, not trans- not, not, not transactional. Like I, I don't know how to describe it. For example, like if you play like sports, like play tennis, play badminton, where you hit the ball, you have a very nice feeling like hitting the sweet spot.

When you do modeling in traditional dom-- When you get the feedback back, you feel like everything sounds right, everything feels right, everything right. But REX-IS and IR is like a problem where it's like you hit the shuttlecock, and you hear a glass shatter like randomly.

It's like you just feel this weird-

Host1:17:00

I see

Yi Tay1:17:00

... world that, that-

Host1:17:02

Cause and effect are too far apart

Yi Tay1:17:03

... like, like it just feels strange. And then sometimes maybe the metric like, I think REX-IS, they use like all the-

Host1:17:08

There is-

Yi Tay1:17:08

... GCGFX and then the BM25 is strong, and then you do this, and then you get like worse than the like the, the BM25. Back in the day when you stack two LSTM to three LSTM, you were like, "Whoa, I see life."

It's just the game-

Host1:17:20

Unrewarding area to work in.

Yi Tay1:17:22

Is, is just weird. Also, the IR community and the retrieval community is also like always behind the mainstream, and then now it's just probably gotten even more worse because of LM and stuff. So okay, I'm getting into hot tech territory, but it's just like certain conferences are just like behind NeurIPS and ICML and stuff like that.

Yeah, some conferences are just like-- They're just like applying things that- ... that this-

Host1:17:44

They're downstream of-

Yi Tay1:17:45

They're downstream. Yeah, they're downstream.

Host1:17:46

Okay.

Yi Tay1:17:47

So it always feels very uninspiring to, to, to, to work on, on this.

Host1:17:50

Yeah. Look, there's a reason that you left and-

Yi Tay1:17:52

But it was like a side quest, like where I work on as a side quest thing. Yeah.

Host1:17:56

Uh, yeah. Okay, un-understand. I still think it's an important business problem, even though maybe it's an unre-unrewarding field.

Yi Tay1:18:04

You can understand why. Because the academic benchmarks for those tasks are just so far detached from-

Host1:18:09

Yeah

Yi Tay1:18:09

... they're so far detached from what industry is. I didn't work on any of this like in, in like-

Host1:18:13

Yeah, yeah

Yi Tay1:18:13

... like the, the, the thing, but just from an academic point of view.

Host1:18:17

Oh, then all you need are online evaluators, right? Yeah, yeah, D-test and like, oh, ooh, okay.

Yi Tay1:18:21

That would have been a different experience, yeah.

GDM Singapore1:18:23

Host1:18:23

That is mostly our sort of topic, research topics, coverage and everything. I think we're just gonna end on a very simple one on GDM Singapore. You, you organize a symposium here with Rajat, Dean, Kwok, and all the others.

Basically, what's the general message or the i-impetus for starting GDM Singapore?

Yi Tay1:18:43

So we talk about the event first. So the event was mostly-- So Kwok and I are going to start the team, and then I think before I came back, we discussed this for some time. Jeff was very supportive of this.

He was in the region many times in Vietnam and Singapore last like, uh- About around the time where I was going to come back. And I think that-- So this event itself was Quoc and Jeff was, were visiting, and we just inspire the community here.

I think that it's also a bit more like a soft-- like setting the tone right for the start of the Gemini team in Singapore. And I think it's a very rare instance where you get somebody like Jeff and Quoc who are the true pioneers of AI i-in the world to be in one room.

And then are you there as well? And I think, like many people told me that-

Host1:19:26

Not a true pioneer of AI, but okay. I was there to, to live tweet.

Yi Tay1:19:29

Yeah, yeah. I think having them all in one room and then giving these talks, like many people came up to me and said they were very inspired by their presence in the region. So starting a team and starting something is also very-- like, there's also no one moment.

There is no, "Okay, I press the button, it starts," right? It's like a process, right? So we hire people, and then people join one by one and stuff, stuff, stuff. Something like that, right? So I think this event was more, I would say, like a, to set the vibe.

I think it's possible for Singapore to be close to the frontier, and I think that with having the true pioneers of AI here, we want to give this, like basically more like an inspiring thing, and also good luck for Quoc and Jeff to meet the people here and then-- So Jeff was here last year, but Quoc hasn't been here for some time, and he's gonna have team here, so it's like also nice to bring him around and meet the people here.

Host1:20:15

Yeah.

Yi Tay1:20:15

So it was a really, like, amazing event. We met Along as well, and-

Host1:20:21

Who a lot of people don't know has a CS degree. He's like one of the few PMs with a CS degree.

Yi Tay1:20:26

Yeah, yeah. I would say that the context of the meeting was more like partially like also he, he wanted to learn more about the IMO stuff.

Host1:20:36

Oh, really?

Yi Tay1:20:37

And then also about Jeff. He want-

Host1:20:38

Oh, 'cause he, they invited you without knowing that these guys were coming or something like that, right?

Yi Tay1:20:44

I was a bit like, it's like there's some-- Yeah, Jeff and Quoc and me were, we went to visit, to chat with at the Istana. And we discussed a little bit on Deep Think, discussed a bit on IMO.

And then I think the rest of it was more like Jeff and Lee Hsien Loong was talking more about like-- Be-became less about AI and tech, more about very macro, economical, political thing. Which was-- I was very out of element with, so I was just, I, as, I just-

Host1:21:09

You were in a suit.

Yi Tay1:21:09

Mm-hmm. I, I was just talking about the Deep Think and the IMO and stuff like that.

Host1:21:13

Yeah.

Yi Tay1:21:13

But he seemed genuinely quite surprised that AI has reached this point. So I think-- But it was a, was an interesting, uh-

Host1:21:20

Yeah. I, I would say for people, like you have done something that is unique in Singapore's history so far. Like you are, you're establishing a frontier research lab in Singapore, which is an accomplishment. I think the other thing also that I guess I'm still trying to wrap my head around is, does geography actually matter?

Like you're all working on the team. You have your London people, you have your Mountain View people, and mostly you're just, like, collaborating with them anyway. You've collaborated with them your whole life. I don't e- don't even really know what countries mean anymore when it comes to research or just AI in general, because this thing is just inherently international from the start.

Yi Tay1:21:57

That's a very good, like, question. Also, it's related to, I think, about identity, right? Because I think you also moved from SF and Singapore quite a bit, right? I was in Mou- like Mountain View this, like one, two weeks ago, and like I'm here, but like almost all my-- Like my-- If you just look at my sibling, aside from my family, like everybody I talk to is like somehow in the Bay Area or like just because of work and everything.

I think the geog-geography matters. Okay, firstly, the, the most like thing is like logistical is like probably the time zone.

Host1:22:26

So you literally want the twenty-four-hour coverage around the world.

Yi Tay1:22:29

No, there, there are-

Host1:22:30

Sun doesn't set

Yi Tay1:22:30

... there, there are advantage-- No, what, what I'm saying is that the difference is possibly, pro- like, it's also more like, like people define the location more than the location define-

Host1:22:39

Okay

Yi Tay1:22:40

... the people somehow. Okay, the time zone, we'll get to the time zone a little bit later. Like there are pros and cons. I think there's pros and cons, right?

Host1:22:45

So you bullish on Sing- on Asia, Singapore people, the talent pool.

Yi Tay1:22:49

I think we managed to find like really amaz-amaz-amazing people. But I would also have to say that this type of things is more like talent attract talent. I think most of the time, like people are very excited. Like the vibe I get is that people are very excited because it's like Quoc's team and my team, and we're working on like very core things related to, to, to AGI.

So I feel like the talent we can get from the region is re-really good, but it's only, it's only because it's us that we can unlock this talent. Otherwise, might join some other place like that.

Host1:23:22

Yes, and move to, to the US.

Yi Tay1:23:25

Yeah, yeah. About the identity-wise, I would say that I definitely agree with you that, like, why it doesn't matter. Like that, I think that the advantages of S-Singapore itself or like just anywhere, like it's also that you are like, okay, the world is very global, so technically you can interact as much as you want.

You can also move there. But I think Singapore has this advantage where you can go close, and you can go far. I think the Bay Area is like so much about-- I have s-friends in London and New York that would just never move to the Bay Area.

So I'm not against Bay Area. I think Bay Area is a great place, right? But it's just AI, AI, AI, AI, everywhere, right? I think sometimes if you have some s- like mental space and energy to like, to have some other culture, and then like, you know, like London, Singapore, New York, they have their own culture, right?

But the Bay Area culture is just like AI, right? Like you just go anywhere, you, you just hear AI everywhere, right? Even the billboards and stuff like that.

Host1:24:17

It can get a bit much.

Yi Tay1:24:18

Yeah.

Host1:24:19

Although I did see some billboards down here-

Yi Tay1:24:21

About AI also?

Host1:24:22

Yeah. No, I was like, "What is this-

Yi Tay1:24:24

It's, it's, it's-

Host1:24:25

... culture infecting Singapore?"

Yi Tay1:24:26

I do think that to some extent, if you want to do research in, you need a little bit of peace and quiet somewhere, right? So this island may be good for that, but then you can-- You're still, we're still like able to m- like be connected, right?

Host1:24:39

Okay.

Yi Tay1:24:40

So I think that's mainly... Talent-wise, I think people are strong here. Uh, yeah.

Host1:24:46

So far enough away, but you're still connected. You have strong talent. What are you hiring for? You're still hiring, right?

Yi Tay1:24:52

We're hiring, like, like my team will work on like RL and reasoning for Gemini and Gemini Deep Think. I think we care more about like talent density now, so we're not like also like growing that big this small Team first, just because co-compute per capital is probably, like, important.

And yeah, so I think that's something that, that we're hiring for now. Basically, I personally just... There's a lot DMs, but I think, like, generally there's a lot of people, like, who are very capable. But I think what I'm looking for mainly is either you have, like, a track record of RL research or, like, some...

Even not necessarily the RL, but, or, or, like, some exceptional achievement in, like, coding competitions or, like, some exceptional achievement somewhere, then that's, like, the kind of people that, that we, we want.

Host1:25:40

Yeah. 'Cause you don't strictly require... I do remember something about your record days where you're like, you like to train your own traders from scratch, right? So you don't-

Yi Tay1:25:48

I forgot if I said that or not, but to some extent, to, to, to some extent, yeah, I think we'll, we'll definitely be very happy with people that are, like, very high stats and just, like, even without much-

Host1:26:00

Statistical knowledge?

Yi Tay1:26:01

No, no, stats.

Host1:26:01

Yeah, yeah.

Yi Tay1:26:02

The stat points.

Host1:26:02

Child with int.

Yi Tay1:26:03

Like this high tech- ... stat points people. Like this raw-

Host1:26:07

IQ, yeah

Yi Tay1:26:07

... high ta- tat- talent people like this. I think all, like, strong engineering skills, ML. ML can be learned easily. Our, the, our knowledge can be learned easily.

Host1:26:16

Yeah. I think maybe one version of this is can it be done on a student budget? And where they, where can you do something, and interesting, anymore on a student budget? I would say rele- relevant to the point where conferences are quieter these days.

I did do an inter- interview with one of the Best Paper winners where they worked on 1,000-layer neural network RL, and that was done on a student budget. They... It was very cleanly executed thesis and paper and, and good findings.

Look, I'm not sure if production models will ever go to the 1,000 layers, but they stretched it in an interesting direction and found some good recommendations, and the guy immediately got hired by OpenAI. And I, I think that's encouraging for the grad students in the market who are like, "Okay, well, do I need to know somebody who works at these labs in order to get in?

My uncle works there. I get the internship," or whatever. No, actually you can just do it on a student budget with, with good advisors.

Yi Tay1:27:03

Oh, I actually think one thing interesting is that for most of the people that I actually went to recruit them, like, personally, right? So it's, and you, you see their work, and then you send them DMs, right? So I get a lot of-

Host1:27:15

For the Best Paper?

Yi Tay1:27:16

No, no, no. Like, like, for hiring generally.

Host1:27:18

Oh.

Yi Tay1:27:18

So generally I almost... Like, and to your point, it's like you almost, like, you can just do good work, put it online, and then somebody will contact you, right? It's actually super easy but super hard at the same time because-

Host1:27:29

No, I, I can tell you, like, I talked to a few of these grad students. They don't know what good work means, right? Because they don't know... There's so many things. Their professors have the agenda they're forcing on them, which, like, may not be right, 'cause it's not like their professors know what to work on either.

So yeah, they just need guidance. They just need, "Hey, work on these, like, five things. Show me an interesting result in any of them."

Yi Tay1:27:50

Okay, so if somebody comes up with something and then does something that you feel that is very tasteful, and it aligns with what, like, researchers in the labs like, like want-

Host1:27:59

Mm-hmm

Yi Tay1:27:59

... and they come up with that independently, you know that the function that produces it-

Host1:28:02

They need to be self-driving

Yi Tay1:28:02

... is, is good, right? Like, if you just, uh, go and tell somebody to do this, like, you, you, you get, you just get the signal that these guys can execute, right? So I think there's, there's some value in people that-

Host1:28:13

Yeah, they demonstrate taste, research taste.

Yi Tay1:28:14

Yeah.

Host1:28:15

Re- research taste.

Yi Tay1:28:15

Yeah.

Host1:28:15

Yeah. It's very interesting, and I feel like I could give people... Yeah, I, I do care about this. In some ways, the research directions work that I do is a little bit of that. Like, it's low accountability for me because obviously it's just thought experiments.

But I think for a lot of people it's like their career is bounded by can you demonstrate research taste with this, like, short three, four years or that you have and just do it. Yeah.

Yi Tay1:28:39

Yeah, I would say that this is, it's more there's so much competition just because of, like, the, everybody wants to get into AI. It's just more of, like, how to... Like, mostly it's more, like, how you gonna prove yourself that your, your-

Host1:28:50

Mm.

Yi Tay1:28:51

Right. Yeah. It must be hard these days to be a grad student trying to prove yourself. It's definitely harder, but yeah.

Host1:28:56

Yeah. Not your job. Okay. That was it. Do you have any other sort of rants or topics that you had queued up before we wrap?

Yi Tay1:29:03

No, I don't. Yeah. But it was great. Yeah, it was. I had a great time.

Host1:29:06

It was fun chatting, man.

Health & Wrap1:29:06

Yi Tay1:29:06

Fun, fun chatting.

Host1:29:07

Yeah. E- even, I even love last time we were supposed to do, last time we meet, uh, at the symposium, we were supposed to record. But even we just ended up hanging out and chatting. It's just nice to get the brain dump of what's going on in your world 'cause, yeah, working on really important stuff, man.

Yi Tay1:29:20

Always great to chat with you and, and see you. Yeah.

Host1:29:22

Good to chat. Parting words on the sort of weight loss and workout journey, 'cause that's also a big thing for you.

Yi Tay1:29:29

I think being healthy is important to be, to do good research, right? And I think, I think I've been probably in one of, like, probably in the peak physical health now.

Host1:29:39

Yeah, you look great.

Yi Tay1:29:40

Yeah. Thanks. And I think it's also impacted my work in a good way.

Host1:29:43

You did the sort of Karpathy inspired, like, biohacking.

Yi Tay1:29:49

I didn't go to extreme, but I was- ... like, also quite data driven when I came con- con- I will, like, have my own email, and I will track this. The, I, I was, I'm still supposed to make a blog post about this, but I, I feel like I'm not, like, really at the end game yet.

So, like, when I get there, I will. But yeah, just to, just for people who don't know, like, I think I lost 23 ki- kilos.

Host1:30:10

This year?

Yi Tay1:30:10

Actually across one year.

Host1:30:12

Yeah.

Yi Tay1:30:12

One half years.

Host1:30:13

Yeah.

Yi Tay1:30:14

Uh, so 23.

Host1:30:15

Yeah, basi- literally from the last podcast to now.

Yi Tay1:30:17

Yes. Yeah. Yeah. It's, that's the abolition study now. 20 kilos. Yeah. And I think, like, my HRV heart rate variability has went up by two, two times, and my resting heart rate has dropped by 30 beats per minute.

Host1:30:28

30 beats per minute?

Yi Tay1:30:29

Like, it was like 80, 90, and now it's 160.

Host1:30:32

Oh, yeah, 80, 90 is super high.

Yi Tay1:30:34

Yeah, I was unhealthy. Yeah, yeah, yeah.

Host1:30:36

Okay.

Yi Tay1:30:37

Yeah. So I think-

Host1:30:37

Things, like, when it's hard, like, what, do you have a thing that kept you going? You know, a lot of people, they maybe, they, uh, they're focused on AI, and including myself, right? I do prioritize work. I enjoy work.

I don't enjoy the fitness side, but I obviously it, it feeds into your intellectual work, like, to sort of log off and go for a walk, eat better, all that kind of stuff. It obviously feeds in. But, like, people seeing a positive example like you, they will get inspired to do the same thing.

So I think it's good to set yourself up as an example.

Yi Tay1:31:08

I think definitely helps. Like, when I do these things for my health, I just think that it's also part of work because it helps me to get better at my job.

Host1:31:15

Yeah, yeah.

Yi Tay1:31:15

So it's important as well. I think it's important as well.

Host1:31:17

Yeah. I like that your HRV off the bat. I have no idea what mine is, but yeah. There's a general question about what is productivity and how do you measure it, what, what really matters, and it's still unclear to me.

But I do think general energy level and hunger almost. Like, you almost have to, like, experience physical hunger in order to have intellectual hunger, and I don't know if that's like a thing.

Yi Tay1:31:44

No, when I'm hungry, I just think of food. So- ... I think to me it's like these disjoint things. But when you, when, when it's hard to do work when you're hungry. Yeah.

Host1:31:52

Okay. Thank you so much.

Yi Tay1:31:53

Yeah, thanks. It was really great. Yeah, have a great time.