Origins0:00
Welcome, Yi Tay, to LaneSpace. Uh, it was-- This is a long time coming, but I'm so excited to have you here.
Yeah. Thanks for, thanks for inviting and, uh, excited to be here. Talk about a lot of stuff. Yeah.
Um, so you're interesting to research and in-introduce. Um, you are now chief scientist of Reka, um, which is, uh, whi-which is a super inter-interesting model lab. But before that, you were at Google Brain. Um, you were architecture co-lead on PaLM 2.
You were inventor of UL2. You're a core contributor on Flan. Um, you're a member of the Bard Core team, and you also did some work on generative retrieval. Um, that's a very, very illustrious three-year career at Google Brain.
Yeah. Thanks. Thanks. Thanks. Yeah.
Uh, and then since then, Reka, you joined in March 2023, announced a $58 million Series A in June 2023. I don't know if you know the, uh, post-money valuation or the pre-money valuation is public. Um, so it's-- Crunchbase is, is, is-
Oh, okay, okay
... uh, 200-
I, I did not know that. Yeah
... 50 something million. So you, you don't even have to leak it. It's, it's on the internet.
Okay.
Um, in February-- Uh, so, so the Reka, Reka's stated goals were to work on universal intelligence, including general purpose multimodal and multilingual agents, self-improving AI, and model efficiency. Um, in February, you released Reka Flash. Um, in April, you released Reka Core and Edge, and then most recently you released ViBEval.
Yep.
Is that a good summary of the last six years? Yeah. And we, we can go deeper into, like, the, the specific papers.
Six. No, it's not s- five, four years.
Four years.
Yeah.
Oh my God.
Okay.
Okay. We're talking about AI a long time.
Yeah, I was wondering, like-
Yeah
... since when did I, like, step into a time machine or something.
Yeah. Okay. So, uh, can we just talk about, like, your, your transition into-- You know, you, you did your PhD, and we can talk about your PhD. Um, transition into, into Brain and research and, and all that. Uh, you know, I saw you do some work on recommender systems.
I saw you do some work on, um, quaternions. What the fuck was that?
Okay. Let, let's, let's forget about that.
Okay.
We just-
De-de-describe your path into modern LLMs, right? Because you were-- you weren't in-- you didn't start there.
Yeah. Okay. Sure. For-- I, I do, uh, it-- the world also didn't start, start there, right? I mean, I think in, uh... So I joined Google in 2019, end of 2019, and the world looked, like, really different at that time, right?
Uh, and I think, like, uh, I think at-- that was around the time the first GPT was released by-- GPT-1 or something was released by OpenAI.
I like that frame.
So research, like, like ML research and, and NLP research, like, um, like, looked very different at that time. Uh, so I was mostly... I, I, I identify, I iden-identify as, like, a, like, a language researcher. Uh, I don't like to use the word NLP.
Jason will kill me if I use the word NLP. But, like, I was like, okay, a language researcher. I, you know, like-- But I was more like a architecture, model architecture kind of researcher. Uh, and when I joined Google, I was also-- I, I continued on this, like, as the model architecture research.
I did-- I, I worked a lot on, like, efficient transformers. Uh-
That was your first viral paper.
Uh, yeah, yeah. And like, uh, you know, I, I worked, like, long range arena. I, I spent quite a lot of time looking of like could we do without a-attention. Like, there was a synthesizer paper back in 2020.
I think that was, like, my early days in Google. There wasn't, like, like a... At that point of time, like, transformer research was mainly like, like WMT, like machine translation and, like, perplexity and stuff like that. It's not really about, you know, there, there wasn't like, like I think it was-
Reasoning
... few-shot, few-shot learning and few-shot in-context learning came only about like, you know, when GPT-3 came out and beyond, right?
Yeah.
And then, uh, so I think that at that time, the, the meta, I would say, the meta looked very different. And, uh, at, at that, at that time, a lot of the work were focused on, like, fine-tuning things like T5 or BERT or some-something like that, right?
So, uh, I think a lot of the research, not, not only myself, but, like, around me or, like, even, uh, the broader community were working on those kind of things. Uh, and I think like... Yeah. So I think that, that was-- I, which I feel that in hindsight today is actually really useful to like, kind of, uh, think about because, uh, a lot of people came into, like, AI and then into, like, right after, like, ChatGPT came out, right?
So they, they, they saw AI as like, kind of like, uh, they, they... You know, I think there's a lot of benefits of like, like, uh, you know, understanding how, you know, transformers and like, like if you, if you tr-- Like I've broken this thing apart, like, so many times trying to, to...
It's like these things actually, you know, help to improve intuition and, you know, I think like, uh, it's not like totally disconnect. Like, I think a lot of things are still, uh, relevant today and, and it's just the scale has gotten, uh, much larger and also like the, the paradigms shift a little bit from like single task fine-tuning to like generally s- do everything kind of universal-
Foundation models
... foundation models, right? I think it's just a slight change in paradigm. But fundamentally, I don't think like the, like the, the, the stuff has actually, like the underlying principles of research hasn't really changed that much except for like compute.
The, the, the-
Compute, data.
Yeah. Yeah, yeah.
Uh, yeah. So basically algorithms stayed put and then compute and data scaled.
So I, I have, I have a, I have some, some thoughts about this, right? So I think like back, back then, like, uh, a lot of the academic research, I think people have talked about this, like, like Sasha Rush has talked about this or like other people have talked about this.
It's like the, the, the conferences were always organized by like applications, right? They were always organized by like, oh, like question answering, this kind of thing. And even in twenty nineteen-
Wisdom, which you, you did some papers on
... like, like, like it was, it was al-al-always like this, right? I think there was-- there's like a bit of a transpose going on. Things become universal and then becoming like, like, okay, there's a data work stream, there's a model architecture work stream, and then people work on improving like a, a, a, a universal model and, and, and general purpose algorithms to improve this model rather than finding domain specific tricks.
I think th-this has-- I think for even in 2019, I think I've already been Like focusing on, like, works that are like, like, you know, you could improve on general architecture. At that time it was like, like maybe LSTMs in 2017 or something, and then you would try on like 10 different tasks-
Mm
... and that kind of thing, right? But like a lot of the research community have been focused more on like, um, you know, like, how do I get that extra 2% on question answering or like, and then, or sentiment analysis.
I think that, that there was this phase of like in 2017, 2018 where this type of work was like still very fashionable in, in, in, in academia, in conferences, right? And then I think the, the big thing about the, the, the ChatGPT moment of like 2022, the, the thing that changed drastically is like it completely like it was like this sharp, like, like make all this work, like, kind of like obsolete if, if, uh, like, uh-
You've r- so November 2022, you're saying? It's-
In the Chat- ChatGPT, I think-
Yeah, ex- exactly ChatGPT launched, because I feel like if you're in the research community, this was coming.
Yeah, yeah. That's what I'm saying.
Okay.
I'm saying that like, uh, in, in like the big labs and stuff like, like, or p- people have already been moving-
Yeah, yeah
... towards general. Even T5 was already like general purpose.
Yeah.
And the, the thing, right? But like there was, it's like there's a bit of a time, like, like lag to, like for like, okay, like p- places like Google and, and, and, well, Meta, OpenAI, we will be working on things like three years ahead of everybody else, and then suddenly like then academia would be like still working on like these task specific things.
Got it, got it, got it.
And then like the-- I, I think the, the forcing function was the, the ChatGPT moment actually really like, like i- i- it was coming, it was coming. It was just like the, the final, the, the last, last straw and then it's finally like-
Yeah.
Yeah.
Now it's serious.
Yeah. Now, now it's really the, the, the thing, the thing completely changed. But-
Yeah
... yeah, so I think that, that was, that was-- I, I don't know how it turned from my, from my background to like, um-
It's okay. Cool
... talking about the meta.
Uh, I think that you navigate the meta very well and I, and part of my goal here is to also isolate how you think about the meta for other people to reflect on, 'cause I think obviously you do it very well.
Oh, thanks.
Uh, a- yeah, somewhere around-- So I'm looking at your papers published. Somewhere around 2021, you had a hard cut to 22, YOL2 and PaLM, and you did YOL2, PaLM, emergent abilities, DSI, reci- recitation augmented generation, all in the same year-ish.
Mm.
Um, so like there was-- Were you-- Did, did you change teams? Did you, did you like have a research focus? Um, like when did you become-
Oh, so you're saying that like my research became-
... the language model guy
... my research became emergent, right? The emerge, emergence.
It was like it's very obvious.
No, I, I, I, I don't think I, I, I don't think, I, I don't think I'm like a person that like, like I'm not like super, super great at like foreseeing a trend like two years a- ahead and then like, like, and then like by special- especially like plan for that, right?
Yeah.
I think I smoothly and as like kind of like as-
Okay
... the field moves-
To you, it was smooth
... like very like, like, you know, it's, it, it, it, it is-- it, it didn't felt like I never actually had a time where I said, "I'm going to pivot myself into like this, this, this, this, uh, uh, you know, this..."
I, I never actually really thought about this, this way. I just did, like at every step I just like optimize for like what I found to be most impactful and most promising. And then that gradually... And also i-i-it's also a lot of influence by, by talking to people, right?
I think, uh, at that time I started working more with-- I had some close collaborations with Jason and, and other people.
Few people.
Like, I mean, Google is a very-- you can work with anybody you want, basically. So, so you, you, you're kind of like also like partly it's like the environment shift. And I think the environment shifts like, like very quickly, but like I, I was also always very like-- I was always like polling for like the, in the environment.
I was not like, uh... I think it's, it's, it's always good to like have open mind and move along with the field rather than like, "Okay, this is my research area. I'm going to get stuck with it two years."
I think I just move along to find like things that interest me and naturally I think like that turned out to be like the things that were most impactful at that time. Uh, I mean, I, I think, I, okay, I mean, if, if, if you put it that way, it's like, okay, I kind of in retrospect I kind of did well, but like I never actually really s- like saw it as like the intentional-
Sure.
I, I didn't like do anything like really intentional except just like doing what I find like interesting, actually. Uh, yeah, uh-
Cool. Well, well, we'll, we'll just talk about the, the main work at Google Brain and then we'll, we'll move to, to Reka. Um, so out of YOL2, PaLM, emergent abilities, which, which of these came first? There is, there's Flan as-
Google Breakthroughs10:13
Right
... Flan as well.
Well, actually I-
Flan was 2023
... wait, I need, I need-- I can't really actually re-re-remember.
Okay.
So-
We, we can just talk about YOL2 then.
Okay. So, so, uh, YOL2 and DSI, I, uh, the, the differentiable search index, I, I, I, uh, I was working on it like, uh, in the December of, uh, 2021. Uh, and, uh, so like at Google there were like projects that are like big efforts that are like, uh, you know, like, uh, you, you like a researcher will be like part of the effort and then this would be, uh, kind of top-down-ish, uh, to some extent, right?
And then there were, there were also like bottom-up like research that one could do. Like I, I, I can't speak for the Google now for sure, but like at least, at least at that time, right? Uh, so YOL2 and DSI, differentiable search index, were like works that I kind of tinkered with in the December break where like nobody was around.
Okay.
And then I was just like working aro- o- o- on it. So YOL2 and DSI were like kind of PaLM, like they, they-- PaLM also there's this, uh, kind of, uh, uh, diff- differentiation because there's PaLM 1 and there's PaLM 2, right?
So PaLM 2 I was actually like the co-lead of one of the work streams, but like PaLM o- PaLM 1 I was more of a contributor, and PaLM 2 I was like... So, so they, they were like, now I have to think back of like, okay, what's the timeline?
Which came first, right?
Oh, yeah.
Uh-
Well, you don't have to-
No, no, it's fine. It's fine.
No, no, it's not like a-
But, but I mean, in, in general there were like, like kind of three categories of works. One is like broader efforts that, that, that are like maybe like org level efforts, and then there are some that like YOL2 and DSI were like my own-
Projects
... like projects. I used the compute that, that I had, and then I just played with it.
You accidentally le- left YOL2 running for a month.
Yeah, yeah, yeah. That was in the paper. Yeah, it, it, it was fun. I think it was really fun, I think. Uh, and then there was also like a third category where those were like- The efforts that my good friends were driving and I contributed.
So FLAN was just one of the... I, I know, like, maybe on, on... I, I would like to just maybe say this publicly, like a lot of people, like, like because I, I, I-
You're very publicly-
I, I talk, I, I talk a lot about FLAN but-
You're FLAN shill number one.
But, but like, yeah, but like the, the first author is actually Xiong Wan, who is great, and then, like another guy, Le. I was a co- core contributor, but I, I mean, I, just because I'm a little bit more visible, so I, I kind of accidentally like took a little bit-
Sure, sure
... more credit for that. But I think, like I was a co-contributor, but I was not like the network-
The, the lead authors are obvious.
Yeah, the, the, yeah. So, so, so I just, you know, sometimes I get like accidentally, uh, uh ... But I think in general like, yeah, so the third category was like projects that my friends... Like Emergence was also like, uh-
Jason's
... emergent abilities.
Jason's paper.
No, actually it was the pa- the, the paper was actually supposed to be only me and Jason on the paper actually.
Okay.
And I, I actually became friends with Jason from that pa-
Because of that
... from that paper.
Ah.
And then that led to like this streak of like, I don't know, 10 papers or something together with Jason, and now we are like super good friends and stuff like that.
The ultimate bromance.
Uh, but that was like the emergent, emergent paper. But, uh, the, the emergent paper was also like a, belong to be like a, a bottom up kind of like a thing and, and yeah, I think, yeah, w- fun times.
Yeah, it was fun. Yeah.
Okay. Um, yeah, all right. So maybe I'll pick on PaLM 2 because I feel like... I'll pick on PaLM 2 and Emergence-
Mm
... 'cause I really wanna make sure I tell those stories. Th- those are important stories. PaLM 2, I think it's a career story that you, uh, effectively became a co-lead on, on the second version of a very high profile, um, company-wide effort.
Um, how did that happen? Um, you know, I think, I think people would, would like to know, to, to know how to, uh, you know, what, what's like the, the sort of career strategy there.
Uh, so like, uh, to be clear, like I was one of the co-leads, but there were a lot of co-leads also. So, so I, I don't want to take like too much credit for that. But like, I think, uh, so I- my involvement with PaLM 2 came from the, like after UL2 was working well and then, uh, it was gaining some visibility within Google.
Uh, and then, uh-
Just a, just a documented re-note. Uh, was UL2 the largest model that Google had released at the time?
20 B, the, the open source?
20 B.
Yeah, I think so. That was the largest.
And you just ... It was a personal project.
It was a personal project, yeah, yeah.
Isn't that unusual?
Like, I'm just like, how can it be like one person's decision to like suddenly release something that, you know, effectively changed-
I think how-
... the trajectory of Google Brain
... how, how, how it, how it worked was that, I mean, 20 B is not that much larger pa- from 11 B to 11 B, uh, T5. Actually, at that time there was 13 B MT5, right? Uh, so I think UL2 is encode-decoder 20 B model.
I think when we got it approved, it was like, uh, kind of, you know, it was released as like kind of like a, like the big brother of T5, you know? Kind of like a, okay, we, we updated T5 with like a new objective and trained this new model at 20 B and we want to...
And it uses the same like pre-training data set and everything, right? So like from-
Pure C4.
Yeah, from, yeah. That was the easiest because there was precedence, right?
Yeah, yeah.
It was like, okay-
But, but yeah, you ... There was some archite- architecture, like the mixture of denoisers, uh, difference.
Yeah, yeah. Uh, so, so back to PaLM 2, I think my involvement with PaLM 2 came from the, the, uh, the work to, to, to, to add UL2 to PaLM 2. Uh, and then like I, I, I mean, it was from the top-down point of view.
I, I thought, I mean, the leads were decided in a top-down, uh, manner. It's not like, like what there was not much like fighting or like, or any, uh, major things, right? It was like, uh, it, it wa- it was a mixture of like bottom up, top down-ish, like half-half situation.
And then like, like from the top it was like, okay, like, like, uh, these are the people who are the most visible in, in contributing to this work stream. And then, okay, how about, uh, he and this other guy becomes, like will be in charge of this like modeling work stream and something like that, right?
So I think it was just hap- it just happened that way, uh, like organically. Uh, and, and, uh, and yeah, I think that that was h- like how, how I, I, I, I kind of, uh, was co-leading the, the modeling work stream of PaLM 2, yeah.
I think in retrospect, you understand now that this is very valuable experience. And I think now today it would be much more competitive to, to get the, the job that you got, whereas you didn't, you know, two years ago, you didn't have to try that hard to, to get it.
Or, or like you kind of lucked into it with UL2 and then, and then like, it just compounded from, from that initial good decision.
Uh-
Do, do you think that, do you agree with that or?
I, I, I think it's very hard to like counterfactually analyze these type of things. Uh, like, uh, I... It's, it's, it's, it's hard to s- Like, okay, I, I think it's, it's definitely true that there are more people like working on generative AI now and, and you know, if you are in a big company, it's way harder to navigate like these type of things, right?
Uh, I wouldn't say that there were like nobody also like wanting to work on this like at that time. Uh, in fact, there were like actually-
But you were the obvious choice.
Uh, there were less, there were less people there. There were, there were def- definitely less people, but I think it was also like, uh, uh, l- uh, how, how do I, how do I put it? Like it's, it's ...
I would say that it maybe is slightly harder now, but like it's also not like it was easy at that time. Yeah.
Yeah. Yeah. I, I imagine it's sensitive, but also like, y- you know, in my mind this, like this is now the most valuable on-the-job training in the world, and so people wanna know how to, how to get it.
This is what I'm trying to, uh, figure out.
Uh, I mean, it, it might may not be, it not may not ... Yeah, I, I also agree that like, actually indi- like individually, like, like you can, you, we also cannot take like somebody else like, like experience and then try to replicate it on...
Plus everybody's like circumstances, their initialization point, their, their thing is kind of also like in different-
Yeah
... uh, uh- In, in dif- I think this is not only f- uh, true for, for LLMs in general, right? Because a lot of times, like, like, oh, okay, you, you, you, you, you did this in this position, and then because of this...
It's, like, it's very hard to, to, to, to, to, to trace all this down to, like-
Right
... to find the causal path for these things. So, uh, yeah, I think everything in life, there's some luck involved, I think.
Yeah, there, there is. Um, emergent abilities.
Yeah.
Uh, very influential paper, uh, subsequently contested by the Mirage paper. Uh-
Oh, yeah, yeah.
Uh, so before we get to the Mirage, uh, was there a story behind emergent abilities? Uh, yeah, you know, I'm sure it's Jason's, like, thesis or... Like, w- just tell, just tell more about, like, the behind the scenes.
Uh-
Like, was, was, was there a discussion that led to it that, you know-
Okay, I, I d- I have to, I have to, like, really be... Like, this, this one was, like, this, the idea, the inception of it was, like, m- mostly Jason.
Okay.
Right? I think, uh, like, I, I helped out to, like, you know, uh, shape up a little bit of the paper, uh, get some, uh, stakeholders involved and stuff. I was, like, like discussing quite a bit with Jason, but this, the idea itself was, like, like, like, like Jason itself.
So actually, when the Mirage thing and everything came out, like, I, I didn't like... Okay, I wasn't just like hot takes for the sake of hot takes. I, I didn't, like, feel... But I, I believe in emergence. Like, I have to just, like, go on the record and just say, like, I mean, I, I believe in emergence.
And then, like, but, like, I was not, like, feeling very strongly because I think that, like, uh, I, I can't speak for Jason, but I would just imagine that he would be maybe personally offended because- ... because I, I'm, you know, I know, I...
Jason is a, is a person that takes a lot of, like, feedback, like, very well. He's a very, like, like, he's not offended by harsh feedback, and he, he, he rebuts well, like, online a- as well, right?
Yeah.
But, but, uh-
One of the most thoughtful writers
... but, but he, he, he, he... Like, I would just imagine he would be the, the one that is the most, like, uh, uh, actually the most, uh, affected by, by criticisms of emergence. I was, like, believing in, in, in it, but I have to say that the paper-- I mean, that's why he's the first author and I'm second.
But, but, like, that was mostly, uh, Jason's thesis. And I have to really say that, like, uh, uh, like, uh, Jason has really good ideas and, and, uh, and I, I was more of, like, a, a support role-
No, yeah, sure
... for, for that, for that, for that paper, yeah.
Sure. Yeah. Um, yeah, cool. Uh, you, you know, lots more to, to discuss there, but you believe in emergence. That's, that's, uh, enough-
Yeah
... enough for me to, to work with. Um, okay.
No, I, I also think, I also think that, I also think that the, the, the Mirage paper is mostly, like, uh, I don't know who-- Actually, I, I don't even remember who wrote it. Like-
Wellen Schafer. Yeah, I, I covered him on, on my NeurIPS podcast. Um-
Okay, okay.
He's a very good speaker, and the paper was well done. It's just that people drew the wrong conclusions from the paper because it had a very good title.
Do you believe in emergence?
Of course.
Okay. High five.
I, I mean, how can you read any paper, uh, read any, the, the progress of LLMs and not believe in emergence? It's so stupid. Like-
Yeah.
... just because you recan-- you can repara- para- per- reparameterize some benchmarks and evals and, you know, make it linear, uh, doesn't mean, like, emergence is completely gone. Uh, and even in the Mirage paper, they, they, uh, acknowledged that there were some, uh, metrics that were true, genuine emergence according to them.
I think it was something like 25-ish percent in the ballpark. That's not the exact number-
Yeah, yeah, yeah
... but it's in the ballpark. So I was like, okay, fine, like some m- some benchmarks you disagree with, but on the whole, there is emergence. It's just, now we're just ta- talking about the magnitude.
Yeah, yeah, yeah. For sure. For sure. I, I, I think, I think, I, I, I, I, I don't think, like, the authors of the paper had really, like, very, like, they, they didn't-- I mean, nobody-- We, we should just assume people don't have bad intentions, right?
But, like-
No
... they, they definitely were just doing this. But, like, the, the, I think I was-
The popular media
... I was more, like, like, annoyed by, like, the, the NeurIPS Best Paper Award. I mean, okay, Best Paper Awards, let's take a, take, take it with a grain of salt, right?
Yes.
But, like, there were people coming at me like, like, "Oh, you should care about this paper because it's the NeurIPS Best Paper Award."
It's been disproved.
Like, because it's-- Like, they were like, "Okay, because it's the NeurIPS Best Paper Award." I'm like, "Does, does Best Paper Awards, like, mean anything actually?" It doesn't mean anything, right? Like, but, like, I think that was more of, like, my, my, like, where my angst was coming from, right?
I, I don't, I don't think, like, I really had a-- I don't even remember who were the authors of that, that, that-
Yeah, yeah
... that, that, that paper, right?
Well, I'm sure they're do- they're doing well for themselves and-
Yeah, yeah
... uh, you know, it's, it's, yeah. We don't have to dwell too much on that.
Okay, okay.
Uh, okay, so a couple more things from Google, and then we can go to Reka. Um, Kwok Lew was a manager.
Yeah, yeah.
Mentors & Twitter22:33
Uh, what is-
I, I had another manager called, like, Don. Like, I, I had two managers during my time at Google.
Uh, so I'm just basically gonna ask for quick hits from, like, what did you learn from Kwok? What did you learn from Jason? What did you learn from Hyung Won and-
Oh, okay. Very interesting.
Yeah, like-
Uh, uh-
Your, sort of like your mental embeddings of, like, who they are, who, who they represent to you, like, how they advise you and all that.
So, so Kwok as a manager, he was more like a friend than, uh... And, and we would, like, talk a lot about res- I think Kwok is a very research-y person. He has a lot of, like, good in-- Like, he's more of, like, intuition, uh, uh, person.
I, I learned a lot about, like, like, from him about, like, uh, you know, like, uh, it's not very ex- like, e- explicit or like, or it's not, like, exactly like, uh, there was no, like, con- concrete. Like, it was more of, like, over time and it was very implicit, soft kind of feeling.
But I think, like, a lot of research sense, we would, like, brainstorm a lot about, like, you know, like, uh, uh... I, I, I quite, quite like that, you know, when we were-- There was this new palm paper that, that didn't, like, get much, like, as much attention that, that, that, that I feel it deserves.
But, like, I think that was one of the works that I, I kind of, like, like, discussed with Kwok quite a bit and, like, at that time we were releasing the Flan 2 stuff and everything. And then, like, I think Kwok has a lot of good sense about, like, what makes a work, uh, uh, a, a, a, a good hit and, like, uh, you know, publicly a good hit and, like, a lot of re- research sense about, like, what, like, uh, what makes, like, like, research, like, cool, you know?
So I think he has good, like, intuition as a researcher, and I learned quite a little bit about... And I also, I also say that I think Jason also probably learned, like, quite a bit from Kwok and, and this also influenced his, his taste, uh- So I, I would guess like, like more of like it was, it was not only just like me getting influenced, but like there was like a, like Jason getting influenced-
Yes, critical mass
... and then Jason influenced me.
Yeah.
And then there was this like... So I think, uh, overall what I learned from Cox probably is more of like intuition, research, taste. We would like chat about AGI sometimes, singularity and stuff like this. Like, it was like, you know, I, I, I, I learned like quite, like he- he's, he's nice to talk to as, as, as a, as a friend, manager, kind of, uh, uh, uh, like f- he's like kind of a friend figure to me and, and a research, like a researcher.
He was, he was very, he was, he's very much a researcher more than like a, uh, like a corporate, you know, m- manager kind of thing.
I-
Yeah
... totally expect that, yeah.
It, it was fun. It was fun. Uh-
Uh, since you mentioned AGI, uh, we actually don't cover AGI on, on this podcast.
Yeah.
Mostly because it's very hard to be precise or make falsifiable claims. Um, do you perceive differences in the way that AI researchers discuss AGI compared to the re- regular population?
Uh, so I've just, I don't think that we were make- making any progress in quantifying it like that.
Okay. I can skip that question.
There was a lot of like, uh, uh, this, you know, uh, uh, a fun, fun chatter around it, but it was not exactly like, uh, uh, yeah.
Yeah.
Yeah.
Um, Jason Wei, uh, what, what did you, what do you find, or what do you learn from him? What is your distillation of the Jason-
Jason-
... cognitive function
... Jas- Jason, Jason, okay, Jason is very interesting. So, uh, I, I learned, like in, in my career, I learned like two or three things, I, uh, major things from Jason, right? So, uh, I think the first thing I learned from him is that like, so Jason was actually...
Okay, I, I'm going to talk about the more casual, more fun stuff first. Jason was the mo- like, uh, was like, uh, more spicy on Twitter first before me.
Mm.
I, I, there was an era where I was like a goody two shoes. I only had my main account. I would only tweet... Like, my only tweets would be like, "New paper alert," you know?
Yeah.
Right? And then Jason was starting to post like, like-
Hot takes
... hot takes, right?
Yeah.
And, and I just thought to myself, "Oh, damn," like, like you know, like and there were, there were times that I was like, "Jason, you should not post this. You're gonna get canceled." Right. And he, he, he was fine.
He, he always braved through the storm and everything.
Yeah.
Until I, I looked at him and I'm like, "Okay," like, uh, maybe it's not that bad after all to-
Yeah
... to, to, to just be, be-
People love it
... right. Right.
Yeah.
Uh, so that was like kind of like the, the... Which is very interesting because Jason is much younger than me, and I saw this, uh... And the other thing also, we, our alt accounts, right, we created them around the same time, right?
Yeah.
And the interesting story behind it was that like, so, uh, Jason's alt account and my acc- alt account has our own, our, our original idea. Like, it was not like a anime character alt that nobody, like, like know who is it.
We have our identity.
It's pseudonymous, yeah.
It's pseudonymous, right. And then I asked Jason like, "Why do you want to, like have a, a pseudo, like why, why don't you just make like..." Right. And he told me this thing which was quite true, was that like if you cannot, like okay, you can post a take that is spicy and it's hot, but if you cannot stand by the opinion, then you should not have the opinion in the first place, right?
Wow.
Right. So that was something that, oh, okay, I thought that was profound because so far this, I mean, there are times where, okay, I post something and it's spicy and then, okay, it gets a little bit bad ash- And then I, I, okay, I kinda agree that, okay, this is bad, then I will retract it.
But if I could stand by the opinion, then I would just stand by it because like that's the point of making it like, like-
It should be said.
Right. It should be said because I, I can put my name behind it, right? So that was a like the, the, th- this is part of the first bucket about like, like, uh, like how, uh, uh, you know, uh, l- like kind of influence like my, my, my online persona like a little bit.
Like and then, uh, I, I, I mean, and then it, it turns out that like now AGI Hippo is so much more spicy than- ... than, than the Koala is. Koala is just hibernating somewhere. He's not even around, right?
So, uh, I think that, that was w- something that, that, that, uh... I mean, Jason also is more constrained because he, he works for, like he has like a, like an actual like employer.
OpenAI, yeah.
Right. And he has to be a little bit more-
My God, I, the worst thing about Twitter is that, you know, anytime anyone from OpenAI tweets anything, they're like, "Did you see this researcher from OpenAI said something," blah, blah, blah. Like, and they, they read tea leaves that are not there, and it makes you very cautious to tweet anything, and so it kills the golden goose is what I say.
Right?
There was one tweet, I mean, at, at the time when somebody was, people were speculating the GPT-2, GPT-2 chatbots, right? And then Jason p- just posted something like on his main account, like something like, uh, "I can't, I, I'm like excited about like new experiments being run."
Like just like a random, like just... And then people like screenshot that and like post like, like-
Yeah. That, I, I hate that. Yeah.
So I think, I, I now I think like for, for his alt account it's mostly like personal, like personal stuff. Like, you know, like very pers-
Yeah
... like I, I, I think he would stay away from like-
Non-work things
... like, like a non-work thing. So I think-
The, the, the golden goose has been killed because people on Twitter cannot control themselves from like drawing random conclusions from, you know, all these like hints and all that. Like it's really-
Yeah, yeah, yeah, yeah.
Yeah, it's, uh, but-
Okay, but, but, but like going to like the actual, like this- ... this, this is like filler, filler. This is filler.
It's okay.
This is not canon, it's filler, right? Uh, I, I think the second thing I learned from Jason is more about like the, like as from my c- uh, you know, kind of like from my own career is like the importance of like, like, uh, like marketing and, and PR.
So Jason is actually like super good at like, I mean, I would just... Like, he was actually like really... You, you know, the emergence, like how many blog posts he wrote about the, the emergent abilities and how many talks he's given about, about emergent, like a lot, you know.
Like probably like the other day I was just at this, uh, uh, uh, Webcom keynote, and he was giving a keynote again about emergent abilities, and it's been two years, right? So I, I think one big success of him is that like he, he does the work.
He, he, he, he, he, he thinks a lot about like marketing the work itself, right? I, I, I did not like in, in my early parts of my career, early parts in Google, right, I was, uh, I, I think I, I was putting out a lot of work, but I didn't put in a lot of like effort in like, like thinking about the, like how the work is gonna be received.
I'll just be like Here's a paper, here's a paper, here's a paper," right?
Mm.
But Jason will be like, "I'm gonna write this paper, and I'm going to, like, market the shit out of it."
Mm.
Uh, so I, I, I think I, I learned-
Mm
... a lot about, like, like, uh, like every single-- So every single first author paper that, that, like, Jason writes in the last y- has, like, 1,000 citation in one year.
Oh, my God. Okay.
Like, like, no, I mean, not every, but, like, most of it that he leads. Right.
So his hit rate is very high. Yeah.
His hit rate, like impact density, like, is very high, right?
Yeah.
So it is, it's pretty interesting, like, it's pretty interesting, like I, I, I kind of, uh... So Jason is way more like young-- Yeah, he's way younger than me and more, like, like technically, like so-called more junior. But I kind of see him as like a, a peer, and I learn a lot from his, uh, uh, basically, uh...
Some, some people are just like talented in, in, in, uh, different ways. And, and I think that, like, I, I looked at how he, he markets his own work and markets himself actually, right? Uh, I think that's such a, such a, uh, something that, that, that I could learn from, from, from, from, from, from that.
If someone is starting from zero, like no Twitter presence, what is the second-best thing to do if, if you don't have a Twitter presence for-
You mean as a researcher?
... for marketing, yeah.
I, I, I think you would like the, the, the most obvious thing to do, like if you're like a research-- Like, say hypothetically, you're like a researcher in like a place without visibility or without-- and then you have no personal visibility.
The, the first goal is always to, uh, try to f- to find a mentor or co-author that is like within this circle, and then you start from there, right? Because-- And then you, you get people from, like who has a visibility and following to-
I see
... to, to retweet. So you would, you would like work with them. Like you-- The, the, the, the, the big goal is not about like, uh... I, I, I learned this. I mean, this, this is like probably a career mistake in, in my, my early days w- was that like, you know, instead of like focusing on like so-called people, like, okay, if you do good work, you know.
It's more of like, okay, how am I going to like say, uh, say I, I see this, uh, uh, visible researcher from DeepMind, right? Or how can I collaborate with this person and then like kind of, uh, uh, uh, uh, uh, uh, do something that like they feel is cool and like I, I can win their respect and that they would like, uh, you know, they would be willing to co-author for me.
Because the exercise itself was so about how to... You're not trying to please reviewers or anything. You're just-- If, if you can find one semi-visible, even-- It don't even have to be like a famous person, just like a semi like few tens of, uh, not tens of, like thousands of followers, has a good reputation of research, and then you collaborate with this person and then like, uh, like when you post the work, you are co-author of this person, and then like you get the person to like, like, uh, vouch for you or like just portray.
Over time, this would like-- It could be from internships. It could be from like-- It could be from, uh, uh, you know, just DMs. I think, you know, people, people are nicer than like, than like some people they, they seem scary, but like if you DM them, they're actually willing to collaborate actually.
Uh, I was scared of you actually.
No, no, no.
And when I DM'd you, you s- you turned out a lot nicer than, than I, than I- ... feared. So, uh, thank you for being nice.
Yeah. Okay. Okay. I'm sorry for-
That's good advice. No, no, no. I mean, obviously I, I, I, uh, we didn't know each other before, and then, you know, now, now I think we're, we're getting a bit more, uh, friendly. Um, cool. That's, that's, that's really great advice for, for people.
I just wanna leave that out there-
Yeah, yeah, yeah
... for people. Um, for others who follow, uh, the work that or the career advice that I give, uh, the title topic of this is Pick Up What Others Put Down, uh, and specifically pick up, pick up what your mentors put down.
Like mentors always have more work to do than they have personally time for, uh, the, the high visibility mentors. And if you can, uh, show that you're a good collaborator with them, they will lift you up accordingly. And, uh, you know, that's, uh, that's a, that's a pretty good formula for career growth.
Mm.
Um, should I ask about Hyung Won or I, I, I don't know how close you are. It seems like-
Oh, we're still, we're still good fr-
Yeah
... good friends. Yeah.
Yeah. So the, again, like, you know, one thing that, one thing that you learned from Hyung Won.
Hyung Won is a great engineer, and he's very systematic in the way he thinks. Uh, I think Hyung Won is, uh, uh, uh... I mean, without, without going into detail too much, like I, I, I, I still spend a lot of time talking to Hyung Won e- even like in the-- e- even after we, we both in, are different places about like very interesting algor- algorithmic ways to think about life.
Like, you know, he would even think about things like... Okay, I, I should not like diverge too much about-
Okay
... per- personal stuff, but like I think he's, he's like talk, talk-- Like Hyung Won himself, I, I learned a lot about, about his way of thinking, uh, like more of like very interesting like, like perspectives on life rather than research.
But Hyung Won is a great, uh, engineer. And the, the, the one thing that scares me about Hyung Won is that like he doesn't, he, he, he doesn't have multiple monitors. He just quotes me one small screen, and he does everything with like very hyper optimized.
Mm.
Uh, and then back in-
This is like one of those U curves where like one screen, one screen, and then many screens.
Yeah, yeah, yeah.
From the top.
So, so I think Hyung Won, Hyung Won scares me because it, it's, it's like, I think that was at NeurIPS 2022, like we, we were doing some work at, at, at, uh, New Orleans, and then he would be like coding like perfectly fine with like this, you know, 13-inch MacBook with like one terminal, and then he would be like-- He keeps telling us like, "Okay, it's more optimal to like, like, like, like keep like using keybind-- like key- keyboard is more optimal than moving your head because if you can switch your screen fast enough, it's faster than your head, like being-- moving to different screens and stuff like that."
I, I, I did not like actually distill that because it's too painful to, to, to do that.
Yeah.
But like, I mean, he, he, he's, he's, he's, he's very, uh, uh, interesting in a way that like, uh, he belongs to one of those like hardcore people with like one monitor and like uh-
Maybe this is a relevant question be-- to, to just close out the Google side. Um, what do you think is a good programmer for, for AI research? Like, um-
Researcher DNA36:11
You mean like setup or like eating-
No, not, not setup. Just like-
... lifestyle?
Not, not even lifestyle. It's more about skills. Like what, what should people have? Uh, what do you interview for maybe, right?
Uh-
What do you see the, the high performers do differently than the, the less high performers?
I, I mean, okay, like generally there's like, I think like for, for AI researchers, like being a strong IC is like probably like the, the, the thing that I feel like, uh, is, is like important for, for AI researchers.
Like, like not, not r- like I think like, uh, you know, there, there are people who like, like there's certain level of like sacrifice to be like a, like a, like a, a AI engineer/AI researcher, especially if you're training like LNs because you cannot really be detached from, from ...
Like your, your, your jobs are could die on a Saturday at 4:00 AM, right? And then, uh, you, there, there are people who like would just leave it dead until like Monday morning and then, or like they, but there will be people who will crawl out of bed at 4:00 AM to restart the job or to check the, you know, TensorBoard or something like that, right?
Uh, I, I think like a lot of like being a successful AI researcher is like about like, uh, how like, how much you're willing to go to like ... A- and it needs to come naturally because you cannot be like, if you're not there, like you don't have like this like inductive bi- You don't, you're not like the, the kind of person.
But you cannot, if you force yourself to do this, you become miserable, right? Like, uh, I, I, I think a lot of it is about like, like, uh, uh, uh, I wanna say like passion is also like the entire thing, but it's more of like just the a, a, a kind of personality that, that, uh, like or like just the abi- that's maybe that's the ability of like if you're j- if something there's a bug at like 3:00 AM on like Saturday night or something, right?
And then you would like be like you, you couldn't go back to sleep unless you, you ... I, I'm not, I'm not s- this is very unhealthy by the way. Like people should not do this for, for, for, for, for, for, for a long time.
Uh, but I think it's, it's, it's like, uh, uh, and you know, I think these kind of things actually like, like, uh, allows people to make progress, uh, uh, faster, but it's unhealthy. So I, I'm also not even sure like what's like the, uh, I think, well, I don't care.
I do, do, that's on the record, I don't recommend this style of lifestyle. I don't want people to, to, to, uh ... But I, I think like a lot of people who, who are, are, are like, not, okay, not a lot, not everybody, like, but I just think this, this kind of, uh, attitude is like, uh, important to make progress.
I, I mean you c- you cannot be like checking out on like Friday, Saturday, Sunday and like work a, a nine to five if you want to like make progress or like some people are just so good at detaching like, okay, like, like, you know, like 8:00 PM I'm not going to ...
My job can die and then the, the chips can stay idle for like the whole night, but-
Yeah. This is-
... you know, I wanna watch Netflix, right? You can, you cannot, you cannot, like I, I, I think there's a level, like it's, it's like a sport, right? It's not like, like if you, you cannot win an Olympic gold if you want to like have like perf- like super ultra good work-life balance, right?
Yeah. I-
So I, I mean I just, I just think this is kind of like the, the-
Passion, intensity, dedication
... intens- yeah, intensity, right.
Yeah.
But, uh, I think, uh, the, the, the, the, the thing pe- like also need to know how to like kind of regulate and make sure that like people don't like die from this type of like-
Yeah
... not, not die per se, but like actually like burn out from this type of things. Yeah.
So, uh, those are really good, uh, personal qualities. Um, and just technical qualities wise, how much of the stack should people know, you know, if I-
Okay, so that was the question.
No, no, no. But that was important as well.
Okay.
Like it's, it's just harder to interview for because you really just see it on the job, you know?
Mm. I think stack is not like, not, not, stack is not that like because-
Like should I know CUDA kernels?
I don't know CUDA kernels.
Exactly right. Okay, good. Just for all, all you listening out there, you don't have-
No, no, but, but, but-
You don't have to feel like an imposter.
No, but, but you need to be willing to learn if you have to, I think.
Well, you haven't had to so far.
Yeah, I haven't had to so far, right. Uh, but-
So if I like Sling, Sling PyTorch, okay, great. Uh, you know, what kind of like, uh, do, do, do I know like distributed systems? Like do I know like what, what, um, what is the, what is the stack that you recommend for people that like, you know, get, gets you like a well-rounded end-to-end researcher?
I don't, I, I don't think there's any specific thing. In fact, I will try to be as like, like agnostic. Like, like I, I don't reg- like I don't really say like, "Okay, you need to learn JAX, you need to learn this."
By the time you finish learning, there's a new framework out anyway.
Yeah.
So, so it's more of like staying like constantly like trying to like, like being able to continuously learn and update like, uh, uh, uh. I, I don't think there's a single like, like single stack or like a single, uh, uh, single like workflow or single like, uh ...
Yeah, I don't think there's a single one. Yeah.
Got it. Cool.
Reka Launch41:10
Yeah.
Um, well that, that leads us to Reka.
Yep.
Uh, what's the founding story?
Oh, okay. Uh, so,
so I, I met some of my other co-founders while we were collaborating at, at DeepMind. I, I was at, at Brain and they were like at DeepMind. Uh, and then, uh, we wanted to, uh ... So I, I, I see myself as like a, a, I was not like, uh, I was, I'm not like a, a, a startup person.
I, I, I, I identify even today as a scientist and a researcher more than like a startup person, right? Uh, I think, uh, my, my co-founder Danny, uh, uh, started this story, right? And then, uh, f- this, this, uh, uh, Reka was like i- in the works from like late 2022.
I, I finally left in 2023. Uh, it was like, uh, I was like, uh, Danny kept asking me he, he wants to do something. Uh, do I want to go with him and do it? And, and it took, took a while f- like for me.
First I was like kind of the last co-founder to, to like, uh, to kind of, uh, form the, the, the, the-
Was the plan always for you to leave at some point and join him?
No, no.
He was, he was just like convincing you to join.
It was like, it was like a, an, an, an, six months more or less. In fact like, uh, I think more than six months period of like, like, uh, and I was like, like I always had this, uh, at the back of my mind for- Since like, what, August, like, uh, uh, I, I, I said no, like I, I, I didn't ...
Like, actually I didn't want to, to do it in the first place, but like, uh, but I think eventually, like in March, I, I felt that like, okay, it's time for me to experience something new. Uh, so I there's, there's like a, a like there's, there's a, a, a, uh, f- from my side the, the found like kind of like my leap of faith was more of like, I want to experience something new.
Uh, I've ... Okay, I've, I've like wrapped up this PaLM 2 work at Google and then like, uh, you know, and then more of like, okay, let me experience this new life and see where we can go with this.
Uh, so I think that was mainly like, like the, the, like all from my perspective, that was the story of like, uh, uh ... And I also, I know we, we, we don't have a, a, a lot of like, you know, I mean, I, I personally, we- I don't have a lot of like, like, oh, okay, like I, I, I, I ...
Okay, the funny thing was that like, like, uh, many, many years ago before my PhD, I wanted to do a startup actually at that point. And then over time I realized that like I was better off as a researcher, and I just forgot about the startup thing.
And it's quite funny that today I end up doing a bigger startup, right? But even until now, I, I actually don't, like yeah, as I said, I don't really ... I still kind of, uh, like iden- identify more as like a researcher and scientist and, and, and, and, uh, like, yeah.
So I, I think this is, this is mainly, uh, the ... It is, it's a, it's a very realistic, like down to earth, grounded founding story. Not- nothing too, too, uh, nothing too fancy. N- no, no app- like, no- nothing, uh, nothing fancy.
It's just, yeah.
Well, I mean, uh, it's not y- when you left Brain, like you already had a high profile coming out of Brain. You could have gone to any, uh, startup out there. They would all have wanted you, right? Um-
Yeah. Okay. Okay. Yeah.
So like why did you choose this one basically? Like, is it just 'cause of preexisting relationships? Uh, because it wasn't obvious to me, like, uh, you know, a lot of it, your other coworkers went to OpenAI, others went to, you know, like the, if you're, if you're fair, you went to Mistral, you know, that kind of stuff, right?
Like, um, Reka, no- Reka was like not on the, on the map.
I, I, I think it was to, for me it was a decision between staying at, at, at Google and like co-founding something. I, I didn't want to like, I didn't want to be, uh, like, uh, like it was more of the experience of like being a co-founder that like, was attracted me.
Mm.
Right? And, and wanting to experience that. I wouldn't have left like for Inflection or something like that. Like, I mean, Inflection is gone now.
RIP.
No, I-
They're, they're still alive. They're selling, they're selling themselves as a model foundry or some- something. Um, some- they, they like ... I, I don't know. They're, they're, they're a services company now.
Yeah. No, but I also think that like, like for example, like if you have to join like another like ... It, it, it would be like a very big tech experience again, right? I, I don't know. I felt like the, the experience I get is very complimentary to what I have.
Like, basically what, what I have I experience now is very complimentary to what I, uh, like, uh, that's the, the, the experience I had at Google, right? But if I were to join like something else, right, then I wouldn't have like ...
It would ... I, I would have just stayed at, at, at, at, at Google, to be honest, because to me it was very clear, like just two decisions that, that, that, that, uh, I didn't really con- Like, I was talking to a bunch of other startups, but I didn't really actually had the intention to like go.
Uh, uh, I was happy at Google actually, to be honest. Uh-
I'm sure.
Yeah.
I'm sure they, they make, they have a lot of things to keep you happy.
I, I was happy at Google, yeah, actually.
Um, so you describe yourself as GPU poor, but also you had $60 million to, to, to play with. Uh, you got a whole bunch of GPUs. Uh, I, I think you disclosed somewhere, but I don't remember the exact number.
Uh, and you had a, you had a, a good training run for Flash and then Coronage.
Mm.
Um, how would you tell the, the, the sort of the story, like people can read the technical report, but also like, uh, you know, what was that overall experience like? Uh, y- and I should also point people to the blog post that you wrote.
Um-
Damn. Uh, so
there were a lot of interesting things that happened along the way that like, like led to our ... So I think I left around like Ap- early April, like March, end of March, April and everything. Right? But most of our compute actually came in December, actually.
Yeah.
Uh, and there were delays.
Yeah.
So H100, there were major delays, right? So we were sitting around, right, bunched with like just-
And to be clear, you don't own the compute, you are renting.
Yeah, yeah, yeah.
Yeah.
So we, we, we, we, we were sitting around like w- with, you know, for, for a long period of time we had 500 A100s because we, we, we, we, we, we, we made a commitment like, uh, and, and they were constantly being delayed, I think because of H100 supply, demand, whatever, like, like reasons that, um, and it was also very hard to get like a lot of compute, like in one place, right?
Uh, and then we were locked in, uh, like, like, uh, for, for ... And, and we had to wait for, for the compute to come, right? So I think it, it was very painful because like, even when the compute came, it was mostly broken most of the time, and it was broken to a very bad extent that, that, that, that, that, uh, that ...
So, so it, it was actually, I, I, I, you know, I, I've ... B- before I, I le- I left Google, I was like, even I, I, even the early stage, I was very optimistic about like, okay, this compute translates to, to, to this amount of flops, to, to this model, right?
But I never expected the, the, the reliability s- to be so poor that, that it, it just threw off all the calculations about like, uh... And then we had to, you know, like, uh, uh, work like 10 times harder just to, just to ma- like make the thing, uh, go smoothly.
So I would say that like the, the, it was a, like bearable pain. I think the pain was like bearable, but like, it was just way, way more than, than, than, than, than, than, than, than expected. Uh-
Y- I think you did address this in your post, but, uh, the temptation would've been just to run everything on TPUs, which is the stack that you already know very well, uh, that, that works very well.
Uh, no, no, no. So, so, so TPUs outside Google and TPUs inside Google are probably very different things, I think.
Oh, how come?
Uh, okay. Firstly, it's like- ... infrastructure. Like, there was, there wasn't, like, a lot of, like, good code bases, like, outside Google that was, like, still-
Yeah, okay
... right? And, and, uh, the, the, the code base that I was most familiar with was, like, T5x. It was a JAX space. It would have been, like, by, by the time we wanted to consider it, it was already, like, deprecated, like, for nine months, right?
Uh, and then- ... TPUs, like, uh, I, I mean, I, I, I, I, I'm ... We, we weren't sure about, like, the, the ... I mean, the avail- availability of TPUs were, were also not great, great. Like-
Oh, I... My perception is it was a lot better.
Uh-
It's just that people have the learning curve.
Yeah, but we, at that point of time, we had our infra set up. We were training, already training models and, like, it would be so much cost to-
Yeah
... like, switch to TPUs. Uh, so I, I think TPUs, the experience of TPUs inside, outside Google, I have not actually run a single TPU job outside Google, by the way. Uh, but just, like, looking through documentation from what I see outside and from, like, like, how much I think that people inside Google don't care about what people think outside Google, uh, like, I kind of feel like, okay, we were a bit, like, like...
I, I didn't, I don't think we, we considered, uh, uh, uh, uh, l- I mean, not, not, like, forever not considering this, but, like, just, like, uh, at that point of time it was like, like-
The obvious choice is just stick to PyTorch
... no, no, let's stick to GPUs and, and, and PyTorch and make, like... Uh, I, I mean, it's not as if the chips we, we, we, we, we, we, we, we, we ordered were not there. They were there.
They're just not in the best shape.
Reliable.
Right. So, so yeah. So I, I think it was too much, like, uh, work to, to kind of migrate suddenly, uh, to TPUs. Yeah.
Yeah. For, uh, for those who haven't read the report, the, you had a very traumatic, uh, description about the chaotic and stable phases of, uh, various compute providers and, uh, I was just wincing when I was reading all those things.
Uh. Yeah, no, that, that was, like, a Three Body Problem reference-
Yeah
... the chaotic and stable phases. I mean, I was watching Three Body Problem at the time, and I just thought it would be ... It was, it was fun to-
This is a good reference.
So there was a lot of, like ... I think we had a lot of fun adding a lot of references and memes- ... into the, into the tech report. I think, like, you know, it, it, it goes to show, like, how fun the, the environment is w- within, within Reka, right?
Uh, we had a lot of fun with this. But, uh, I think chao- like, so, so I think chaotic and stable phase mostly is, like, we, we actually found that, like, uh, usually when, like, a provider, like, provisions new nodes or they would, like, uh, give us the-
Yeah, you don't want to be the first to use it.
Yeah. It's usually, like, like, bad. Like, like, dog shit, like, at the ti- like, at the start. Uh, and then, uh, it gets, like, better as you go through the process of, like, returning nodes, like, and, and you know, like, uh, uh, like, draining them, giving it back to them.
They will send it back for repairs and everything. Uh, like, and then, like, over time ... Because it's more of, like, it's, it's a more of, like, a numbers game, right? If there's one bad node, it kills the entire job, right?
So, like, the, the fact of the, the game became, like, just eliminating bad nodes from the, from the thing, right? And then, you know, uh, uh, I mean, just because of, maybe because of the supply issue or something, uh, when, when the deadline comes to ship this, this ...
For example, you, like, hypot- like, I just give rough numbers. Like, say you order 1,000 H100s, right? They will not be able to ... Usually they don't meet the demand of, like, 1,000 H100s at the date. They will give you, like, 500 first just not to piss you off, and then they'll give you, like, another 100.
Like, every over, like, two, three weeks they will just like, okay, I added, like, four nodes, added, like, like, eight nodes, that kind of thing. And then over time you reach, like, the capacity, like, that you ... Or you actually, maybe you never actually ever reach the capacity that you ordered for.
Uh, and then, like, as they add these nodes, right, like, sometimes these nodes are bad, and then they just kill entire training runs. And the, the thing which I feel that, I mean, like, for those, all, all those people, like, trying to sell GP- there are a lot of people trying to sell GPUs now, like resell, shell, package, whatever, GPUs, right now.
I think the most important thing that, like, that, that they are like, obviously they are like SLAs, all this in, in, in the contract and everything, and obviously, you know, you, you m- might be, like, like, entitled to something, something if, if something goes wrong, right?
But, like, the, the thing that, like, for g- like, large model training runs is that, like, one bad node kills the entire job, right? So should the compute provider be liable to pay for all the node wastage then?
No way.
No, it's, it's un- because it's unlikely, because otherwise-
It's unrealistic.
Yeah. Yeah. But-
No one would take that on.
No, no, no one would take that on, right. So I think that's also, like, a, a tricky thing. Who, who is taking the risk? Is the, the LM startup taking the risk, or is the compute provider ta- taking the risk, right?
I, I'm, I'm, I'm, I'm sure ... I, I think that the, the, I mean, this is my sense, I'm not, like, 100% sure, but I think, like, like, uh, uh, th- as, as, as there are more pro- providers trying to sell GPUs, we get all this inbound so much, so much about people trying to sell us GPUs, right?
The, the dif- the key differentiator is actually to find a way to, to balance the, the, the, the, the, the, the risk of node failure with, like-
Yeah
... like, as long as the provider, like, like ... I, I'm, I'm not, like, going to say 100%, but, like, if somebody can come and tell me that my nodes are so stable that I can share some cost with you if your no- job dies, this is, like, green flag.
Green flag, right? The moment they start to, "Ah, I cannot," like-
Do, do any of the big clouds do that?
No. I, I think as, as, as far as, as I know, no.
'Cause they have the, you know-
It's very hard to-
... the size to-
It's also very hard to-
... to guarantee that
... well, as far as I, like, to the best of my know- my knowledge, I actually don't, don't, don't know if any- anybody, like, like, like, like, does that. But, uh, I think, like, for anybody who is watching or if you do, like, a, like, a compute startup or anything, the, the biggest green flag would to be to, to share the cost of node failures with, uh, with, uh, your, your customers, right?
Because-
You mean the whole run?
No, no. Like, if the node-
Okay.
It's very hard to go, because you need soft- you need software to, like, you need, you need software to, to, like, like ... So let's say you, you run it for 12 hours, right?
Yeah.
It dies after 12 hours, right? You, you get 12 hours of, of throughput, right? But then you, you get, like, some wastage because of, like, uh, the, the, like, you know, the downtime and everything, right? Uh, you know, I, I think it would be fair to find some, like, middle ground to, to, to kind of split the cost-
Okay
... of the failures, right? And this bring back to my point about, like, like, like work-life balance. Because if the node- ... fail so, fail so badly, right, like, it, it actually, like, basically, right, your engineers cannot sleep at all You have babysitting rosters and everything, but you're living life with, like, constant anxiety because, uh, uh, e- even if...
Okay, even in the, in the case, right, where the n- the node failures are refunded, right, you still lose time. You lose three hours.
Sure.
You lose everything, right? So i- i- it's, it's a, it's a, a, I, I don't know how to, how to, how to, to, to, to, to, to go around this, but I think, uh, if there are a lot of compute, compute, like, compute providers like fi- like, like, like, like fighting over, like, the...
I think a good, uh, good, good thing to, to, to do is to figure out, like, this pain point, otherwise, uh, uh, or at least, you know, like, figure out some hot swapping, like, mechanism to, to, to, uh, uh...
But, but so far, most things we, we, we, we, we... Most of the providers that we tried don't, don't have this. They will also get confused when they, when, when you try to ask them like, "So my job is dead."
Like, can you pay for the full... Like, can you, like, like, refund for... Or at least they will get confused because, like, like, this is a, a LM specific thing that the, the large nodes, like, none of those-
They don't care about, yeah.
Yeah. They, they, they will get confused about, about this, right.
So s- current status quo is the LM startup pays for everything, right?
Mm, you c- you won't, you... Maybe you could gen-
Unless you can make some negotiations
... negotiate some, like, like, like refunds, but usually they are, they will not be so generous to like, like, like, uh, uh, pay for like the, the, the... Say you run 500 GPUs, right? If, if you break for four hours, then one node break for four hours, right?
They, they will, in their mind, they will be thinking, "I should refund you for one node." But in your mind you just think that actually they should refund you for, like, the full, full job, right? So-
So, okay, I, I need to... Everyone who is from my background is going to be asking this. How is it so fragile? Like, how is it so brittle? Like, w- what's your frequency of checkpointing?
Uh, so our, our, our, our checkpointing is kind of like we, we s- we see how stable the job is, and then we decide... Because checkpointing takes a, a... W- without a good file system, checkpointing takes actually quite long.
Uh, so it could be-
It's like a few hundred gigs, right?
Mm.
Max. But, but-
Yeah. I, I, I, I think so. I think so. I, I, I, I don't remember offhand, but, but-
That doesn't take that long
... but no, no. But sometimes i- uh, if your, if your file system is slow, right, your file IO is slow, your checkpointing could, for a 20B model, could be like, what, 30 minutes or something.
Okay.
I, I, I'm, I'm... Okay. I, I'm, I'm not, I don't know this by ha- by, by heart
Sure. Sure, sure. But it's not hours.
Uh, if you go larger, what, what if it's like a 200B model, right?
Sure. Okay. And still, like, y- okay, so, like, y- you should have some kind of ideal checkpointing to run ratio that is not catastrophic if you run into a node failure.
Yeah. No, so we, we, we see of it as like, like a MFU, like, because you can average out your, your flop utilization, and then you can see how many percent hit, like, how much slowdown.
Mm.
Right? So you probably go for something like if it's like you're taking off 1% off your speed, 2% off your speed, so basically it's actually fine to just checkpoint more, more, more, uh, regularly, right? Uh, that, yeah, so, so I, I, I think checkpointing, like, you, you will never also, like, fully...
Like, you also never fully, uh, uh, uh... You, there'll be like, you, you, you can get like from the clean slate, like nothing, right? You, if... As you optimize and, like, engineer, like, the, the system to automatically restart everything, you, you get some, like, of the time back, but you'll never be like, like, like perfect, perfect.
Like, so you still lose, lose, uh, uh, uh, stuff like that. If you checkpoint too often, like, what, every 30 minutes, then your file system is going to blow up, right? If you're going to checkpoint every, like, like, uh...
So I, for us, we just see as like how much average-
Storage is cheap compared to compute.
No, when your la- when model is, like, very, very large, your storage can, can, can easily-
Okay. All right
... uh, uh, blow up. So, uh, yeah, yeah, yeah. I think that there's still this pain point, uh-
Okay. Um, going on to the models. I, I feel like I, I digress so much about all these fun side things.
You like compute, right? You like, you, you like hardware and compute.
I love hardware and compute. Oh, I'm, and also I'm an orchestration guy.
Yeah.
Um, so one part of we... The question, one of the questions I s- I'm skipping right now is, um, you know, there's, uh, I came from Temporal. I'm familiar with Kubernetes. I've used Airflow. These are all the, like, the data eng cloud or cloud engineer type tools.
It, it's surprising to me that you guys don't have your set of orchestration tools that you, that it solved, right? You, you wrote, um, in your blog post, you had like the pain of multi-cluster setups and, like, to this, to the rest of us, it's, this is completely solved.
Okay.
I don't know if you know that.
No, I, I, I don't, I don't think... Like, so, so we, we, we, we, we use Kubernetes for, for a bunch of stuff, but, like, I think, like, for experimentation and, like, stuff like this, it's still not fully...
Like, we, we, we didn't have, like, the time to actually like, like, like, build something that is like-
It should exist in open source. Someone should have done this.
Okay. Okay.
It's, uh... No, I'm not s-... It is what it is, but I'm surprised. That's all.
Okay. Okay.
'Cause it seems like a valuable problem, um, and someone should do it.
Okay, okay, okay. Yeah, yeah, yeah. Good to know. Good to know.
Um, okay. So, uh, Reka Flash Core Edge, um, you know, congrats on beating a whole bunch of, um, of state-of-the-art models, uh, especially much bigger than, than, than each. Uh, people can see the papers for all the other stuff.
Was this your expectation from the start, that you would basically definitely, uh, be frontier? Like, how do you s- how do you, like, from the start of, like, you, you haven't trained anything yet, and you're about to kick off the runs.
Like, how-- are you able to s- to, like, call your shots and say, "We will be GP- we will beat GP- 3.5"?
Uh, nobody can predict the future.
Okay.
No. Uh-
Because it will-
How much confidence? Okay. We were confident. Like, we were confident.
How? Yeah. W- why?
Uh, I, I don't... So, uh, I, I think we, with, with, with, with, uh... Like, okay, how, how, how, right? It's a good question. Uh-
'Cause it would be, it'd be a shame to do a whole bunch of work and then end up sort of middle of the pack, which a lot of people end up, right?
Uh, I, I, I, I... We, we were confident. I think that we, uh... A, a lot of it was like YOLO. I mean, I'm, I'm, I mentioned in, in, in, in, in, in the thing. I think we would, like, require a lot less iteration than, than...
Just because of our prior experience in, in, like, training these models. Like, so I was confident in, in, in myself About like, uh, like the, our models will turn out to be, to be, to be, to be good.
Uh, and like I, I, I... about ex- exactly how I actually don't really like pinpoint to a particular reason of like... I mean, we, we de-risk stuff, right? We de-risk stuff. So, so the, the, a lot of, a lot of part of it is like, like de-risking and like, okay, you run like 4B es- e- evaluations and you can see, okay, this is like my sp- sp...
If, if you run 4B and your loss is like going crazy, you know that, okay, this is gonna be a shit model, right?
Mm-hmm.
But I think it's like we train enough like, okay, we don't have a lot of compute to do a lot of evaluations, but we did s- enough experiments to know that, ah, okay, our, our, our infrastructure and our, uh, like everything is set up to be, be, be, be good, right?
Obviously, you know, uh, the field moves, right? So the, the what- whatever we, the, the, the field moves. So, so I, I wouldn't say that everything was like smooth, like the first time round it's like smooth and everything, but I, I think we were confident in our ability to like make the least...
Like we're not like really, like, uh, we're more confident about like, uh, the ability to like move with as little steps as possible to the goal, more so than like we are more confident about this ability, more so than like my model is going to be this like level at, at, at, at this time, you know what I mean?
It's more of like, uh, you know, like for example, we, we, we, we, let's say we, we run the first round of human evaluations, right? Uh, and then we see our number as this, right? And then we are confident that in five more tries we will get to this, you know, kind of like get, get to like, like, like, like this.
It's more of the like that kind of confidence rather than actually like, like, uh, uh, uh, you know, the, the... You know, it's also a little bit of like, you know, you see a new leaderboard, hypothetically, like in acade- like in, if in, as, as a researcher, you see a re- release, a, a, a new le- le- leaderboard, right?
Uh, you, you ex- you, uh, you approach it like a puzzle. You don't know like whether you... At the start of it, you might not have the answer to the puzzle, but if you're good at solving puzzles, like generally, right, you know that within one hour I'll be able to solve it.
You know, that kind of confidence.
Yeah.
Like, it, it's like, you know, it's the ability to, to hill climb or the ability to, to improve up over arbitrary things, right? Rather than... I think we were confident more about that rather than like, like, uh, like, uh, you know, I mean the, the, everything is different, right?
The stack is different, the, the, the infrastructure is different. The data is also different from what, what, I mean, we have a lot-
Which you haven't talked about, right? It's just 5 trillion tokens.
We, we, we have a lot of, yeah, we have a lot of ex- experience from prior, like our jobs, but like, it's not gonna be the... Like we, we don't have actually exact, like exactly the s- the, the, the same thing because, you know, like dif- different companies have different stacks, different everything, right?
Uh, so it's more about de-risking, uh, uh, uh, being confident in like solving the general problem of like improving over things. Uh, which is why also I think that the team is valuable in the sense that we are not like valued by our mo- our, our model itself, but we are just valued about like, like how we can see one problem and we can just like so- like solve it like s- super quickly.
Yeah.
Right? And that's what we are confident about, right? It's more of like, like, uh, than actually like the artifact itself. Uh-
Uh, you mentioned that, mentioning your team, you said, uh, at the largest your s- your team was three to five people on the pre-training side. Uh, it, it, was that the team that you recruited? Was, was it all your ex-colleagues?
How did you, how do you find people that, uh, you know, would have this kind of solid intuition?
Uh, so I think that like, uh, some of the, the people in our team were like, I worked with them at, at Google as colleagues and stuff, right? Some of them were like fresh hires, like they were like, uh, fresh PhDs or like, uh, and everything.
Uh, I think that, uh, uh, everybody helped out and worked like quite, like they, they, they did what they, they were, they, they, they, they, they were like the best at. Uh, and like, uh, I, I think, uh,
y- yeah, I, I, I think we, we, yeah.
Architecture1:05:02
Okay.
I, I don't know whether I answered the question, but yeah.
I'm, I'm-
I was getting like a-
I'm always looking for like, you know-
I was getting like a-
How do people get hired at Reka or like what, if other companies are looking to hire like you have hired, I think you've hired successfully well, you know, you small, small team that with impactful results, what should they be thinking about when hiring, right?
So these are useful takeaways for people if they're listening in. Um, but if you don't have any, if, if it's all vibes, it's okay. It's vibes.
Vi- yeah. Okay. Good vibes only. Good vibes.
I, I, I understand. I understand. Uh, okay, so, uh, I do wanna comment on Noam architecture.
Okay.
Uh, so Swig- uh, if you wanna like, people have variants of all these, SwigLU, GQA, Rope, RMSNorm, um, and then obviously the big one is encoder-decoder versus decoder. Uh, could you comment on each of those? Like, uh, were you just like, "We're confident that Noam got it right," or did you actually do a evaluation of each of your architecture choices?
Oh, I, I mean, like, okay, ar- architecture-wise, it's something that I feel like I'm ea- easily able to like I've run so many architecture ex- experiments- ... that like, you, you know, like, as in like I look at architecture and like, okay.
I, I, I don't want to be like overly like, but it's, it's, it's, it's, it's like I, I, I think it's very hard to outperform the-
OG Noam
... the, the OG Noam tr- uh, uh, uh-
Why? It can... I mean, on the surface of it, like we have to have learned something in the last seven years.
No, all the changes, all the changes that, like, like s- like SwigLU was this like, okay, SwigLU is like probably one of my favorite papers of all time, just because of the, the divine bene- benevolence, like that Noam actually wrote like, like, uh, we owe this success to divine benevolence.
Oh, yeah, yeah.
Like that was like a m-
Yeah
... it's always a, a meme thing, right? Uh, and, and like, okay, so like GQA, uh, uh, MQA was always like the, the multi-crit was always like a, a, a, a big, uh, controversial thing because MQA usually you get a hit because it's MQA and everything.
So people kind of know that like it was a very-
Hit in what? Hit in performance?
Like hit or miss. Like it was like you could, you could get a hit in the performance from MQA, like m- like, like, uh, uh, MQA alone. MQA was always like- You know, a, a, the, the choice, right?
It's always like, okay, should we use MQA, should we not use MQA, right? G- when GQA came in, right, it became like a no-brainer to use GQA because you don't get the hit anymore, and then you just get the fast, like, inference benefits of GQA, right?
So I think GQA, uh, I mean-
Which Llama 3 now the most, yeah
... yeah, yeah, yeah. So, so I think Llama 2 already. I'm not very sure. I don't know
Llama 2, uh, the 70, 70B
... G- G- G- GQA, right?
Yeah.
But, uh, I mean, the reason why we call it Noam architecture because, like, MQA came from Tome and, and GQA was, like, a follow-up paper by some of my colleagues at, at, at, at, at Google, right? So I think GQA was became a point where, okay, this is already accepted.
Like, i- it's good en- like, it's a no-brainer to use GQA. SwiGLU was an interesting thing because there was a very period, very long period of time. Also SwiGLU was a single author paper by Noam, and very few papers was-- like SwiGLU had very few citations, like, at, at the start because it was, like, a very, like, it was very-- it was obscure.
Like, only Google papers were citing SwiGLU at one time and, and a lot of them was like-- like I was ci- like at one point I was, like, probably, like, like 30% of SwiGLU citations. 'Cause every time, like, like SwiGLU became popular because of the, the, the updated T5, the T5, uh, 1.1 that uses SwiGLU, right?
And nobody actually really cared about SwiGLU for a long time, 'cause I was checking why is this, like, underrated paper, like, like, not getting much citations.
Mm-hmm.
And then I think probably now it has, like, a few hundred citations by now. Uh, but I think S-Sw-SwiGLU is one of the things that like, that, that, uh, uh, uh, you know, I, I played around with a lot, like, at, at, at Google.
So SwiGLU really works. Uh, there was also, there was also a, a, a paper we wrote about like do transformer modifications, blah, blah, blah. Like it was a paper with Noam and, and Sharan and Hongwan and stuff like that.
And then we ablated like so many transformer v-variants. Um-
Yes. Yeah, I, I saw that.
And, and-
Uh, some of them matter, but most of them don't
... most of them don't. And then the only thing that matter in, in that two paper was, in that paper was, uh, SwiGLU. I, I forgot which exact SwiGLU variant was it, but, and sparsity, uh, at that time, right?
So, so that was strong enough like to finding, uh, to, uh, uh... Right. So I think SwiGLU is one thing that, that really works, uh-
For, for the listeners, this is the, uh, inductive bias, uh, scaling laws versus model architecture is how this inductive bias influence-
No, no, no, not this one. There was-
Okay. This one. Okay
... another one, like, do transformer modifications something, something, something.
Okay. Got it.
Uh, uh, I think the-- I forgot. The f- yeah, first author, author was Sharan, I think.
Right.
Sharan Arang, uh-
You, you gave the keywords-
Yeah, yeah
... so people can find it.
And then, uh, yeah. So I, I think, I think-
So RoPE and RMSNorm are, are left
... like I think the RM- RMSNorm RoPE thing, like-
Not controversial
... like is, is, is not like con-- like, like, uh, obviously, I think RoPE is probably like it has that extrapolation thing, which is nice. Uh, and then like, like it's also like default now. Nobody wants to add positional embeddings anymore, right?
Uh, and I think, I mean, I like the T5 style relative attention for a bit, but like I think, okay, RoPE is, uh, uh, I actually wrote-- ran that ablation for Palm, like the T5 relative, uh, attention versus, uh, RoPE, uh, uh, and, and, and, and stuff.
Like I think RoPE is similar to other things, but it has this extrapolation thing, which is nice and like, uh, and, and you know, I think it's just better-
Which, which is why your long context version can go to 256. Okay.
Uh, this for all, most of the long context models-
Yeah
... they, they use the RoPE, RoPE extrapolation thing, which is nice property, right? Uh, so, so that, that was for RoPE. Uh, I think, uh, uh, there were also like some things like the, the layer norm, like positions and stuff like that, uh, that, that were like, uh, you know, like it mattered a little bit, maybe not too much and everything.
But I, I think in general there, there was not a lot of like-- there are not a lot of things that people could do to the transformer to, to... It's been like f-four, five years, right? And then the, the-
It's amazing
... the, the, the, the vanilla transformer, I think if you use it as it is today, will not be like that optimal. But like the, the, the, the transformer that we slowly evolve to now is like, like the Noam transformer is probably like very, very, very strong baseline that is very hard to like-- like I, I don't even think that anything like-- I don't even think that anything that, uh, uh, like, uh, I don't even think that like...
I think we need a drastic shift to, to beat that, right, rather than-
The state-of-case model type things
... like or you could find like a, a more like, like, uh, s- like SwiGLU is a small change, right? You could find like a, like some small change that are like, that are like, uh, a big enough impact, like widely like that, that don't cost a lot of like...
'Cause like a lot of architecture ch-changes, right, the moment they are like tedious to implement, like nobody-- like SwiGLU is a simple thing, right? This PD and then get the thing. It is a very simple thing to implement.
Maybe that's why it's caught on because it has like a, a, a, a, a additional boost that's for the simplicity of it, right?
Mm-hmm.
So there's also like a bit of like implementation lottery, if you will, right? A little bit of like if you, if you propose like some very complicated thing for like-
Yeah
... for like 0.1%-
Easy as PyTorch.
Yeah. Nobody will use that, right? So, uh-
Uh, the biggest, biggest, I mean, I, I can't, can't believe we have-- we're taking so long to come to this topic, but the biggest Noam architecture decision is encoder-decoder versus decoder only.
No. So encoder-decoder is not like a Noam, Noam-- the Noam architecture is more like the-
The, okay. Maybe-
... the, the, the-
... like more, more old school transformers. Like, I don't know. But how, how-- so just maybe you want to just talk about the decision on encoder-decoder versus decoder only.
Uh, so I-- okay, I wouldn't be able to comment about like exactly our, our setup, but like I think encoder-decoder are kind of very misunderstood from, like are, are kind of very misunderstood thing, right? So there's encoder-decoder, there's a non-causal decoder, which is a prefix LM, and then there's a decoder only model, right?
Technically, a causal decoder and a non-causal decoder are very similar in the sense that it's just a bidirectional mask, right?
Mm-hmm.
And then a prefix LM and encoder-decoder has only, uh, the, the only difference is that encoder-decoder, uh, splits the inputs and targets into different, uh, uh, non-shared Transformer stacks, and then, uh, there's a, like, there's encoder bo-bottleneck in the end, right?
So te-technically pe-people al-like kind of always associate, like, encoder-decoders with, like, like, like BERT or, like something like... Like, you know what, people get confused about these things, right? But I think in the UL2 paper we really, like, kind of, uh, explored this, and also, like, maybe some of the big, big science papers they also talk about this, right?
Is that, uh, prefix LM and, and causal decoders are very similar, just the mask. Prefix LM and, and encoder-decoder are actually also quite similar.
Mm-hmm.
At the end of the day, they're all autoregressive transformers. There's actually, like, really, like... The, the only big benefit of encoder-decoders is that it has this thing called, like, I mean, what, what I like to call it, intrinsic sparsity.
Okay.
So basically, a, a encoder-decoder with, like, N params is, like, uh, like, basically if, if, if, if it's like, it, it has the cost of, like, a N over two decoder model. So it, it's a bit like a sparse model because you actually spend the same amount of flops.
It's just that you have two sets of parameters, like, for encoder-
Right
... and decoder, right? So it's actually f-flop match with a decoder model of, like, half the, the, the, the-
Parameters
... the parameters. So like a, like a U- like UL220B is actually about a 10B decoder only model, right? So you get free sparsity, like free, free, free, free sparsity from that. It's, it's something that... Okay, the, the, the, the OG TP paper talks about this.
You can, you can look at it. There's this complexity chart. I, I di- I didn't like, like, like, like cover this, but when doing the UL2 paper I kind of like was mind blown by like, like, "Wow, encoder-decoder is so much more, um, uh, uh, uh, ef-"
Expressive.
No, not expressive. It's so much more, uh, uh, powerful compared to decoder model on the same flop match, right? There, there's a table in the o- in the OG TP paper. This was 2019, actually.
Yeah.
There, there, there was like, uh... So I, I think there, there actually isn't really much to... The, the, the only thing about the encode-decode architecture is that it's, it has, uh, it provides like a 2X, uh, uh, intrinsic spasi- like free sparsity, right?
But then the question is that if you go to MoE, there's this, there's this still, there's this still hole because actually MoE also kind of... You, it's like the flop parameter ratio that you kind of, like you, you kind of change the s- the, the, the, the, the, the, the, the, the, the, the flop parameter ratio of like, like a, the...
And then, like, encode-decoder is like a 2X or that. So it, it, it's just like that. The difference in architecture is just that. It's not, it's not that complicated. Like, people don't need to overthink this, right? The other thing, though, is the objective function of the, the, the, the, the people always asso-associate-
Architecture
... encoder-decoder with the thing, right?
Yes.
It's not the same thing. You can train encoder-decoder with regular language modeling, and you will... A-actually, to be honest, like, a lot of the retrieval augmented language models can also be seen as some form of encode-decode because you have the, the retrieve documents as, like, the encoder.
They could get compressed and-
Oh, okay
... they, they, they actually not very-
They're not in the model, but you can insert-
Yeah, yeah
... them in a context.
No, it's, it's not actually that... I, I mean, people are kind of overthinking this like, like encoder-decoder, like decoder only thing, right? They are actually, at the end of the day, like autoregressive models.
So, so the context becomes the encoding s- element. Yeah.
Yeah. It's also, uh, the, the, that's how you, you, you think about, like, encoder, like for, like for, you could, like for example, a decoder only model, right? You have the prompt, like there's inputs and targets. Like, you just think of it as like in- like, like the target is like generation input.
It's like the prompt, like what context you can retrieve documents, whatever you just, like, p-put that in, right? Uh, you could also put that inputs into the decoder instead, and then you just continue generating from decoder. Or you could just put, like, the inputs into the decoder, but then you put like some extra not so information, not so important information into the encoder.
The advantage of this though is that by splitting into encoder-decoder, your encoder can actually, uh, do some a little bit more funky stuff. Like, because you, you don't, you, you are not bounded by, by, uh, you're not bounded by, uh, the causal mask anymore.
A lot of the efficient transformers, like a lot of the sparse, the, the, the sparse transformers, like, I mean, the ear-early days, there's like, I don't know, like Linformer and like whatever things like this. They, they cannot maintain the causal mask, and that's why like, like the, like, uh, you cannot train a language, like proper language model with this, right?
But with like if you separate out your very long context into encoder, this encoder has no loss, right? You could just do like aggressive pooling. You could do some cr-crazy sparse attention that has like, that, that, that is like, you know, like final transformer, something like that, right?
And then you could make that smaller than the decoder. You could make that faster than the, the, the decoder. You could also do like... Uh, so I mean, that are just some of the advantages of like, like, uh, like why like splitting into encoder-decoder is actually like, m- uh, could be beneficial to like, uh, just using like a, a decoder only, uh, model.
But, uh, but fundamentally, I mean, like i-it's, it's a, it's a, at the end of the day, the, the, the, the decoder in encode-decoder is a language model.
Will, will cost, yeah.
It's still a, it's still a regular autoregressive language model. So there's actually like, I mean, it's not, not that much different from like a retrieval augmented language model that you pass, like retrieval-
This is, this is news to me. I don't know if you've ever, ever expressed this, but yeah, this is actually makes sense.
Okay. Okay. Yeah, yeah.
Uh, I, I don't know, unfortunately I don't know enough to, to push back on this, but, uh, it, it ki- it, like, on the surface of it, it seems to make sense. Would you make the same choices if you were not focused, so focused on multimodality?
'Cause like, you know, that's one of the ways in which I was thinking like, "Oh, encoder-decoder makes sense, that it's more natively multimodal."
Uh, yeah, I was, I would, I just have to say that it's, it's relate- it's, it's, it's also-
It's relevant
... relevant. Yeah, it's relevant. Yeah.
Uh, yeah, i-i-it wasn't that ob- uh, that obvious to me. And I don't know if you want to compare your approach versus Adept's approach, uh, because they published some things on FuYu. Um, I, uh, I, I don't know if you consider them competition or, or not, but like, obviously they're also trying to push the, uh, the, this, this kind of similar models that you're, you're also releasing.
In, in s- in the sense of like small, medium, large multimodal models
No, I, I'm, I'm, I'm thinking whether I should say something about, about
It might be a hot take.
Uh, there's a l- uh, so we, we compare with FuYu-8B, the release one.
Y- yeah. Um, you know, yes, they, they, they maybe don't, don't do as well on the benchmarks or whatever, but I'm just thinking about the architecture choices, um, because a lot of people are commenting on, um, FuYu and Fu-
Oh, okay. I, I, I think we, we were, we were-
Okay
... not, not confident-- uh, not, not, not comfortable talking more about this. Yeah.
Yeah. 'Cause the-- yeah, the, the vision encoding was interesting. Uh, okay. Uh, anything else we should talk about Reka that we haven't covered?
Uh...
And if you wanna drop hot, uh, hot news, we can embargo until the news is public.
No, no, no, there's nothing. Yeah, yeah, we can move on. Yeah.
Cool. Um, then we can move on to broader trends in LLMs, just commentary on just like-
Yeah
... ecosystem stuff, like completely, uh, uh, independent from Reka. Um, you commented on a few things like Llama 1 to 3 glowed up a lot. I, I call this the Llama 1 to 3 glow up. Like they improved into like an actual top tier, uh, open source model.
Yeah.
Um, Phi-1 was, was, uh, had a lot of criticism, but it seems like Phi-3 is getting a lot of love. Uh, do, do you just generally see like in your open model tier list, like, uh, what, what's going-
Oh, damn.
What's, what's going up and down?
Uh, so I think, I think, uh, Llama 1 and Llama 2 are like quite mid, right? But Llama 3 actually got good, right? I think Llama 3 is actually strong, right? Uh, I don't really follow Phi much, just that like I just don't follow, f-f-like follow like, like, like-
Their whole thesis is the textbooks is all you need thing, right? Like that we can-- well, we can use way less data than everyone else and still-
But I think you cannot cheat the scaling laws, right? Because like you-- I, I remember seeing like vaguely saying that like, like, oh, they match like Mixtra 8 by 22 or like something like that on like some... Okay, I, I don't think these academic benchmarks are like that meaningful anymore, right?
So but then like, then when you go-- they go on LMC, they get like what, 47, and then they get like maybe just like sim-- slightly maybe it's like-
That was Phi-2. I don't know about Phi-3.
Oh, there's Phi 3? No, I think-
Phi-3 was just released like yesterday.
Oh, I don't even-- I didn't even-
Yeah.
Yeah. But, but I, I, I don't know. Uh, I, I think there's, there's, there's some, uh, like I, I, I don't follow Phi that much, but I, I don't... I think that like, like synthe- like a model that is synthetically...
I, actually I don't even know this where, where like I didn't even read the paper. But I think that like, uh, a model that is like, like based on the premise of like, like distilling and stuffing, something like that is like not that interesting to me.
Mm.
Uh, I, I-- okay, like, you know, like, yeah. So I think I don't really follow like Phi much, but I think that like Llama 3 actually shows that like, like kind of like Meta got a pretty like a good stack around training these models.
Uh, you know, like, oh, and I've even started to feel like, oh, they actually, uh, you know, kind of maybe caught up to Google now, right? They're kind of feeling, uh... That's also maybe a hot take itself. But, uh, but yeah, I mean, Phi, Phi I don't really like-- I don't really kind of, uh, uh, uh, follow it that much.
And, uh, yeah, I, I just, yeah, I mean, there's too much, too much things to, to follow. So I think it's like I, I, I think like Llama 3 is probably like the most-- the, the, the first most legit open source.
When you say these kinds of things, like most legit, um, it, it-- obviously there's some of this vibes eval or whatever.
Yeah.
Um, but like I feel like a lot, um, people-- the, the very common feeling is MML is kind of saturated.
Yeah.
Evals1:22:57
So like what do you look at now? Is it just LMSYS?
Okay, so I-- okay, like I think that LMSYS has its problems also.
Yeah.
So LMSYS is not like exactly like, like, uh, I mean, it's, it's probably better, better than all these regular, uh, benchmarks, right? But I think like, uh, serious LLM labs create their own evals and, and they-- a good eval set is one that you don't release.
Mm-hmm. Yeah.
Right? A good eval set is the one that you like-- okay, you release some of it, but like, it's like you, you don't like, you know, let like-- let it be contaminated by the, the, the, the, the, the community.
Uh, so I think like, uh, yeah, I think LMSYS is probably the most legit one, like out of all the... I mean, like, you know, the things like GSM8K human eval, they are-- the coding human, they're all like-
Contaminated
... like not, not-- I wouldn't say-- they, they're all like saturated, contaminated. No, like, you know, GSM8K whether you're 92, 91, like no one cares, right, kind of thing, right? Uh-
But we still report three decimal places in all of our reports.
Yeah, yeah, yeah. But it, it's kind of like almost like this like, uh, uh, uh, obligatory thing to do. You have a table of, of, uh, of, uh, uh-
Numbers of, uh, your, your thing at the bottom.
It's, it's, it's, it's interesting to see how, how, how the, the field evolves, uh, also over, over time for, for, for this type of like benchmarks. But I think evals are gonna be important and it's on the-- actually interestingly, it's on probably, probably on the academics to, to, to, to, to set the correct like, uh, uh, set the correct, uh...
I mean, they, they, they, they have like they, they've been-- academics have always been like, like, "Oh, we have no computer," but like, okay, this is your chance to like steer the field in the right direction, right?
Yeah.
Uh, and then, yeah.
I think the, the, the challenge is getting attention. Um, so you know, now MMLU, you know, is reaching its end of its life. Like what, what is next, right? There's MMU or there's MMLU Hard, which, uh, someone recently released.
It's Pro, MMLU Pro, I think.
Pro?
Yeah, it's called MMLU Pro.
Oh, yeah, that's right. That's right.
MMLU Pro.
Um, but, but like that only lasts you like a year, right? And then you have to find something else. So I, I don't really know what, what is that. Well, so one, one thing, you know, you had a, you had a comment I think in your Reka paper about, uh, there's two types of evals.
Or this is a Vibe eval paper. Uh, one is, uh, LMS as judge, uh, and then two is Arena, is Arena style, right? That's as sort of the, the two way-- two ways forwards for just general evals that cannot be gamed.
Although there's also- Uh, there's also like human evals that you-- like instead of LLM as a judge, there's also like human evals that-
Human evaluations
... that, that you run. Like that's kind of similar to arena, but kind of different to some extent also. Different in the sense that like-
By the way, do you, do you use your own staff to do that or do you like hire an outsourcing firm?
No, we don't. We, we, we, we, we have a like-- we work with third-party data companies to like-
Okay
... there, there are a bunch of these like around, right? But like obviously we don't like eval them ourselves. Like
I don't know. Like, I don't know how much, how m- how many evals you, you wanna do, right? Like, uh, uh, I do think Andrey Kopytov mentioned that sometimes you-- like the best researchers do their own evals.
Yeah.
Just to-
Looking at the outputs and, and stuff-
Yeah
... is something that, that like, uh, researchers should, should, should do.
Um, well, th-there is one element of parametric evals, which I, I'm hoping that more people can, uh, come up with, wh-where like you kind of, you, you, you generate a f- the, the eval's kind of like a formula.
Sorry, the benchmark is formul- is generated from a seed, let's say, let's say. And you can, you can, you can withhold the seed or like you can vary the seed, like you can report, um, how, how your model did on the benchmark, given it-- given a certain set of seeds or whatever, and you can maybe average them.
But in that way, it, it becomes harder, much harder to, uh, to contaminate. I, I wonder if that is, uh, possible.
Wait, do you have like a what, like what, what, uh, do you have like is there like-
An example of this? Uh, not specifically. This is just something I'm wondering for myself. But, uh, I did-- someone did recently put out GSM 1K, uh, which was, uh-
Oh, the scale thing or the-
I think, I think... Is it scale AI?
Yeah, yeah.
Yeah. Uh, which this is s-sim-similar in that respect. Like put-- uh, make it easy to make variations of a, of a one little benchmark, but like it-- that is more likely to be withheld from, uh, from training data.
Seems like a-
Yeah, yeah, yeah
... possible thing.
But eventually those would, would like... So it's, it's always this like-- like even we put out like eval, we also are quite, are quite like upfront with like if the, the more people use it, there's a lifetime. It's like a car, right?
After you drive, run, run, drive, run a certain miles, it's, it's, it's, it's, there's-- it's time to sh-shelve it, right?
Yeah.
So I, I don't think there's like a actually like a, like a good solution to, uh-
Yeah
... in general, I'm also like a bit like, like, uh, I, I mean, I, I, I think this is like important for, for, for the community to, to, to, to, to think about, right? But like, is it, is it like a fundamental limitation that any benchmark that goes out...
Like also there's also one thing is that in the past people used to like withhold test set, right? Like squat or something. They used to withhold test set. But then like the, like after a while, I think people also realized that like when you withhold, like MMML-
Kaggle, Kaggle matching
... no, like when you with-withhold the, it's like so, so much extra work for like the, the, the community to like eval on, on this, that they just don't do that, right? It's either your data set become-- your, your benchmark becomes unpopular or I think it's also incentive things, right?
So if you-- let's say you are, you, you want to run like a contest, right? And then your goal as an academic is to get as much citations as possible on this benchmark paper, right? Like then you or, or like this, this-- you want to be as, as famous as possible.
You will not want to withhold the test set because if you withhold the test set and then people have... Like there was once like i-i-in, in, in, I mean, like many years ago, there were even some benchmarks where you had to like, like package your model and send it to them to run.
Like, and, and, and this, like, these benchmarks never, ever like, like never, ever like-
Took off
... like took off just because like... So at the end of the day, right, it's like the, it's the root problem, like incentives. Like it's the-- also the, the benchmark, the benchmarking problem is also like an incentive problem, right?
So like, it's also like, like, like people want to show their model is the best, and then the, the game masters want to, want to, to gain as much clout as possible. And I think also LMsys also get into, caught into some-
Trouble
... I, I don't have a, I don't have a take on this, but like there's, there's, there's like people who also feel that they are also optimizing for hype, right? Their own clout, right?
Definitely.
So, so there's all this-- I, I think there's a lot of interesting like-- I don't know what field this would be, but I don't know, social logic. I, I don't know, like-
Yeah.
Like, like I, I think there's a lot of papers to be written, right? I mean, about, about, about how these incentives, like rewards and incentives like kind of, uh... But it, it might not be, it might not be, it might not be, be, be, be solved.
So, uh, yeah, I don't know. This-
I would say SWE-bench is probably the, the, the one that's kind of broken out this year as like, uh, now a, a thing that everyone wants to compete on is if you're a coding agent.
Mm. I see.
Uh, I don't know if you have a view on it, but it's just like you have-- it should, it should be known to be hard, and it should be, uh, you should be able to make progress on it quickly.
That makes you popular and cited a lot.
Yeah, yeah, yeah, yeah, yeah.
Um, okay. Multimodality versus omnimodality. Uh, so this is a little bit of commentary on GPT-4o and Chameleon. Um, uh, I don't know if you saw the M-Me- uh, Chameleon paper from Meta.
Briefly saw it. Uh, yeah. I, I'm not-- I didn't really take a look.
Basically, the general idea is that most multimodal models, uh, like are, like LLaVA or Flamingo, which are late fusion, uh, which is you freeze, freeze, and then you sort of join together, um, versus early fusion, where you do it properly, where like everything is, uh, all the modalities are present in, in the, in the early pre-trained stage.
And it seems like, uh, things are trending from late fusion to early fusion is, is the general thesis with GPT-4o being very obviously early fusion. Uh, you guys, I, I would class it as early fusion. Um, I, I don't know if you have commentary on whether this is obvious to you or this is the, this is the way, or they will just be-- they will coexist, um, anything like that.
I think whenever possible, like early fusion is better. Uh, but like I think there will still be a lot of works that do late fusion just because of like it's a-
GPU poor.
No, no, not, not GPU. Okay, part-pa-partially, right. I see this as like an art-- as an artifact of the, the, the, the, the line between language research, researchers and vision researchers and more of like, okay, like people who are training language models, they put out like a Llama or whatever, and then somebody takes it and then do late fusion on top of it.
It's more like a Like, i- it's just, you, you eventually everything-
It's Conway's law. It's shipping the org chart.
Uh, yeah, yeah. Yeah, yeah, I think so. I, I don't know what, what law was it?
Conway's law.
Okay. I didn't, I didn't, I didn't know about that. But it's, it's kind of like an artifact of the org-- the, the, the organization kind of thing. Right? Like-
No, it's just because people don't have money to train things from scratch. I don't know.
No, no. I mean, e- even in big, even in big companies, right?
Okay.
Like, in-- I mean, I, I don't know how things have evolved in ma- many companies, but, like-
You're talking about Flamingo?
Like, language and vision teams don't used to be the same thing, right?
Yeah, yeah.
Uh, so I think this is like a artifact of, of this. But as early fusion models get more traction, I think the, the, the teams will start to get more and more, um, uh... Like, uh... I, it's, it's, it's a bit like how, how all the tasks, like, unify.
Mm-hmm.
Like, from 29-- 2019 to, like, now, it's like all the tasks are unifying. Now it's like all the modality is unifying. Uh, and then I think, like, eventually everything move towards, like, early fusion.
Yeah. Something I don't understand is, uh, and I don't know, you know, if... Feel free to pass on this if you're not confident, but, um, tokenization of images to the same latent space as, uh, as the, as the language stuff.
Like, I feel like early f- like, I-- Is there a paper that I should read on, like, how this is done?
Oh, then I should pass on this. I, I-
Okay
... I'm, I'm not a-
Yeah, yeah
... a vision person. Yeah.
Okay. Uh, the other, the other element of multimodality I'm interested in, and that came up in the ADEPT paper... Oh, yeah, please, please. I, we've been talking for an hour and a half. Um, is-- I, I've been calling this screen modality, uh, screen vision versus general vision.
Mm.
In the sense that, uh, ADEPT is, like, very, very focused on screens, tables, charts, blah, blah, blah, and most vision models focus on things in the, in the real world and, and, uh, embodied, uh, sort of images. Do you have a view on the usefulness for this?
Uh, it should, should it all just be part of a mix? Um, anything of that nature? 'Cause I s- I, I see this as the, the primary division now in multimodal focuses that, uh, I came away from when I talked to David in, uh, for, for the Adept episode.
Like, I came away really impressed with that idea that actually the more valuable thing should be screens.
I, I don't think there's, like, a huge, like... I mean, I think at the end of the day, like, maybe screen intelligence is, like, more useful in general. But, like, what if you have, like, a natural image in the screen?
Yeah. They should, they should be part of a mix.
I don't... I mean, no, no, I think at the end of the day it should be mixed, right? If a model can do natural image as well, it should be able to do screen well and everything. I think at the end of the day, like, the models would become, like...
I don't, I don't see that there will be, like s- like, screen agents and, like, natural image in, in... Humans, like, you can read what's on the screen, you can go out and appreciate the scenery, right? You're, you're not, like, say, "I only can look at screens."
and it's like-
Yeah, yeah
... right? So I, I mean, uh, I think eventually the models would, like, be this good on everything. Uh- ... I, I, I don't, I don't feel like, okay, there's a, there's a, there's a, there's a, uh... I think we, we-- Like, I look at it from a point of, like, capabilities.
Uh, and screen is, like, you know, there's even screen, there's also like, you know, like, like mobile phone screen. Uh, and there's also, like, you know, laptop screen, like, also, like, you know, different type of interfaces and everything, like reading emails, whatever, right?
But, like... Or, like, reading a page from a website or, like, you know, buying something from Ama- like Amazon or something. Like, all kinds of things, right? Or, and then even in, in the picture of like a shopping website, there could be like a natural...
Or like, for example, like picking Airbnb, right? To, to buy. Like, there's, there's a natural image in there, then it's like you have to understand, like, how nice is the scenery, right? Or like, you know, like, like, like where is it, right?
Like-
Yeah
... so I think at the end of the day, it's probably like the same-
If you want to build a general model.
Yeah, yeah. But, but I think the natural images is, like, way easier. Like, as, as in just way, like the models currently. Current models are actually already very pretty good at, at, at, at, at, at, uh, at, uh, these natural, natural images.
Uh, and I think like screen images are just something that people need to like, uh, enhance the capability a little more. That's why there's like some focus on, on that. Yeah.
Got it. Okay, excellent. Um, I'll touch on three more things and then we, we'll just go to career stuff.
Mm.
Uh, scaling laws. PaLM 2 was Chinchilla, which is one-to-one scaling of model parameters and data. Now you are training a 7B model with 5 trillion tokens. Um, what are you thinking about the, the trend in scaling laws for data versus brands?
Scaling & Context1:35:36
Chinchilla's scaling laws are just like optimal for-
Compute
... like with this amount of compute, how much do you have to think, right? But like actually the optimal, like there's no... I mean, this is something that, that even before I left, like we already, you know, we, we already knew that like Chinchilla's scaling laws are not the, the end or, or like of, of it, right?
Obviously, there's also a inference optimal scaling law, which is obviously you take a small model and then you just blast it with as much compute and data as you can, uh-
Until?
Until you saturate on everything that you care about, right? Uh-
Okay
... right. So I think, I think like, uh, like Llama 3 is what? 15t tokens or something, right?
Yeah.
So I, I think-
Which is ridiculous.
It is, it is, it's ridiculous to be honest. But at, at, at a certain point of time, your value per flop is like not great anymore because you just... You know, your models eventually gets, get, get like saturated.
But then the, the, the problem of like, the question of like where is this saturation is also like you always find like some metric that you still continue to improve a little bit and then you're like, okay, maybe like, like, like, oh, 100K more is worth it to continue training, like just a little bit more, right?
But, but then it's like, like where does it end, right? But I think at the end of the day, like, like, uh, the f- the thing about Chinchilla's scaling laws is that like it was a bit misunderstood. Like, uh, I, it's not really like...
There, there was not any like bad intention in the way it was framed. It's just that the, the, the, the, the, the... It got misunderstood as, as though like, like this model you need this compute and, and, and if you train the Chinchilla's scaling law of like, like, uh, like you, you kind of like...
Like it w- I don't know why so many people had this, had this idea that, that you will not improve past the Chinchilla's scaling law. And then people make so much big deal about- About, uh, uh, like, uh, uh, you know, trading past Chinchilla scaling law.
Like, oh, Llama 2 is the first model. It's like, like T5 base, right, was one trillion tokens. That was already so much beyond Chinchilla scaling law, right? Because that was T5 base, right? So I, I don't know why so many people are so surprised about, about going past Chinchilla scaling law when, when, uh, when-
I think OPT and, uh, GPT, m- uh, maybe s- maybe set that as an industry standard. As GPT-3 specifically. Uh, I don't know. That, that's my initial thought.
Uh-
No, sorry. G- wait, GP- GPT-3 was not Chinchilla scaling, right?
No, I think like OPT and Bloom, right?
Yeah.
Models like this-
Yeah, yeah
... they were, they, they trained a large model and with a very small number of tokens and, and the model turned out to be bad.
Yeah, yeah. So, um, th- I'm talking about Kaplan, the, the, the pre-Chinchilla one, um, the Ka- uh, Kaplan scaling laws.
Oh, okay, okay. I see.
Uh, that one was, was from OpenAI.
Yeah.
Anyway, uh, death of Chinchilla covered. Agreed. Uh-
But Chinchilla is still a cool paper. I think Chinchilla is still a- ... it's still important paper
Any-- I love any scaling laws paper, to be honest. It's like such a ch- service to the community in, in general. Um, yeah. Uh, the Hugging Face recently did one, DataBlations, uh, which is like a data scaling laws paper.
Mm-hmm.
Um, looking at data constraints, uh, which was, which is kind of nice.
I see.
Um, long context. Y- um, people are talking million token, uh, context, uh, 2 million token from Gemini. Magic is talking about 100 million token. Um, what-- h- how important is it, do you think?
I think we, we, we need to solve benchmarks first before solving the long context, right?
We have your benchmark.
No, no, no. Not, not, not, not, not, not like benchmarks for long context.
Okay, yeah.
Because like you-- like the needle in haystack is, uh, basically like, like, like a M- MNIST like or it's like a unit test for, for, for this type of things, right? But like I think like, uh, like there's, there's a one part about like hitting the context line and other part about like actually like-
Utilizing
... utilizing, right. I, I think G- Gemini's long context is surely like amazing, right? But I think like for the community to move forward in this, then it comes to a problem of like how do you evaluate this?
Uh, I think I see some long context benchmark like coding one like and, and, and stuff like that. Like I think making those are important and, and, and for the, for the community to hill climb. Uh, but I, I think long context is, is, is important.
It's just that you don't have a very good way to like, uh, like measure them like, like properly now. Uh, and yeah, I, I mean I, I think l- long context is definitely the f- the future rather than RAG.
Uh, but I mean they could be used in conjunction like, like-
Definitely rather the future.
I, I-
Okay.
Yeah, yeah. I don't know.
That's an hot take.
Which part of the-- which part, which-
Context, uh, long context is the future rather than RAG. Like you would-- they will coexist, but you are very positive on long context. Wait, I would, I would put myself on the other, on the-
Okay. So you-
... mirror, mirror image, which is like long context is good for prototyping, but any production system would just move to RAG.
There are, there are a lot of, uh, application use cases where you want a model to take that time and then come up with the right answer, right?
Sure.
Because RAG is like-
But you'll use those sparingly because they're expensive calls.
Yeah. You-- it depends on like the nature of the, the, the application, I think, because if in RAG, right, like you-- there's a lot of issues like, okay, how you-- like y- the, the retrieval itself is the issue or like, you know, you, you, you might, you, you get fragmented.
Like, you know, you-- if, if it's like what if it's like a, a, a, like a very complex story, right? That you-- like a storybook or like a complex like thing, right? And then, and then like re-- like RAG is very like y- you kind of-
Chunks
... chunks and chunks, right?
Yeah.
The chunking is like-
I get it
... uh, and you definitely have lots of information, right?
Yeah. Yeah.
So there-- I think there are a lot of-
That's true
... application use cases where you just want the model-- like you are like, "Okay, like 100 bucks, like take your time, take one whole day." "Come back to me with like the answer," right? Rather than like I, I pay like, like-
Yeah
... one cent and then like get back a wrong answer. So I think there's, there's like, uh, uh-- and then it is, it's actually very easy to show that RAG is better than long context because there are a lot of tasks that don't need this long context.
You would like-
Right
... like fact retrieval, you just like RAG and then you do th- this thing, right? So like long context may get a unfairly bad rap sometimes because like it's very easy to show like, like RAG is like 100 times cheaper and, and it's very easy to show this.
Mm-hmm.
Right? But then it's very-- it, it's, it's also like not so easy to emphasize the times where you actually really need the, the, like the long context to really make like very, very, very, very, very, very good like decisions.
Uh, so yeah, I mean, I, I think both have their pros and cons depending on the use cases. Using them together is also interesting. Uh, and like at the end of the day, it's like a hyperam that you have to, to, to wiggle around, right?
Uh, yeah.
There's another wiggle on the hyperam, or there's another toggle on the hyperam, which is how much you fine-tune new knowledge into the model. Are you, are you positive on that? Do you have any views?
I can, I can elaborate if you want.
Yeah, go ahead.
Uh, so for example, instead of doing RAG on a, uh, on a corpus and then, and then inserting into, uh, context, um, you would just fine-tune your model on the corpus so it learns the new knowledge, uh, in, in whatever capacity, right?
Like...
Uh, this is cumbersome, I guess. This is cumbersome, and you don't want like, you don't want so many of-- like the point of in-context learning is so that you don't actually have to, to do... I think this one is depending on like the business use case, right?
If, if fine-tuning is actually like the, like you, you, you are very clear, like you want this, this knowledge, and then you just fine-tune once, and then you don't ever have to pay like, like context, like in the context window cost again, then maybe that makes sense.
Uh, but if the domain is keep changing, then you might not like.
Yeah. Obviously, it doesn't make sense if the domain keeps changing. But I think for, for the model to maybe update fundamental assumptions or, uh, you know, re-weight associations between words, uh, for let's say a legal context versus a financial or medical context, like it might ...
work. Th- th- this, this is the arguments that some, some people are talking about. So, you know, in, in t- I, I see this as a trio. Like it's long context, it's RAG, and it's fine-tuning. Like, people always have this like, uh, wh- whether either of them will kill RAG, basically.
Because RAG is kind of the simplest approach.
Uh, yeah, yeah. Okay. I, I mean, I, I could see like, like if you want a, like a model for medical domain, legal domain, then fine-tuning really works. It's always the most, like the, you know, s- uh, domain specialized model, universal model and, and you know, the kind of this tension between both of them.
Yeah.
Uh, I think it definitely like, uh, uh, makes sense. Uh, and, uh, it also makes sense like to, that, that, uh, fine-tuning can also be like alternative to, to, to RAG. Yeah.
Yeah. Okay. Yeah. Th- well, there's some, there's some companies that are set up entirely just to do that for people. So, uh-
Yeah
... it's, it's interesting that, I mean, I, I, I sort of view Reka as like not working in that space, but you could, you c- you, you could potentially offer that if you wanted, wanted to. Um, okay. Uh, I was gonna ask about, uh, efficiency and scaling.
Um, I'll just mention this briefly and, and then I, then we can talk about MoEs 'cause I discovered that you, you were, you wrote-- You were co-author on the Sparse Upcycling paper, uh, which is actually-
MoEs & Efficiency1:44:36
Oh, no, I was just advising on that. Like-
Oh, okay.
Yeah, yeah.
But you can talk about Sparse Upcycling.
Yeah.
It's a topic that's hot. But more generally, efficiency, I- in my mind, uh, when I go to IC- ICLR or I go to NeurIPS and I see efficiency paper, 90% of the chance, like I'm just gonna ignore it because I, I don't know if it's gonna work.
A- and I think this is related to your, some of your scaling work and your induct- inductive bias work.
Oh, okay. Scaling law, isn't it?
Which is like, okay, uh, uh, there was this, uh, Tir Texas, I don't know who this, this person is on Twitter.
He keeps talking about me.
He's fucking amazing.
Oh, okay.
Uh, y- yeah, he, he, he does have some obsessions, but like, he's good. I don't know who he is, but he's good. So he says, "If twentyfo- if 2024 papers are to be trusted, you don't need most attention, you don't need high precision, you don't need most KV cache, you don't need most feedforwa- uh, network layers, you don't need a reward model," blah, blah.
Like, it's like a lot of efficiency papers are just like, "Hey, on this like small example, we cut this thing out, works fine or works great, works better," whatever. Um, and then it doesn't scale, right? Like or ...
Yeah, yeah, yeah.
So, so it's a very interesting observation where like most efficiency work is just busywork or like it's work at a small scale that doesn't, that just ignores the fact that like this thing doesn't scale because you haven't scaled it.
It's just fine for a grad student, but f- as for someone who's try- trying to figure out what to pay attention to, it's very difficult to figure out what is a worthwhile direction in efficiency.
Yeah. That's, that's, that's, that's, that's a good point. Uh, I think there's a couple-- I, I agree with you fundamentally that like, uh, that i- it's actually quite easy to tell, like when you see a paper, "Okay, this one doesn't work, this one works, this one doesn't work."
Uh, I, I guess the Hippo account will just tell you that sometimes directly about this thing doesn't work, this thing works, everything, right? So- sometimes it's not like, you know, like i- i- i- it's like you, you can always find a task and a dataset where your efficiency method, m- m- method gets neutral results, right?
You can always find one thing that has, okay, I have comparable p- perplexity at... And you know what's the most, the cutest thing ever? Every time pe- some people propose like this, they run like some zero-shot score on like some LM eval harness or something like that.
And you know, like at 1B scale, all the numbers are random, basically. Like all your BLU kill, class, they're all like random chance performance, right? And they'll be like, "Okay, I get like fifty versus fifty-four. I'm better." But like, dude, that's all random chance, right?
Like you know, I've, I've seen papers that-
Oh, okay
... that, that they run experiments at like-
Okay
... and then it's ra- random.
That's, that's a good tell. Yeah.
Uh-
It's a good tell
... right. So, uh, I, I, I think it's very-- Like the, the sad truth is that like, it's very hard to tell unle- until you, you, you, you scale out. And, and sometimes the benchmarks that we have don't even probe entirely about, uh, what, you know-- Like there, there's-- I mean, especially all the works about the, you know, the, the transformer alternatives, right?
You can always find like this alternative that, that at 7B scale, at one, 3B scale, you kind of like, "Okay, I met transformer on this and this, this, this," right? But then what's the implications when you go to like 200B?
What's the implications when you go to 100B? No one knows, uh, uh, that, right? So, uh, I think that's, that's one thing, right? And, and, uh, and, uh, yeah. I, I, I, I think developing your own intuition of like what works and what doesn't work is, is, is, is important.
Uh, and, and like for example, some, if somebody's like, uh-- I, okay, to be honest, all researchers like sometimes are also like, like guilty of this sometimes because you cannot test on like everything. They cannot test on everything, right?
So sometimes you also just want to show your method works on this. But it depends on the objective. If the object-- If the, if the objective is to write a paper to ICML, sure, you can, you can find two datasets your, your, your stuff works, right?
But will it get adopted? I, I'm not sure. So...
Yeah. Well, uh, you know, researcher meta game is one thing, but as a consumer of research, I like, I'm also trying to figure out like, what is-- how do I know what is a, uh, what is a useful direction, you know?
That, that's the interesting thing. So, for example, um, MoEs seem to have worked out in a s-
Yeah, yeah.
Uh, I, I, I'll go so far to say it's the first form of sparsity that worked. Like
Okay.
'Cause there's, there's so much sparsity research like we can, you know, chop, chop, chop, chop, chop all these parameters and, uh, look, we still, still perform the same, but then it, it never actually works. But, but MoE is really-
Oh, you mean like the pruning line of work.
Pruning, pruning line of work.
Okay.
Sorry. I, I should have used that word. Um, uh, so like, you know, it's-- it-- I don't know if you have any commentary on like, uh, Mixtral, DeepSeek, Snowflake, Qwen, uh, all these, um, proliferation of, uh, MoEs, MoE models that seem to all be sparse upcycled because, you know, you, you were advisor on, on the Sparse Upcycling paper.
The, the Sparse Upsize- Upcycling paper was mostly vision focused with a little bit of T5-
Okay
... experiments. So it was a-
So this is much more-
It was, it was, it was like the, the-- it was a very like- Like, like early stage of like sp- sp- large-
Yeah
... tracking. But it was good that Google was already thinking about this, like, long ago.
And, and Nome also had a-
That, yeah
... paper on it, right?
Yeah. Uh, and then, so I think... Wait, what was the question again? Like-
So, like-
What I think about-
Yeah, what do you think about MoEs?
I think MoEs are the way to go.
Is it very promising?
I think MoEs are the way to go.
I- is it, like, 100 experts, is it 1,000 expert, you know? Like, for some reason, the, the community settled on eight.
Now you probably get more gains from-
Yeah
... from more, more than eight. But, like, I think in general, it's like MoEs are, are just a trade-off with, like, param and, and flop, right? And then you're able to, like, kind of make-
Active param. Yeah
... like, you cannot make that, that, that in- like, that, that scaling law increase from, from that, uh, additional like, like, uh... So you, you can keep a low flop but kind of have more parameters. It's just changing the flop parameter ratio.
Mm-hmm.
Uh, in fact, I think-
Keeping in mind there's a lot of inefficiency between the experts.
Yeah, yeah. But I, I think, I think that it is, is, is, is, is, is, is, uh... How, how do I s- how do I say? I think as a architecture itself, the flop parameter ratio makes it, like, worth it, right?
But I think the, the thing that is not very well understood is that, like, how does, like, MoE... Like, like for me as a res- research question is that, like, when you... Like, how does it, like, uh, uh, relate to capabilities and stuff like that?
Like, does this inductive bias actually, like, uh, you know... Like, for example, when you, when you do, like, massive instruction tuning, I think there was this paper, like, like Flan MoE or something. Like, they show that like, you know, like instruction tuning.
I'm not, I'm not like f- fully sure ab- about, I don't recall fully, but, but like when you do in- massive instruction tuning, like MoE models are like they behave differently from, from dense models and stuff like that.
Like, I think it... Okay, like, fundamentally, I just think that MoEs are just like, uh, the way to go in terms of like flop parameter ratio. They, they bring the benefit from the scaling curve. If you do it right, if you, they bring the benefit from the scaling curve, right?
And then, like, that is like, that's the, the, the, the, the performance per flop argument, like activated params, whatever. That, that's like kind of like, that's a way to slightly cheat the scaling law a little bit, right? By having more parameters, right?
The, I think the, the, the, the more interesting thing is about like, like what trade-offs do you make in terms of capabilities because of this new architecture.
Mm.
I think that's actually like the, uh, the, the question that, uh, I, I think, I, I guess all the frontier labs are, they already know this, but nobody's writing papers anymore about this, so, like, you just have to live with, with, uh, with, with, with what's outside.
But I think MoEs are the-- I think Mo- I'm, I'm, I'm bullish about, about, uh, about MoEs.
Yeah.
Yeah.
I had to, uh, I made an exercise for myself on rating research directions and what their asymp- asymptotic value is.
Mm-hmm.
And I put MoEs pretty low because I think you have a good base, you have a good base model, and then you upcycle it, and it bumps you a little bit, and I think that's it. Um, but like I, I'm always seeking to invalidate my hypothesis, right?
Oh, but, but like, uh, from scratch MoE is also promising, right?
From scratch MoE is promising. Uh, sure.
I think in the I- in the I- in the IU case you would do MoE from scratch, I think.
Yeah. Actually, yeah.
I think in the IU case you would do, uh-
Yeah
... uh, MoE from scratch.
Upcycling is just a-
Upcycling is just a-
... upcycling
... a, a commit, like a-
I think people still harbor, uh... So there's some rumors about the architecture of GPT-4, where they had pluggable, um, uh, experts, in the sense that the vision model was p- vision expert was plug- plug- pluggable. I don't know if that makes sense at all, but this is something that was said.
Mm. I see. I see.
Uh, I, and, uh, I mean, it could just be as simple as just sw- s- uh, swapping out the, the, the MLP side of MoE. I don't, I don't know. Um, okay, cool. Yeah, this, it's all speculation. I, I...
Okay, the, the last part that makes me uncomfortable about MoE debate is, uh, it, it actually is related to another paper that you wrote about, um, uh, the efficiency misnomer, in the sense that, like now people are trying to make the debate all about the active parameters rather than total parameters.
Ah.
Uh, but it seems like, it sounds like that's something that you're comfortable with, like flops at, at inference, um, i- is, is a relevant metric. Um, and it's, it's not that-
Well, thanks for like actually reading all the, like reading the papers there.
I'm trying, man.
Thank, thanks for-
It's very hard to co- it's very hard to... You have a lot of papers.
No, I'm actually very impressed that like, oh, you are bringing up these, these papers very, very-
Yeah, I'm using attention context.
Okay. Yeah, thanks. Thanks. Uh-
And also, like, I mean, I'm interested in efficiency that works. It's just very hard to find efficiency that works. Uh, and so, like any, anything that helps me have high signal on efficiency is helpful.
So I, I think, I think, like efficien- for the efficiency misnomer, by the way, I, I love the paper, by the way. It's-- I had a fun time working on it. Uh, I think efficiency misnomer was like we found that like a lot of people, like they, they use params, like especially like, like to do kind of like...
Right. And then MoEs, uh, was not very hot like in the, like, community at that time, right? And but MoEs were like a thing long ago, like at, at Google, right? Uh, so I think using active params, I'm comfortable with using a- active params to, to kind of approximate like cost for the model.
But like in the efficiency misnomer paper, we actually m- made it quite clear that you should always like look holistically about like-
Yeah
... like your... Because, you know, like you, you have serving, like additional serving costs.
Yeah.
You f- like fitting in, in, in GPUs, like fitting on a single node and something like that.
An interesting one was speed. Like, and, you know, not nobody, nobody really, uh, talks about speed, but your, your paper actually talks about speed.
Oh, okay. I, I have a, I have a interest-- I, I have something to say about speed.
Okay.
Throughput, right?
Yeah.
There are so many methods, right, that are proposed about efficiency, right? They are like theoretically like faster-
Okay
... because of like complexity, like something like that.
Okay.
But because there's no way to work around the implemen- implementation-
Mm
... or like your implementation becomes so hard, it becomes like 10x slower.
Okay.
There are so many papers around-
It's not hardware aware
... like, it's, it's, it's-- It could be hard-- It might not be-- It could be hardware. It could be like, it could be like, uh, just the way that... Like you, you have a conven- like you have a convenient way to like, like i- i- in, in, in, in this, like, in its mathematical form, it's actually like, okay, linear complexity, like whatever, and it's actually theoretically faster.
But like just because you have to like do a scan or something like that, like, like, and like, and then it becomes like- Actually, like 10 times slower in, in, in practice, right? There, there are a lot of things like-- No, not a lot, but like there are some things that are like some methods that are like, uh, like, like, like, like, like this where you don't take it into account throughput, right?
Which is also the problem of like sometimes like the incentives of like, like people who gain effic- effic- efficiency. You can easily just like sell a paper as like more efficient, and then people-
Ignore throughput
... people will not, like, like people will not suspect that like, like, uh, uh, like... Because the, the, the, the reason why we wrote the paper is that so many people were confused about like efficiency itself, right?
Yes.
And then they will be like, okay, like, uh, a lot of these unsuspecting reviewers, especially like even academics or they, they, they don't have like that, that r- r- real feeling. They were less like, "Okay, less parameters, more efficient," right?
So you could have a method that's like less parameters but like three times slower because, you know, a lot of times when you add things to the model, it, it becomes slow. You add-- Every time you add complexity, especially if it's like something that is not hardware optimized, no kernels or like something that is like bad for TPUs or whatever, your, your model just become like, like slow.
Oh, that's a temporary issue.
People can-
You can fix it
... pe- people can fix it, but some things are not like so-- like some things may not be like so easily fixed or like it just adds a lot of ha- like, like sweet cost to, to, to optimize it and everything, right?
But then it's always marketed as like because I save param so I save.
I see.
Like, right. And then also like the params where you add a different place of the model. Like for example, like, uh, for example, like for examp-- if let's say you, you, uh, uh, e- even, even in the case where you param match models, right?
If, if I take out like some params from like FFN, right? And I put it to like embedding layer, right? Embedding layer is like a, it's just, it's, it's a cheap operation for embedding layer, right? But my model becomes like lopsided, right?
I could say I param match this, but it's not- ... flop match. It's not throughput match, right?
Yeah.
Uh, because the, the-
It's unbalanced on one side
... it's, it's, it's unbalanced on the, the, the, the, the side, right? So there's a lot of this type of tricky things that like, that when makes com-- model comparisons like very, very, very, very, very difficult, uh, and, and because you cannot even put like flop throughput and speed, uh, flop params and speed, like actual speed, right, in the same plot, right?
And then there's always like one money shot in the like-- there's always like a Pareto like kind of co-compute, like whatever plot, right? Uh, like for marketing in papers or something like that. It's always very easy to like, I mean, not intentionally, but like to subconsciously like show one story when it's actually like there's like all these other things to consider.
Yeah, yeah.
Right.
It's a selection bias, self-bias-
Yeah
... whatever. Um, very cool. Very cool. Uh, okay. Well, that, that was mostly of, most of the technical side. Um, we have one commentary that will happen today on the future of open source models. Uh, the, the basically Founders Fund said like the future is closed source.
Open vs Closed1:58:16
Uh, you were, you were, you were, you were agreeing with it. Um, and a lot of the open source fanatics, uh, you know, are up in arms over, over this. Uh, I don't know if you, you care to comment about just-
Oh, okay, okay. Uh-
... open, open versus closed and closed source whatever.
So, so, so I, I mean, I, I don't really like-- When I mean like, uh, if you're, if you're, if you're referring to the tweet that I wrote, but like I wrote something about, about, about, about-
But this is huge. Like so many people are commenting about it 'cause they, they are personally physically offended that open source cannot catch up.
Okay, wait. No, no, wait. Okay. So I, I, I want to say this. It's like I'm not-- Like I contributed to open source in the past, so I'm not like against like open source-
Yeah
... per se. But I'm-- The, the, the thing, the, the interesting thing that I want to talk about here is that like, uh, there's a difference between-- Like I, I, I draw a line with like, uh, like open source as in like, okay, like to me Llama 3 is like, it's, it's like Meta has an org that is like, okay, hypothetically very similar to, to, to like Gemini or something, but they just didn't decide to release the weights, right?
Yeah, it's open weights.
It's, it's open, it's, it's open weights everything, right? I think when most people try to say that like open source is-
Yeah
... catching up everything, they kind of mean like this grassroots like-
Uh, yeah, the distillation
... like this, this, this, no, this, this bottom up people that are like, that are like, uh, uh, like this, uh, indie developers that are like coming together to like, like fight. Like it's, it's romanticized and it's dramatized to some extent.
Yeah.
Just to fight against like this.
It definitely is.
Right. Uh, and to be very fair, I think that there isn't really much like-- Like so far, if you just look at like the, the fractions of people, the big labs are just pushing and pushing and pushing. The academics like Stanford and stuff, they came out with DPO, they came out with things like that.
They, they make some like-- But they, they're kind of in be- in between the line of like open source community and, and, and, and then there's also like the developers that are like fine-tuning on GPT-4, Distil, Distil models and everything, right?
So I think that like, uh, uh, uh, I, I don't-- I think the, the, the open, open source underlying like thing about like, uh, uh, collectively improving something. I, I, I'm, I'm, I'm, I'm not like criticizing it for the sake of criticizing it, but like I'm just saying that like in, in order to make progress, right, the-- I think the incentives of open source are like-- What I observe is that like, like people like to do things like they like to take somebody else model, they rename it, and then they, they take-- They make a quick-
Yes
... they make a quick, quick, quick win from that.
Yeah, I think we have to close up in the next 10 minutes. Yeah.
Uh.
That's it.
Uh, they, they will make, they, they, they'll make a quick like... And then like but you notice that like when people realize that like this, uh, like turning on the GPT-4 tab and running some DPO is not going to give them the reward signal that they want anymore, right?
Then all these variants gone, right? You know, there was this era where there's, what, there's so many of these like I cannot even-- I lost track of this, like all these model variants. But, but, uh, but now they're all gone because people realize that, that you cannot climb LMsys because you, you need something more than just something that is lightweight, right?
So I think that was just my, my overall like-
Honestly, the Hugging Face leaderboard contributed to most of that. It's not LMsys. I, I-
No, no, I think LMsys probably they realized that they could not.
Oh, yeah.
Yeah. Right. The op- the OpenLM leaderboard is like probably like the, the, the, a bigger, a big like- ... problem, to be honest. Uh, so
We- we're talking to Clementine in, in, uh, in, in one of our future episodes, so-
Okay, okay, okay. Yeah, yeah
... they, they dedicate a lot of... I mean, there's so much attention to them, it's, it's, it's a tough problem. But they're p- providing a public service, for sure.
Yeah, yeah. I mean, good intentions are always good. I mean, good-
Yeah
... like, good intentions are always good. Yeah. Uh-
Yeah. Rather have them than, than not have them-
Yeah, yeah
... is, that's what I'll put it. Um, okay. Uh, you know, due to, to, to cut short on time, um, I'm interested in, like, just like, just career-wise, uh, what is your productivity practice, or... And so I'll split it into three, three things.
Career Advice2:02:23
Re- k- keeping up, uh, like reading papers and whatever, and just the outside world. And then two, like, uh, how you organize your own work. And then three, like, m- work and life. Uh, just use any, any... Take that in any order that you wish.
Uh, I, I, I d- I, I don't have much of a life, actually. Uh, but I am trying more to have more.
I mean, you're a father now and, yeah.
I, I, I have a baby now, so, like, I'm trying more to have more life, uh, and, and, and everything like this. Uh, productivity-wise, I would say that, like, I, I just-- I think I, uh, uh, I think f- the productivity hack that I have is just, like, I didn't have, like, a boundary between my life and my work-
Okay
... like, for a long time. So I s- I, I think I just cared a lot about working most of the time. Actually, for the last, like, doing my PhD, doing my-- at Google and everything, I w- I'll be just, like, working all the time.
Uh, it's not, like, the most healthy thing, like, like, ever. Uh, but I think that, that, that was actually, like, one of the biggest, like, productivity, uh, like, uh... And I spent, like, I, I like to spend a lot of time, like, writing code.
Uh, and I, I just enjoy running experiments, writing code, and stuff like that, right? So you k- you kind of-- if you enjoy something, it's not work, right? So, like, it's, like, it's very strange. It's like, it's like I, I would get distracted by...
Like, sometimes I have to watch some Netflix series because, like, my wife asked me to, like, watch it. Like, or some- somebody tells me that, like, what, what, like, I, I, I, I've, I, I'm, I'm back on time on some, some shows, right?
Uh, but then I get distracted by my experiments running, and I just end up, like, like, writing code instead of, like-
Wow. That's great
... like, so, so, so, so, so, so things like, like this. I, and it's not the most healthy thing, but I think that's one.
I, I'm looking for, like, a practice where, like... Okay, uh, so Andre recently had a thing where, like, before-- when he wakes up, he doesn't look at social media. He only goes straight to co- straight to work.
Damn, I check Twitter the moment I wake up.
I know. See? Like, which is, which is something I do as well. But I'm like, "Damn, damn, that's, that's a smart rule." And, like, I'm looking for, like, rules like that. Like, do you have a rule-
No, he doesn't check social media because his phone is exploding all the time.
All the time, yeah. I'm sure.
Right? I don't have so many likes and followers, so, like, it's fine for me.
Yeah, you get there.
Um-
Uh, but, like, rules like that, mantras that you've developed for yourself where you're like, "Okay, I must do this." So for example, recently for me, um, I, I, I've been trying to run my life on calendar for a long time, and I found that the only way that I work is I, I write things down on, on pen and paper, and I cross them off individually.
And, like, that, that physical action really, really helps me, helps me, uh, you know, get things sorted. Uh, or, and, and that's, that's work, work-wise. Reading-wise, uh, I, I don't know if you know, but I've been running this, like, AI newsletter, uh, that, like, auto-summarizes all Twitter, Reddit-
Yeah, yeah
... Discord and all that. Uh, so that helps me keep up because I, I have, like, a socially graded, um, and, and I, I, I r- I personally, uh, vetted the entire pipeline from beginning to end. So, like, this is my input algorithm, and I know how to keep up with news because I, I now have a, a information condenser.
Mm.
So, like, what I'm trying to figure out, what is your algorithm or what's your rules for keeping up?
Oh, I got something. I got, I got something. Uh, so for keeping up, so I, I used to, to, uh, check archive, uh, like, every morning when the gate opens.
Wow.
I just check archive. I will wake up 9:30 a.m. Singapore time, the archive gate opens, right? And then I'll be very sad if there's no papers to read. But you usually just pick one paper or two papers that you find interesting.
I don't read them. I just, like, skim, like-
I see
... the, the, the thing, right?
Yeah.
Uh, so I used to that. I don't do that anymore. I, I, I mean, ever since, like, I have in the startup, I s- I read, read-
Yeah, you have, you have a real job now.
I, I read, I read, I read less papers, right? But I used to camp at the door of archive, uh, quite frequently just to see.
That's, isn't that-- That's not a good use of time. I, I'll, I'll come out and say it. It's, it's not a good use of time.
No, no, no. I said-
It's a newness bias. Sorry, go ahead.
No, no. It's just because, like, I ran, ran out of things to, to-
I see. Yeah
... to... It's just that, like, the new stuff comes out, right?
Yeah.
Like, and, and, and then, and like, the new stuff come out, right? So that's how I keep up to date to like-
So in the space of three years, you read every-
No, no. I didn't read everything. It's just that-
... AI, NLP
... it's just, it's just that-- But these days, I realize I don't have to do that anymore just because if the paper is important enough, Twitter will show it to me.
Sure.
Right? So that's true, right? You actually don't have to follow anything. If the paper is important enough, the Twitter algorithm will give it to you.
Yeah.
Uh, so I, I, that's, that isn't really, like... And one thing I do is that I, I actually don't read papers like that, that much anymore. I just, like, skim them-
Yeah
... like, al- almost, right? The, uh, so that's, that's for keeping up, like, with, with papers, research and everything. Uh, and the other thing more of, like, just, like, a productivity point of view is that I used to always keep, like, the, like, you, you know, the, the, the text, like the, the, the overleaf or, like, whatever you call, like, for, like, uh, uh, like I, I usually start writing the thing while working on that thing itself.
Like, so I'll even, even, like, like, like, let's say if, like, like if you want to launch something, like, then the, the end goal is, like, a blog post or shipping something, everything, right? I like-- Or not, not, not really a launch, let's say, or like, like just papers or I always like to look at it from, like, what's the, the story in the end, and then I just, like, figure out what I need to do to get to, to, to kind of-
Mm.
Right? So I think-
Work backwards
... for, for, for as a researcher, like, this is something, like, I would have, like, like, so many drafts of, like, like, uh, like when I'm start, I start a project, I don't know the experiments yet, everything, right?
But I like to imagine, like, what the title will be.
Yeah.
Right? And then I always vibe check, like I always, like... So I, I mean, my friends at Google will know that I always have like, like a, a, like the overleaf, uh, uh, a draft of like, like so many, uh, and then I will just spend time looking at it, like tweaking the title.
Is it better two second line? So I, I care about, I used to care about a lot of this, but this actually helped my productivity because every time I look at it, I'm like, "Okay, this is the final product."
I'm, like, working towards it, right? Because I think a lot of researchers, they, they tend to, like, they, they swoo around in their experiments and they never, like, ship the final story.
Yeah.
It's like the shipping, like, like a-- I mean, a, a startup will ship products, but, like, as a researcher, your, your, the product-
It's a bit like product management, yeah, in some ways
... you, you are shipping the, the thing. So I like to I like to hang around a lot in my, in my drafts and, you know, like I get motivated from that and, and that's like one productivity thing that I did as a, as a, as a, as, as a researcher.
And, and, uh, uh, yeah, so I think that, that, that's, uh... Other than that, I, I don't really have any, uh, uh, like I don't really have any w- like things that I do that probably different from, from others.
Yeah.
I- probably you don't know it. This is unconscious competence versus
Okay, okay.
Um, okay. Uh, I, I, we probably have to, uh, three more questions. Uh, what do you used to strongly believe that you've changed your mind on?
Whoa.
I was not prepared for this question. Let's skip. I don't have like a good answer for this.
Okay. These are the ... N- I've reserved the Singapore questions to the end.
Yeah.
Uh, what's it like just NTU PhD, uh, you know, the, the, just the story of like what ... like how was it coming up from NTU, which is i- which is like a good school but like not, you know, not typical target school for like a big lab.
Uh, I, I, I, I did my PhD unknowingly. Like, I, I didn't have very ... Like, when I was ... I was a very regular undergrad. I had d- decent grades, but not the best grades. I was not like super smart in school or something like that.
Uh, I, I, I was, uh, I wanted to do a PhD just because I was, like, curious and, and I, I mean, like, and then, uh, I wanted to stay in Singapore at the time, so I just like naturally just did a PhD there.
I didn't even vet my advisor. I didn't even think too much. I just like fell into the PhD program.
Okay.
And then that was when I realized that, oh, actually, I can do research. Like, I, I'm, I'm, like, pretty decent at research. Like, I just fell into a PhD, like, like unknowingly.
Yeah.
Uh, and, uh, I definitely like NTU leaves a lot to be desired. Actually, to be honest, I think that, I mean, Singapore leaves a lot to be desired in general. Like, the research community here is like, like probably not great.
Uh, I've also-
So how, how, how did you, like, break out? You know, like i- if I was you, I, I would have, I would have no idea how to break onto the international scene and-
I, I think, I think it was ... Okay, to be honest, like in retrospect, it's a bit of, like a bit of a, a, a miracle or like, I, I mean, it's, it's not easy to ... I, I think I, I can, I could not pro- if, if I had like a pro- like a, a, someone to mentor, I probably could not replicate like the same ...
Like, I could not like tell somebody how to replicate the same thing that, that I did. It's much easier now maybe compared to in the past, but like, actually maybe not. That one I ma- may not be very sure about that.
But, uh, I think like, uh, I, I've, I've been mostly self-supervised doing my PhD. Like my advisor was basically like, like Grammarly. Uh, like a free paid plan of Grammarly. Uh, he won't watch this, so it's fine. But like, uh- ...
uh, I, I, I've, I've, I've learned, like, as in I, I, I ... There's, there's a lot of things that, that, that it wa- it was like this strange arc of my life where I was figuring out research by myself and, and everything and, and, uh ...
Okay, maybe going back to the like, uh, uh-
Change opinion
... the cha- cha- change your opinion is that, like, the biggest culture shock I had, like, when like was moving from Singapore PhD to Google, I think my research, like taste-
Which you went straight to Mountain View, is it?
Yeah. I went, I went to Mountain. I started in Mountain View. Like, my research taste and everything, like, like, like I was in constant, like it was a culture, like my b- like it, it, it was so different, like the research culture is so different in, in, in, in, in US and in Asia that, that, uh, that I had to grow so much, like, during my time at Google to like actually, uh, uh, uh, evolve.
And then whenever I come back, right, I still have friends in like faculty in, in here and everything. They would either think that I'm a snob or they think that I'm like being a, like a very nasty person.
Uh, because like I think to be honest, the, the, the research here is like in Singapore is just basically like they just care about publishing papers and stuff like that.
Mm-hmm.
Uh, and then it's not like impact driven. I think a- at, at US it's mostly focused on impact driven and the thing needs to re- make real impact, right? Uh, so it's this shift like, like I, I, I-
And, and what, to be fair, you're also working at an industrial lab versus an academic circle, right? Like you're, you're comparing apples and oranges here a little bit.
Uh, I, I, no, I, I mean, at the end of the day, I think research is like, like, uh, fundamentally, uh, like we ca- I mean, a- as an industry RS is you write papers. Your goal is to advance science and everything.
To be honest, it's, it's all the s- s- you know, the incentives rewards system is like different and, and maybe like slightly different everything. But like, a- at the end of the day, I still feel that researchers are researchers, scientists are scientists, no matter like really like, like, like, like, like, like where you are.
Uh, so, uh, I, I will get so much dissonance when I come back and I talk to people. Like, I will feel like, "Oh, why do you think like this?" But then I used to think like this. So, like the environment shapes like, like a way a researcher thinks.
Uh, the taste is very important. Uh, the environment you're, like very important. I, I feel like sometimes I try to communicate this to people, and then maybe I come across as a snob, uh, to, to, to like the, the local community here, right?
But like, uh, it's, it's just that there's like, maybe there's so much dense information that I want to bring back but like, like there's no like, like-
Receptiveness
... fast way to like, like transfer like all the, the-
Mm.
Like, like transfer all the, the things that I've learned.
Yeah.
Uh, and, uh, I, I got s- uh, also a big culture shock because I was in Brain in the Singapore office for a while, and I'm reporting to, to-
You're the only Brain person.
Yeah, yeah, in Brain Singapore. And then I had like, um, uh, I took on an intern from NUS actually. And the, the, the research like vibes and the thing was so much of a conflict for me that it was almost like my body was rejecting it, you know?
Mm-hmm.
Uh, but th- this person so like grow, grew and became, uh, like I, I, I'm happy with how this person grew with from, from my mentorship. So he's now in a way better situation. Uh, but I would say that like, like a lot of people in the, in, in universities here are like, not like, a bit like- Uh, uh, like the, uh, ignorance is bliss, right?
Maybe sometimes. Uh, so-
Well, no, the, uh, it's exposure. Um, I didn't know any better myself until I went to, to the US for, for college and, uh, then, yeah, my, my world was expanded and it's a little bit of a ch- um, Pandora's box because once you've tasted that, you're never happy.
Yeah, yeah, yeah.
You know? Um, yeah. So, okay, last question would be, um, just, just a sort of Singapore, uh, question. I, so what I, I like to be visible, visibly non-American, um, covering the AI scene because it's very US centric.
Um, every non-American I talk to always wants to be like, "How can we br- build Si- Silicon Valley in my city?" You know, my, my country, my city, whatever, that is, that's not Silicon Valley. Uh, I feel like you have basically just kind of like me, you kind of operate in the US cir- circles, but you just don't live there.
Singapore2:14:40
Um, do you have any advice for like, uh, if, if, if Singapore... Okay, so I'm, I'm wearing a race shirt today.
Yeah.
This is the, this official Singapore government's sort of community group that is, uh, that is trying to guide Singapore AI policy. If we want 100 more Yi Tays to come out, what should we be, what should, what should governments be doing?
What should, what should, uh, communities, ecosystems should be, be doing?
Uh, so I actually think that like sometimes like, like, uh, like not doing too es- like too much is maybe less is more maybe. I, I don't think there's actually much like the government can do to like influence, like this kind of thing is like a natural-
Organic
... like organic natural thing, right? Uh, the worst thing to do is probably like to create like, like create a lot of artificial things that like, uh-
Exchange programs?
Okay. I mean, Singapore used to have a lot of exchange programs. Like they send people to-
NOC used to have a lot, yeah
... to, to, to, to, to, I mean, just talking about AI specifically, right?
Yeah.
I think that, um, like for example, like sometimes like trying to do like too much or like moving in the right, wrong direction is just better than not moving at all. Especially if you, if you accelerate in the wrong direction, you actually get into a worse-
Sure
... state, state than possible, right?
Sure.
So I think it's very dangerous to like move in, in a, in a, in a, in a, in a, in a, in a bad like direction. Uh, about, I think respect your talent more maybe. The government should just respect the talent more.
Uh, and like, I don't know whether this is too much of a-
No, no, no, no. Go for it
... it's, but, but, but not, not, maybe not like
moving in a wrong direction is, to me is a, already a very good thing. So like, I think that, that's my, my, my take is that like, uh, uh, and yeah. I've, I've seen on, on, on, yeah, I think that's basically like the, the overall, uh...
You, you, you need to like, like ask me specific, specific things because I-
Prompts
... I, I, I, you need to prompt engineer me a bit.
Uh, funding, incu- uh, for s- for startups, incubation, um, g- getting, um, holding, holding, uh, academic conferences. I think I, I clear next year is gonna be in Singapore, so people come here and exposed to it. But like, uh, I, I don't know.
This is just, it's very interesting. Like, um, everyone wants to build up AI expertise within their own country and like there's a massive brain drain to the US. I'm part of that, like I live, I live there.
Mm.
Uh, I feel guilty and, uh, I don't see any other way around it. Uh, it's, it's such a huge problem and, and, uh, I, I also do, I also do think that there is like a cultural hegemony, let's call it, like US values bas- basically being asserted on the whole world, right?
Like, because we decide RLHF on, on these models and, and now you, you shall use all our models. Um, and it, it's just troubling for like, national sovereignty should be AI sovereignty and, uh, I, I don't know how to achieve it for people.
It's very scary.
Hmm. Okay. There's, there's a lot to, to unpack. Uh-
Yeah. This is not technical, but I was just, you know, curious. Because obviously like, uh, I, so, you know, we can make, we can make this the ending conversation, which is I think you, you have like, you're, you're an inspiration to a lot of other people, uh, who wanna follow your career path and, you know, very, I'm really glad that we got the chance to like-
Mm
... gonna walk through your career a bit and, uh, yeah, I'm sure this is just the start. So, uh, hopefully there's, uh, there's more to come and I, we want to, we, we, I wanna inspire more of you.
Okay. Yeah, yeah. Sounds, sounds, sounds, sounds, sounds, sounds good. Yeah





