Intro0:00
Hey everyone, welcome to the Latent Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners. I'm joined by my co-host Swyx, founder of Small AI.
And today we have Dylan Patel in the pod at Mini Studios. Welcome.
Thank you for having me, and it was very short notice, right?
Yes, yes. Uh, just hours. I, I was thinking you were in Taiwan somewhere, and I was like, "It, it's gonna be hard to schedule this guy, but I'm sure you visit San Francisco."
Yeah, yeah.
And obviously you just DM me on the day of, and you go like, "Let's set something up."
Yeah, yeah. Well, you know, uh, the, the folks at Tao gave me this hat, and then they mentioned you, and I was like, "Oh yeah, we talked about something." And then, uh, you know, you mentioned from Taiwan. I- you didn't see this.
I was talking to Swyx about this, but this is a, uh, mooncake-
Oh, nice
... uh, from Taiwan that I brought back. So you know, hopefully you'll enjoy that.
Nice. Thank you.
These are amazing.
Mm-hmm. Um, so you are the author of the extremely popular SemiAnalysis blog. Um, we have both had, uh, a little bit of credentials or claim to fame in breaking details of GPT-4. Uh, George Hotz came on our pod and, uh, talked about the mixture of experts thing, and then you, you had a lot more detail than-
Well, to be clear, I talked about mixture of experts in January. It's just people didn't really notice it-
Yeah
... I guess. I don't know.
Uh, you, you went into a lot more detail, and, and I'd love to dig into some of that. But anyway, so welcome and, uh, congrats on all your success so far.
Yeah. Thank you so much. You know, it's, uh, it, it, it's really interesting. Like I've been, I've been doing consulting in the industries, semiconductor industry since '17, 2017 and like, uh, you know, 2021 got bored, and in November I started writing a blog and then like 2022 I was good and, uh, started hiring folks at, for my firm and then all of a sudden 2023 happens and it's like the perfect intersection 'cause I used to do data science but not like AI, not really like multi-variable progression is not AI, right?
Um, but also I've been p- I've been involved in the semiconductor industry for a long, long time, uh, posting about it online since I was 12, right? Uh, and so it was like the perfect like time and place 'cause like semiconductors became important, right?
Uh, you know, all of a sudden it wasn't like this boring thing and then also the shortage, you know, in 2021 also mattered. But like, you know, all of a sudden this all kind of came to fruition, so it's cool to have the blog sort of, uh, uh, blow up in, uh, that way.
Uh, I used to cover semis at Beverly as well. Um, and it was, for a long time it was just the mobile cycle-
Mm-hmm
... and then a little bit of PCs, but like not that much. Um, and then I, I, maybe some cloud stuff, you know, like public cloud, uh, you know, semiconductor stuff. But it really wasn't anything until this, this wave.
And I was, I was actually listening to you on one of the previous podcasts that you've done and, and it was surprising that high performance computing also kinda didn't really take off. Like this is like AI is just the first form of high, high performance computing that worked.
I, I-- One of the theses I've had for a long time that I think people haven't really, uh, caught on, but it's really, really coming to fruition now is that, uh, the largest tech companies in the world, their software is important, but actually having and operating a very efficient infrastructure is incredibly important.
And so, you know, people talk about, you know, "Hey, Amazon is great for-- AWS is great because yes, it is easy to use and they built all these things." But behind the scenes, and no one really talks about it that much, but it's like behind the scenes they've done a lot on the infrastructure that is super custom that, uh, Microsoft Azure and Google Cloud just don't even match in terms of efficiency.
Like if you think about the cost to rent out, uh, SSD space or the cost to rent, you know, offer a database service on top of that obviously. A cost to rent out a certain level of CPU performance, um, Amazon has a massive advantage there.
And likewise, like Google spent all this time doing that in AI, right? With their TPUs and infrastructure there and, and optical switches and all this sort of stuff. And so like in the past it wasn't immediately obvious, but it-- I think, I think with AI especially like with like how scaling laws are going, it's like incredibly important for, you know, infrastructure is like so much more important.
And then like when you just think about software cost, right? Like the cost structure of it. Um, there is, there is always a bigger component of R&D and like SaaS businesses, you know, all over SF, right? Like, uh, you know, all these SaaS businesses did crazy good 'cause you know, they just start as they grow and then all of a sudden they're so freaking profitable for each incremental new customer.
Um, and AI software looks like it's gonna be very different in my opinion, right? Like the R&D cost is much lower in terms of people, uh, but the cost of goods sold in terms of actually operating the service, I think will be much higher, right?
And so in that same sense, infrastructure matters a ton, uh, for that.
GPU Poor vs Rich4:31
I think you wrote once that training costs effectively don't matter.
Yeah. In, in my opinion, I think that's a little bit spicy. But yeah, it's like training costs are irrelevant, right? Like GPT-4, right? Like 20,000 A100s. That's, that's ch- like I know it sounds like a lot of money.
It's 12 million all-in, 5 million all-in. Is that a reasonable estimate?
Um, yeah. I think, I think for the, the supercomputer it's, it's a-- it's slightly more, but yeah, I think the 500 million is a fair enough number. I mean, if you think about just the pre-training, right? Three months, 20,000 A100s at, you know, at a dollar an hour is like, that is way less than 500 million, right?
But of course there's data and all this sort of stuff.
Yeah. So people that are watching this on YouTube, they can see a GPU poor and a GPU rich hat on the table, uh, which is inspired by your-
Yeah, that's you with that. Yes.
Yeah, yeah. Your Google Gemini Eats the World, um, blog post. One, did you know that this thing was gonna blow up so much? Sam Altman even tweeted about it. He said, "Incredible Google got the Sem- SemiAnalysis guy to publish their internal marketing recruiting chart."
Uh, and yeah. Tell people who are the GPU poor, who are the GPU rich? Like what's this framework that they should think about?
So it's, it's, um, you know, some of this work we've been doing for a while is just on infrastructure and like, hey, like when something happens, I think it's like a, you know, sort of competitive ad- advantage of, of our firm, right?
Me, myself and my colleagues is like we go from software all the way through to like low-level manufacturing and it's like, who... You know, oh, Google's actually ramping up TPU production massively, right? Um, and like I think people in AI would be like, "Well, duh," but like, okay, like who, who has the capability of figuring out the number?
Well, one, you can just get Google to tell you, but they, they don't, they won't tell you, right? That's like a very closely guarded secret and most people that work at Google DeepMind don't even know that, know that number, right?
Um- Two, you go through the supply chain and see what they've placed in orders, right? Um, but then you, you know, three is sort of like, well, who's actually winning from this, right? Like, "Hey, oh, Celestica's building these boxes.
Wow. Oh, interesting. Okay, um, you know, uh, this company's involved in testing for them. Oh, okay. Oh, this company's providing design IP to them. Okay. Okay." Like, that's, like, like, you know, very valuable in a monetary sense. Um, but, you know, you have to understand the whole technology stack.
But on the flip side, right, is like, well, why is Google building all these? Uh, what, what could they do with it? And, and, and what does that mean for the world? And the state of the world is like, especially in SF, right?
Like, I'm sure you folks have been to parties, um, people just brag about how many GPUs they have. Like, I've-- it's happened to me multiple times where somebody's just like, I'm, I'm just witnessing a conversation where somebody from Meta is bragging about how many GPUs they have versus someone from another firm.
Uh, then it's like or like a startup person's like, "Dude, can you believe we just acquired-- We have five hundred and twelve H100s coming online in, in August." And it's like, "Oh, cool." Like... But then you're like, you know, going through the supply chain, it's like, " Dude, do you realize there's four hundred to five hundred thousand being manuf- m- four hundred thousand manufactured last quarter and, like, five thirty thousand this quarter being sold, right, of H100s?"
And it's like, oh, crap, that's, that's a lot. Um, you know, so sort of like that's a lot of GPUs, but then like, oh, how, how does that compare to Google? And like, there's one way to look at the world, which is just like, hey, scale is all you need.
Like, obviously, data matters. Obviously, all this stuff matters. But given, given any data set, a larger model will just do better, I think, is... It's gonna be more expensive, but it's gonna do better. There's the, there's the view of like, okay, there's all these GPUs going to production.
Nvidia is gonna sell well over three million, you know, total GPUs next year. Um, you know, over a million H100s this year alone, right? You know, there's, there's a lot of GPU capacity coming online. It's, it's an incredible amount.
And, and like, well, what are people doing? What are people working on? Um, I think it's very important to, like, just think about what are people working on, right? Um, some, you know, uh, what, what, what actually are you building that's, that's gonna advance, you know, what is monetizable, but what also makes sense?
And so, like, a lot of people were doing things that I thought felt counterproductive. In a world where in less than a year there's gonna be more than four million, you know, high-end GPUs out there, you know, we can talk about the concentration of those GPUs, but if you're doing really valuable work as a, as a good person, right, like you're, you're contributing in some way, should you be focused on like, well, I don't have access to any of those four million GPUs, right?
I actually only have access to gaming GPUs. Should I focus on, like, being able to fine-tune a model on that, right? Like, eh, it's not really that important. Or like, should I be focused on batch one inference on a cloud GPU?
Like, no, that's like pointless. Like, why would you do, why would you do batch size one inference on an H100? That's just, like, ridiculously dumb. There's a lot of counterproductive work. Um, and at the same time, there's a lot of like, you know, things that people should be doing.
And so like, you know, kind of you can tier the world into like, hey, like, I mean, obviously most people don't have resources, right? And I love the open source, and I want the open source to win, and I hate the people who want to like, you know, just like, "No, we're, we're, we're, uh, X lab, and we think this is the only way you should do it, and if people don't do it this way, they should be regulated against it," and all this kind of stuff.
I hate that attitude. Uh, so I want the open source to win, right? Like companies like Mistral and like what Meta are doing, you know, Mosaic and, and you know, all these folks together, blah, blah, blah, right? Like all these people doing, you know, huge stuff with open source, you know, want them to succeed.
But it's like there's certain things that are, you know, like fo- hyper-focusing on leaderboards at Hugging Face, right? Like that's just like, no, like Trueful QA is a garbage benchmark. Like you, you, you, you, you, you-- like some of the models that are very high on there, if you use it for five seconds, you're like, "This is garbage," right?
And it's just like you're gaming a benchmark. So it's like there, there's, there was things I wanted to say. Also, you know, we're in a world where compute matters a lot. Google is gonna have more compute than any other company in the world, period, by, by like a large, large factor.
And so like it's just like framing it into that like, you know, mindset of like, hey, like what are the counterproductive things? What do I think personally, or what have people told me that are involved in this should they focus on?
Right? Um, you know, and, and what is the world where like, you know, hey, we're, we're doing the pace of acceleration, you know, from 2020 to 2022 is less than 2022 to 2024, right? Like we are growing, you know, GPT-2 to 4, 2 to 4 is like 2020 to 2022, right?
Is, is less than I think from GPT-4 in 2022, which is when it was trained, right, to what, what, what OpenAI and, and Google and, and, and Anthropic could do in 2025, right? Like I think the pace of acceleration is, is increasing.
Um, and, and it's just good to like think about, you know, that sort of stuff. I don't know, I don't know where I'm rambling with this, but-
Yeah, yeah. No, that makes sense. And the, the chart that Sam mentioned, uh, is about, yeah, Google TPU v5's completely overtaking OpenAI by like orders of magnitude. Let's talk about the TPU a bit. We had, um, Chris Lattner on the show, which I know you know.
Um, he used to work on TensorFlow and Google, and he did mention that the goal of Google is like make TPUs go fast with TensorFlow. But then he also had a post about PyTorch kind of stealing the, the thunder, so to speak.
Uh, how do you see that changing if like now that a lot of the compute will be TPU-based, um, and Google wants to offer some of that to the, to the public too?
I mean, Google internally, and I think, you know, is, is obviously on Jax and XLA and all that kind of stuff, right? But, uh, externally, like they've done a lot-- a really good job. Like, I mean, I wouldn't, uh, you know, I wouldn't say like TPUs through PyTorch XLA is amazing, but it's, it's not bad, right?
Like some of the numbers they've, they've shown, some of the, you know, code they've shown for TPU v5e, which is not the TPU v5 that I was referring to, which is, uh, in, in the sort of the post, the GPU-poor post I was referring to.
But TPU v5e is like Uh, the new one, but it's mostly, mostly an inference chip. It's a small chip. It's a little bit l-- it's about half the size of a TPU v5. That chip, you know, you can get very good performance on like, of, of Llama 70B inference, right?
Like, you know, very, very good performance. So like when you're using PyTorch and XLA. Now, of course, you're gonna get better if you go Jax XLA, but, uh, I think Google is doing a really good job after the restructuring, um, of focusing on external customers too, right?
Like, hey, like TPU v5e will probably won't focus too much on TPU v5 for everyone externally, but v5e, well, we're, we're also building a million of those, right? Hey, a lot of companies are using them, right? Uh, or will be using them because it's gonna be an incredibly cheap form of compute.
The world of like, you know, frameworks and all that, right? Like, that's obviously something a researcher should talk about, not myself. But you know, the, the stats are clear that PyTorch is way, way dominating everything. Um, but Jax is like doing well.
Like there's external users of Jax. But in the end, like there's the front end, right? The front end is like what... But, but you know, we're referring back to, or may-maybe it's something we do later and you guys are gonna edit after, but the layers of abstraction, right?
Like, uh, you know, it-- for-- the forever shouldn't be that the person doing, you know, you know, like PyTorch level code, right, that high, should also be writing custom CUDA kernels, right? There should be, you know, different layers of abstraction where people hyper-optimize and make it much easier for everyone to innovate on separate stacks, right?
And then every once in a while, someone comes through and pierces through the layers of abstraction and, and innovates across multiple, um, or a group of people. But I think, you know, that's probably... You know, frameworks are important, but, you know, compilers are important, right?
Chris Lattner's-- what he's doing is really cool. Um, I, I don't know if it'll work, but it's, it's super cool, and it certainly works on CPUs. We'll see about accelerators. Um, likewise, there's OpenAI's Triton and like what they're trying to do there and like, uh, you know, everyone's really coalescing around Triton.
Uh, you know, people, you know, third-party hardware vendors. Uh, there's Palace, right? Uh, so I don't know if you've heard about that, but I, I don't wanna mischaracterize it, but you can write in Palace and, and it'll go through...
You can-- it-- lower level code, and it'll work to, uh, TPUs and GPUs, kinda like Triton, but like there's a backend for Triton. I don't know exactly everything about it, but I think there's a lot of innovation happening on make things go faster, right?
How do you, how do you go brr? Because every single person working in ML, ch-- you know, it would be a travesty if they had to write like custom CUDA kernels always, right? Like, that would just slow down productivity.
Mm-hmm.
Uh, but at the same time, you kinda have to.
Hardware Optimization14:23
Yeah. Excellent. Um, w-by the way, I like to quantify things when you say make things over. Like, uh, um, is there a target range of like MFU that you typically talk about?
Yeah, there's, there's sorta two metrics that I like to think about a lot, right? So in training, everyone just talks about MFU, right? But then on inference, right, which I think is, you know, one, LLM inference will be bigger than training or multimodal, whatever, bubble line inference will be bigger than training, you know, probably next year, in fact, um, at least in terms of GPUs deployed.
The other thing is like, you know, what, what's the bottleneck when you're running these models? So like the simple stupid way to look at it is training is, you know, there's six flops, uh, floating point operations you have to do for every byte you read in, right?
Every, every parameter you read in. So if it's FP8, then it's a byte. If it's FP16, it's two bytes, whatever, right? Um, but on training. But on inference side, the ratio is completely different. It's two to one, right?
There's two flops per parameter that you read in, and parameters may-maybe one byte, right? Because it's FP8 or int8, right? Eight bits, byte. And but then when you look at the GPUs, right? The GPUs are very, very different ratio.
The H100 has three point three five terabytes a second of memory bandwidth, and it has a thousand teraflops of FP16, B float 16, right? So that ratio is like... Well, I, I'm sorry, I-I'm gonna butcher the math here, and people are gonna think I'm dumb, but two fifty-six to one, right?
Call it two fifty-six to one if you're doing FP16. Uh, same applies to FP8, right? Because, uh, yeah. Anyways, uh, per pa-parameter, parameter read to, uh, number of floating point operations, right? If you quantize further, then you... But you also get double the performance on that lower quantization.
That does not fit the hardware at all, right? So if you're just doing LLM inference at batch one, then you're always gonna be underutilizing the flops. You're only paying for memory bandwidth. And the way hardware is developing, y-that ratio is actually only gonna get worse, right?
Uh, H200 will come out soon enough, which will help the ratio a little bit, you know, improve memory bandwidth more than it improves flops. Uh, just like the A-A100 eighty gig did versus the A100 forty gig. Uh, but then when the B100 comes out, the flops are gonna increase more than memory bandwidth.
Um, and when future generations come out and the same with AMD side, right, MI300 versus 400. As you move on generations, just due to fundamental like semiconductor scaling, DRAM, uh, memory is not scaling as fast as logic has been.
Uh, and so you're, you're gonna continue to ha-- uh, and, and you can do a lot of interesting things on the architecture. So, so you're gonna have this problem get worse and worse and worse, right? And so on training, it's very, you know, who cares, right?
Because my flops are still my bottleneck mostly. I mean, memory bandwidth is obviously a bottleneck, but like... Well, I-- you know, batch sizes are freaking crazy, right? Like, people will train like two million batch size is trivial, right?
Like, that's what Llama, I think, did. Llama 70B was two million batch size. And like, you talk to someone at one of the frontier labs, and they're like, "Ha." Right? "Just two million?" Right? Uh, two million token batch size, right?
That's crazy. Or sequence, sorry. But when you go to inference side, it's like, well, it's impossible to do-- one, to do two million batch size. Also, your latency would be horrendous if you tried to do something that crazy, right?
Um, so you kinda have this like differing problem where on, on training, everyone just kept talking MFU, model flop utilization, right? How many flops? Six times the number of, uh, parameters, basically, more or less. Um, and then what's the quoted number, right?
So if I have three-three hundred and twelve teraflops out of my A100, and I was able to achieve two hundred, that's all. That's really good, right? You know? Um, some people are achieving higher, right? Some people are achieving lower.
That's a very important like metric to think about. Now you have like people thinking MFU is like a security risk. Uh, but on inference, MFU is not nearly as important, right? It's, it's memory bandwidth utilization. You know, batch one is, is, is, you know, what memory bandwidth can I achieve, right?
'Cause if as I increase batch from batch size one to four to eight to, you know, even two fifty-six, right? It's sort of where the crossover happens on an H100, uh, inference-wise, right? Where, where it's flops limiting you more and more.
But like you, you should have very high memory bandwidth utilization. So when people talk about A100s, like sixty percent MFU is decent, right? On H100s, it's more like forty, forty-five percent because the flops increased more than the memory bandwidth.
Uh, but people over time will probably get above fifty percent on, on H100 on MFU, on training. But on inference, it's not being talked about much, but, uh, mod, B-MBU, mo-me, uh, model bandwidth utilization is the important factor, right?
So if my three point three ter-five terabytes a second of memory bandwidth, I have on my H100, can I get two? Can I get three? Right? That's the important thing. And, and right now, if you look at everyone's, uh, you know, inference stuff...
So I, I dogged on this in the GPU 4 thing, right? But it's like Hugging Face's libraries are actually very inefficient. Like, incredibly inefficient for inference. Uh, you get like fifteen percent MBU, um, on, on, on, on some configurations, like A-A-A- eight A100s and Llama seventy B.
You get, like, fifteen percent, which is just, like, horrendous because at the end of the day, your latency is derived from what memory bandwidth you can effectively get, right? Um, you know, so if, if, if you're doing Llama seventy billion, seventy billion parameters.
If you're doing it int eight, okay, that's seventy, uh, gigabytes a second... Uh, gigabytes you need to read for every single inference, every single forward pass plus, you know, the attention. But, you know, again, we're simplifying it. Seventy gigabytes you need to read for every forward pass.
What is an acceptable latency for a user to have? Um, I would argue, you know, thirty milliseconds per token. Um, some people would argue lower, right? But at the very least, you need to achieve human reading level speeds and probably a little bit faster because we like to skim to have a usable model for, uh, chatbot-style applications.
Now, there's other applications of course, but chatbot-style applications, you want it to be, uh, human reading speed. So thirty tokens per second. Thirty tokens per second is thirty-three... Or, uh, thirty tokens... Milliseconds per token is thirty-three tokens per second, times seventy is, um...
Let's say three times seven is twenty-one, and then add two zeros. So twenty-one hundred gigabytes a second, right? To achieve human reading speed on Llama seventy B, right? So one, you can never achieve Llama seventy B human reading speed on...
Even if you had enough memory capacity on a model on, on an A100, right? Even in, even an H100 to achieve human reading speed, right? Of course you couldn't fit it because it's eighty gigabytes versus seventy billion parameters, so you're, you're, you're kind of butting up against the limits already.
Um, seventy billion parameters being seventy gigabytes at in-int eight or FP8. Uh, you, you end up with, one, how do I achieve human reading level speeds, right? So if I go with two H100s, then now I have, you know, call it six terabytes a second of memory bandwidth.
If I achieve just thirty milliseconds per token, then I'm... You know, which is thirty-three tokens per second, which is two point one terabytes a second of memory bandwidth, then I'm only at, like, thirty percent m- bandwidth utilization. So I'm not using all my flops on batch one anyways, right?
Because seventy... You know, the, the, the flops that you're using there is tremendously ri- low relative to inference, and I'm not actually using a ton of, uh, the tokens on inference. So if with two H100s I only get thirty milliseconds a token, that's a really bad result.
You should be striving to get, you know, so upwards of sixty percent and that's like sixty percent is kind of low too, right? Like I've, I've heard people getting seventy, eighty percent model bandwidth utilization. You know, obviously you can increase your batch size from there and, and your model bandwidth ut-utilization will start to fall as your flops utilization increases.
But you know, there... You have to pick the sweet spot for where you want to hit on the latency curve for your user. Uh, obviously as you increase batch size, you get more throughput per GPU, so that's more cost effective.
There's a lot of like things to think about there, but I think those are sort of the two main things that people want to think about. Uh, and there's obviously a ton with regards to like networking and inter-GPU connection because, uh, most of the useful models don't run on a single GPU.
They can't run on a single GPU.
Is your TPU equivalent of Mellanox?
Networking21:39
Uh, so the TPU's... Uh, so, so the Google TPU is like super interesting because Google's been working with Broadcom, who's the number one networking company in the world, right? So, uh, they, they... Mellanox was nowhere close to number one.
Uh, they o- they only... They were number t- they were... They had a niche that they were very good at, which was the network card, the, the card that you actually put in the server, but they didn't do much.
They didn't ha- They weren't doing successfully in the switches, right? Which is, you know, you connect all the network's cards to switches and then the switches to all the G- uh, all the n- uh, you know, servers. So Mellanox was not that great.
I mean, it was good. They were doing good, and Nvidia bought them, you know, in '19, I believe, or '18. Um, but Broadcom has been number one in networking for decade plus, right? Um, and Google partnered with them on making the TPU, um, and they've, you know, TPUv, you know, all the way through to all TPUv5, which is the one they're in production of now, and uh, six and, you know, all these.
These are all gonna be, uh, you know, co-designed with, with Broadcom, right? So Google does a lot of the design, especially on the ML hardware side, on how you pass stuff around internally on the chip. But Broadcom does a lot on the network side, right?
They specifically, you know, how to get really high connection speed between two chips, right? They've do- they've done a to-ton there, and obviously Google works ton there too. But, uh, this is sort of like Google's like le-less discussed partnership that's truly critical for them.
And why Goo-Google's tried to get away from them many times. Their latest target to get away from Broadcom is 2027, right? But like, you know, that's, that's four years from now. Chip design cycle is four years. So, um, they already tried to get away in 2025 and that failed.
Uh, but, but yeah, they, they had this equivalent of very high-speed networking. It works very differently, um, than the way GPU networking does, and, and that's important for people who code on a lower level.
I've seen this described as like the ultimate limit on how big models they build. It's not flops, it's not memory, it's n-networking. Uh, like it has the, it has the lowest scaling law. It's like the lowest Moore's law.
So hard to outperform them and I don't know what to do about that because no one else has any solutions.
Yeah, yeah. So I think, I think what you're referring to is that like network speed has increased slower-
Much slower than the other two
... than, than flops. Yeah. And, and bandwidth. Yeah, yeah. Um, and yeah, that's a tremendous problem in the industry, right? Uh, but, but like that's why, that's why Nvidia bought a networking company.
Yeah.
That's why Broadcom is, is, is, is working with, uh, on Google's chip right now, but of course on Meta's, Meta's internal AI chip, which they're on the second generation of working on that. And what's the main thing that, uh, Meta's doing interesting is networking stuff, right?
Multiplying tensors is kind of, you know, anyone can... There's a lot of people who've made good matrix multiplier units, right? But it's about like getting good utilization out of those and interfacing with the memory and interfacing with other chips really efficiently makes designing these chips very hard.
And most of the startups obviously have not done that really well.
GPU Poor Strategy24:20
Yeah, I mean, I think the startups point is the most interesting, right? You mentioned companies, uh, that are GPU poor, they raise a lot of money and there's a lot of startups out there that are GPU poor and did not raise a lot of money.
What should they do? H-how do you see like- The space dividing. Are we just supposed to wait for, like, the, the big labs to do a lot of this work with a lot of the GPUs? Like, w-what's like the GPU-poor's beautiful version of the, of the article?
Like, the whole point was that, like, Google... You know, OpenAI, who everyone would be like, "Oh yeah, they have more GPUs than anyone else," right? But they have a lot less FLOPS than Google, right? That was the point of the, like, thing.
And it was... But not just them, it's like, okay, you, you, you know, it's like a relative totem pole, right? Now, of course, Google doesn't use U-GPUs as much for training and, and inference. They, they do use some, but mostly TPUs.
So kind of like the whole point is that everyone is GPU-poor because we're gonna continue to scale faster and faster and faster and faster. And, and, and compute will always be a bottleneck, just like data will always be a bottleneck, right?
You can have the best dataset in the world and you can always have a better one. And same with, you can have the biggest compute system in the world and you can al- but you'll always want a better one.
You know, like Mistral, right? Like, they trained a fricking awesome model on relatively fewer GPUs, right? And, and now they're scaling up higher and higher and higher, right? You know, there's, there's a lot that the GPU-poor can do though, right?
Like, hey, um, we all have phones, we all have, uh, laptops, right? Uh, there is a world for running GPUs or models on device, right? Um, you know, the Replit folks are, you know, trying to do stuff like that.
Their models can't be that... They can't follow scaling laws, right? Why? Because there is a fundamental limit to how much memory bandwidth and capacity, uh, that you can get on a laptop or a phone, right? Um, you know, I mentioned the ratio of FLOPS to bandwidth on a GPU is actually really li- really good compared to like a MacBook or like a phone.
Hey, to run Llama seventy billion requires two terabytes a second of memory bandwidth, two point one, at reading, human reading speed. Yeah, but my phone has like fifty gigabytes a second, right? Um, your, your laptop, even if you have an M1 Ultra, has what, like, I don't remember, like couple hundred gigabytes a second of memory bandwidth.
You can't run Llama 70B just by doing the classical thing. So there's like, there's stuff like speculative decoding and then, uh, you know, Together did something really cool and they put it in open source of course, Medusa, right?
Like things like that that are... You know, they work on batch size one, they don't work on batch size, you know, high. Um, and so there's like the world of like cloud inference. And so in the cloud it's all about, you know, what memory bandwidth and MFU I can achieve.
Whereas on the, on the edge, um, I don't think Google's gonna deploy a model that I can run on my laptop to help me with code or help me with, you know, XYZ. They're always gonna wanna run it on cloud for control.
Or maybe they let it run on the lap- on the, on the device, but it's like only their Pixel phone, you know, kinda like a walled garden thing. There's, there's obviously a lot of reasons to do other things for security, for, for, you know, openness to not be at the, uh, when- whims of a trillion dollar plus company who wants my data, right?
Like, you know, there's a lot of stuff to be done there and I think like, like folks like Replit are, are like... You know, I love it, right? That's exactly like the stuff, you know... I don't, I don't...
They open sourced their model, right? Um-
Yeah.
Yeah. Yeah, they... So they open sourced their model. I think, you know, things like what, what, like Together, what I just mentioned, right? That, that developing Medusa, right? That, that didn't take much GPU at all, right? That's... They're very G- They...
While they do have quite a few GPUs, they made a big announcement about having four thousand H100s. That's still relatively poor, right? When we're talking about hundreds of thousands of like the big labs, uh, like OpenAI and, and so on and so forth.
Um, or millions of TPUs like Google. Um, but you know, still they were able to develop Medusa with probably just one server, right? One server with eight GPUs in it. And its usefulness of something like Medusa, something like speculative decoding is, is on device, right?
And that's what like a lot of people can focus on. You know, people can focus on all sorts of things like that. I don't, I don't know, right? Like w- a new model architecture, right? Like are we only gonna use transformers?
I'm pretty pilled to think like transformers are it, right? Like just because like my hardware brain can only know something that loves hardware, right? Um- But like, so like, y- you know, like, you know, people should continue to try and innovate on that, right?
Like, you know, asynchronous training, right? Like that kind of stuff is like, you know, super, super interesting. Like Tim Demeters.
Distributed.
Yeah, yeah, distributed, like not in one data center.
Yeah.
Um, I think it's Tim Demeters. He had like the swarm-
Swarm papers.
Yeah, there you go. Sorry. Sorry.
Same guy as Aurora. The-
Yes, he had the swarm paper and Petal and... Well, I think Petal is, is whatever. Um, you know, like that research is super cool.
It's like steady at home, right? It's not gonna-
Yeah, I mean, yeah, but like I like research, that kind of stuff, right? Like, you know, hey, like the universities will never have much compute, but like, hey, you know, to prepare to do things to, you know, all these sorts of stuff, like they should try to build, you know, super large models.
Like if you look at what Tsinghua University is doing in China, like actually they open sourced their model to I think the largest like by parameter count at least open source models.
Which one? Kanku?
Uh, I, I don't remember the name.
MOE. Yeah.
Yeah. It's from Tsinghua University though, right?
Yeah, yeah. I think it's, uh, it was like a one point seven trillion.
Yeah. I mean, of course they didn't train it on much data, but it's like, you know, it's like still, like you could do some still cool stuff like that. I don't know. I think there's a lot that people can focus on, uh, because, you know, one, scaling out a service to many, many users, distribution's very important, so figuring out distribution, right?
Like, uh, figuring out useful fine tunes, right? Like, um, you know, doing LLMs that, you know, OpenAI will, OpenAI will never make, you know, sorry for the crassness, a porn DALL-E 3, right? But open source is doing crazy stuff with stable diffusion, right?
Right? Like I don't, I don't mean to... Yeah, but it's like, it's like... And there is a legitimate market. I think there's a couple companies who make tens of millions of dollars of revenue from, from-
Porn.
From, yeah, yeah. From, from LLMs or diffusion models for porn, right? Or, or you know, that kind of stuff. Like I mean there's a lot of stuff that people can work on that will be successful businesses or doesn't even have to be a business but could advance humanity tremendously that doesn't require crazy scale.
How do you think about the depreciation of like the hardware versus the models? Like we covered-
Two years. Two years usually.
Yeah. Like we covered open models for a while. If I think about the episodes we had like in March with like MPT 7B-
Oh yeah, nobody talks about that anymore.
E-exactly. It's like the depreciation is like three months, you know? It's like-
Well, I mean, no one should be talking about Llama thirteen billion anymore-
Yeah, exactly
... because Mistral just showed them up, right?
Yeah. So I, I'm really curious 'cause like, you know, if you buy a H100, sure, the next years is gonna be better, but like at least the hardware's good. If you're spending a lot of money on like training a smaller model, like it might be like super obsolete in like three months and you got now all this compute Coming online.
Um, I'm just curious if, like, companies should actually spend the time to like, you know, fine-tune them and like work on them where like the next generation is gonna be out of the box so much better, you know?
Unless, unless you're fine-tuning for on-device use, I think fine-tuning current existing models, especially the smaller ones, is a useless waste of time, right? Because the cost of inference is actually much cheaper than you think once you achieve good MBU and you batch at a decent size, which any successful business in the cloud is gonna achieve.
Um, you know, and then, and then two, um, fine-tuning, like people are like, "Oh, you know, this seven billion parameter model, if you fine-tune it on a dataset is, is almost as good as 3.5," right? It's like, yeah, but why don't you just fine-tune-
Use 3.5.
Yeah. Why don't you fine-tune 3.5 and look at your performance, right? And, and like there's nothing open source that is anywhere close to 3.5 yet, right? There will be. There will be. I, I think, I think, um, people also don't, don't quite grasp-
Falcon was supposed to be Falcon 140B.
Uh, it's, it's less parameters than 3.5 and, and also, uh, I don't know about the exact token count, but I believe it's less than-
Do we know the parameters of 3.5?
Um, it's not 175 billion.
We know, right? We know it's-
We'll keep saying this. No
... 'cause we know it's three, but we don't know 3.5.
3.5-
It's definitely smaller.
No, it's bigger than 175, but it's... I think it's sparse. I think it's, I think it's-
Yeah.
You know, it's an MOE. I'm pretty sure it's, uh... Um, you know, you can, you can do some like gating around the size of it by looking at their inference latency. Um-
Which is also, you will get upper bounds.
Yeah. You can look at like, well, what's the theoretical bandwidth if they're running it on this hardware and, uh, you know, um, you know, and, and doing tensor parallel in this way so they have this much memory bandwidth and maybe they get...
Maybe they're awesome and they get 90% memory bandwidth utilization. I don't know. That's an upper bound and you can see the latency that 3.5 gives you like, uh, especially at like off-peak hours or if you do fine-tuning and you have your, if you have a private enclave, they'll, like Azure will quote you latency.
So you can, you can figure out how many parameters per forward path, uh, which, which I think is somewhere in the like 50 to 40 billion range, but I could be very wrong. That's just like my guess based on that sort of stuff.
Um, you know, 50-ish. Uh, but then the-
16 experts or? I, I don't-
I have no clue. I have no clue. Uh-
There's no way to figure that out 'cause just stop routing.
Yeah, yeah. There, there actually there, there's, there's, there's, uh, there's someone I've talked to at, at one of the labs who like thinks they can figure out how many experts are in a model by, by querying in a crap load but I...
But that's only if you have access to the logits, the like the, the percentage-
Oh, yeah
... chance, yeah, before you do the softmax. Uh, I don't know. It... But yeah, there's like a ton of like competitive analysis you could try to do. But anyways, I think, I think open source will have models of that quality, right?
I think like, you know... I mean, I, I assume Mosaic or like Meadow will open source and Mistral will be able to open source models of that quality. Um, now furthermore, right, like if you just look at the amount of compute, obviously data is very important and the ability...
All these tricks and dials that you turn to be able to get goodMFU and good MBU, right? Like depending on inference or training is, is there's a ton of tricks, but at the end of the day, like there's like 10 companies that have enough compute in one single data center to be able to beat GPT-4, right?
Yeah.
Like straight up. Like to- if not today, within the next six months.
Yeah.
Right? Like 4,000 H100s is, is... I, I think you need about 7,000 maybe and with some like good like, uh, with, with, and with some algorithmic improvements that have happened since GPT-4 and some of the, some data quality improvements probably.
Like you could probably get to even like, you know, less than 7,000 H100s running for three months to beat GPT-4. Uh, but of course, that's gonna take a really awesome team. Um, but you know, there's quite a few companies that are gonna have that many, right?
Open source will, will match GPT-4, but then it's like what about GPT-4 Vision or what about, you know, uh, 5 and 6 and, you know, all these kind of stuff and like interact tool use and DALL-E and like that's the other thing is like there's a lot of stuff on tool use that the open source could also do, um, that the, the GPU poor could do.
Um, I think there are some folks that are doing that kind of stuff, agents and all that kind of stuff. I don't know. Uh-
That's way over my head, the agent stuff. Yeah. It's over everyone's head. Um, one more question on just like the, the, the sort of Gemini GPU-rich, uh, essay. We've, we've had a very wide-ranging conversation already, so it's hard to categorize.
Um, but I have tried to look for the Mina Eats the World document.
Oh, it's not, it's not-
I was like, "No, no, no, no, no."
Yeah.
I will find this your article. You-
No, so, so Noam Shazeer-
You've read it.
Yeah. I've, I, I read it. So, so Noam Shazeer is like... I, I, I don't know. I think he's like-
The GOAT.
The GOAT. Yeah. I think he's the GOAT. Like obviously like-
In one year he published like-
Yeah, yeah, exactly. It's like, it's like all this stuff that we were talking about today was like, you know... And obviously there's other people that are awesome that were, you know, helping and all that sort of stuff. You know, ju-just to be clear, but-
Absolutely.
Um, there was a couple other papers but like, yeah, so like Mina Eats the World was basically he wrote an internal document, uh, around the time where Google had Mina, right? And Mina was, uh, their, you know, was one of their LLMs that like is a footnote in the history.
Like, you know, most people will not like think about Mina as relevant. Um, but it was like he wrote it and he was like basically predicting everything that's happening now, which is that like large language models are gonna eat the world, right?
In terms of, you know, compute and he's like, "The total amount of deployed flops within Google data centers will be dominated by large language models." And like back then, a lot of people thought he was like silly for that, right?
Like internally at Google. Um, but you know, now if you look at it it's like, oh wait, millions of TPUs. You're right, you're right, you're right. Okay. Uh, we're totally d- getting dominated by like, uh, both, you know, Gemini training and inference, right?
Like, you know, like or whatever, two, three, four plus, plus one, two, three, four for Gemini and all these other things like that's, yeah, total flops being dominated by LLMs was completely right.
So my question was, was he had a bunch of predictions in there. Do you think there are any like underrated predictions that, um, may not have yet have come, come true but you're kind of moving-
I think, I think like if, you know, obviously, uh, I've read the document but I read it on someone else's device. They didn't send it to me so I can't really send it, sorry. Um, and also they, they, they, uh, they were okay with me talking about the document and calling Noam a GOAT because they also think Noam is a GOAT.
Um, uh, but I think, I think like, you know, now most everybody is like scaling law pilled and like, uh, LLM pilled and like, you know, all this sort of stuff and like it's a very clear line of sight.
Was he wrong on anything?
I, I mean like- Mina sucked, right? Like I mean, it was great for the time, right? But like-
It's like billion parameter model. I mean-
I, I don't remember off top of my head, but it's like if you look at the total flops, right? Uh, you know, parameters times tokens times six, right? It's like, it's like a tiny, tiny fraction of GPT-2, which came out just a few months later, which is like, okay, so he wasn't right about everything, but like maybe he knew about GPT.
I have no clue. But he, you know, OpenAI clearly was like way ahead of Google on LLM scaling even then, right? Um, it's just p-people didn't really recognize it back in GPT-2 days maybe, or the peop- the number of people that recognized it was maybe hundreds-
Yeah
... tens, right? I don't know.
You mentioned transformer alternatives. The other thing is GBU alternatives. So the TPU is obviously one, but there's Cerebras, there's Graphcore, there's Madax, Lumerian Labs. There's a lot of them. Thoughts on what's real, who's alive, who's kind of like a zombie-
Um
... company walking?
AI Hardware Startups38:03
So if, if you, if you go back and like, you know, I, you know, mentioned like transformers were the architecture that went out. But I think, you know, the number of people who recognized that in 2020 was, you know, as you mentioned, probably hundreds, right?
You know, for natural language processing, maybe in 2019 at least, right? You think about a chip design cycle, it's like years, right? You know, so, so it's kinda hard to bet your architecture on the type of model that develops.
Uh, but what's interesting about all the first wave AI hardware startups is you kinda have, you know, this ratio of, of memory, capacity, uh, compute, and memory bandwidth, right? Um, and so everyone kind of made the same bet, which is I have a lot of memory on my chip, which is A, really dumb because the models have grew way past that, right?
Even Cerebras, right? I mean, uh, you know, like I'm talking about like Graphcore, uh, it's called SRAM, which is the memory on chip. Um, much lower density, but much higher speeds versus, you know, DRAM, which is the memo- you know, memory off chip.
Um, and so everyone was betting on, you know, pretty much more memory on chip and less memory off chip, right? And, and, and if that... And, and to be clear, right, for image networks and, and models that are small enough to just fit on your chip, that works.
That, that is the, that is the superior architecture. But, you know, scale, right? Scale, scale, scale, scale. So, so Goo- NVIDIA was the only company that bet on the other side of, of more memory bandwidth, right, um, and, and more memory capacity external, external- uh, also the right ratio of memory bandwidth versus capacity, right?
'Cause there were people, a lot of people like Graphcore specifically, right? They had ton of memory on chip, and then they had a lot more memory off chip, but that memory off chip was a much lower bandwidth. Uh, same applies to SambaNova, same applies to, uh, Cerebras.
Uh, you know, they had no memory off chip, but they, they thought, "Hey, I'm gonna make a chip the size of a wafer," right? Like, I can fit, I can fit... You know, fine. You know, those guys, they're, they're, they're silly, right?
Hundreds of megabytes? We have forty gigabytes. There's no way, you know. And then, oh, crap, models are way bigger than forty gigabytes, right? Everyone bet on sort of the left side of this curve, right? Um, the interesting thing is that these, these new age startups, um, like Lumerium, like Madax, I won't get into what they're doing, but they're making much more rational bets.
I don't know. You know, it's, it's hard to say with a startup like it's gonna work out, right? Uh, obviously there's tons of risk embedded. Uh, but tho-those folks like, like, you know, Jay Dawani of Lumerium and like, uh, you know, M-M-Mike and, uh, and, and, and, and Renier, they, they, they understand models.
They understand how they work. Um, and, and if transformers continue to reign supreme, right, you know, they're-- whatever innovations those folks are doing on hardware are gonna need to be, you know, fitted for that. Or you have to predict what the model architecture is gonna look like in a few years, right?
What does it look like, right? Um, you know, and, and, and hit that spot correctly. So, so that's kind of a background on those. But like now you look today, it's like, hey, um, you know, Intel bought, uh, Nervana, which was Naveen Rao's MosaicML's.
He started MosaicML and sold it to Databricks somewhere recently. He's obviously leading LLMs and stuff there, AI there. But, you know, that, the-- In-Intel bought that company from him and then shut it down and bought this other AI company and, and now that company is kind of, uh, you know, got new chips.
They're gonna release a better chip than the H100, uh, within the next quarter or so, right? Um, AMD, they have a GPU, MI300, that will be better than the H100 in a quarter or so. Now, that says nothing about how hard it is to program it, but at least hardware-wise, on paper, it's better.
Why? Because it's, you know, a year and a half later, right, than in the H100, or a year later than the H100, of course, and, you know, a little bit more time and all that sort of stuff. But, you know, they're at least making similar bets on memory bandwidth versus flops versus capacity, kinda following NVIDIA's, uh, lead.
Um, the questions are like, what is the correct bet for three years from now? How do you engineer that? Uh, and, and, and will those alternatives make sense? The other thing is, if you look at total manufacturing capacity, right, for this sort of bet, right, you need high bandwidth memory, you need HBM, and you need large five nanometer dies, you know, soon three nanometer or whatever, right?
Y-You need both of those components, and you need the supp- whole supply chain to go through that. We've written a lot about it. But, you know, to simplify it, NVIDIA has a little bit more than half, and Google has like thirty percent, right, through Broadcom.
So it's like the total capacity for everyone else is much lower, and they're all sharing it, right? Amazon's Train and Inferentia, Microsoft's in-house chip, and, you know, you go down the list and it's like Meta's in-house chip and also AMD and also all the...
So all of these companies are sharing like a much smaller slice. Uh, their chips are not as good, or if they are, they're-- even though they're, you know, I mentioned, y-you know, Intel and, uh, AMD's chips are better, that's only because they're throwing more money at the problem kind of, right?
You know, NVIDIA charges crazy prices. I think everyone knows that. Their gross margins are pro- are insane. Uh, AMD and, and Intel and, and, and others will charge more reasonable margins, and so they're able to give you more HBM and et cetera for a similar price, and so that ends up letting them beat NVIDIA, uh, if you will.
But their manufacturing costs are twice that in some cases, right? In the case of AMD, their manufacturing costs for MI300 are more than twice that of H100, and it only beats H100 by a little bit from, you know, performance stuff I've seen.
So it's like, you know, it's, it's tough for anyone to like bet the farm on a alternative hardware supplier, right? Like in my opinion, like you should either just like be like, you know, a lot of like ex-Google startups are just using- TPUs, right?
And, and hey, that's Google Cloud, you know, after moving the TPU team, you know, into the cloud team, infrastructure team, sort of they're much more aggressive on external selling. And so companies like-- You even see companies like Apple using TPUs for training LLMs, as well as GPUs.
But, um, you know, either bet heavily on TPUs because that's where the capacity is, bet heavily on GPUs of course, and stop worrying about it and, and leverage all this amazing open source code that is optimized for Nvidia.
Mm-hmm.
Or y- okay, if you do bet on AMD or, or Intel or, or on any of these startups, then you better make damn sure you're really good at low-level programming and damn sure you also have a compelling business case, uh, and that the hardware supplier is giving you such a good deal that it's worth it.
And also, by the way, Nvidia's releasing a new chip, you know, in-- You know, they're gonna announce it in March, and they're gonna release it, you know, and ship it, you know, Q2, Q3 next year anyways, right? And that chip will probably be three or four times as good, right?
And maybe it'll cost twice as much or fifty percent more. I, I hear it's three X the performance on an LLM w- and fifty percent more expensive is what I hear. So it's like, okay, yeah, like-
It's worth it
... n-nothing, nothing is gonna compete with that, even if it is fifty percent more expensive, right? And then you're like, "Okay, well, that kicks the can down further." And then Nvidia's moving to a yearly release cycle, so it's like very hard for anyone to catch up to Nvidia, really, right?
Are, are, you know, investing all this in other hardware? Like, if you're Microsoft, obviously who cares if I spend five hundred million dollars a year on my internal chip? Who cares if I spend five hundred million dollars a year on AMD chips, right?
Like, if it lets me knock the price of Nvidia GPUs down a little bit, puts the fear of God within Jensen Huang, right? Like, you know, then, then it is what it is, right? And, and likewise, you know, with, with Amazon and, you know, so on and so forth.
You know, of course, their hope is that their chips succeed or that they can actually have an alternative that is much cheaper than Nvidia. But to throw a couple hundred million dollars at a company, you know,'s product, um, is, is completely reasonable.
Um, and in the case of AMD, I think it'll be more than a couple hundred million dollars, right? But, uh, yeah, I think, I think alternative hardware is like-- It, it really does hit like, like sort of a peak hype cycle kind of end of this year, early next year, because all Nvidia has is H100 and then H200, which is just better, more mem-more memory bandwidth, higher memory capacity.
H100, right? But that doesn't beat what, uh, you know, AMD are doing, uh, doesn't beat what, um, you know, even Intel's Gaudi 3 does. But then very quickly after, Nvidia will crush them, and then those other companies are gonna take two years to get to their next generation.
You know, so it's g- it's just a really tough place, and, and no one decides... You know, the, the, the main thing about hardware is like, hey, that bet I talked about earlier is like, you know, that's very oversimplified, right?
Just memory bandwidth, FLOPs and, and, uh, memory capacity. There's a whole lot more bets. There's a hundred different bets that you have to make and guess correctly to get good hardware. Not even have better hardware than Nvidia, get close to them-
Mm-hmm
... right? And, and, and, and that takes understanding models really, really well. That takes understanding, you know, so many different aspects, whether it's power delivery or cooling or, uh, you know, design, layout, uh, all this sort of stuff, and it's like, how many companies can do everything here, right?
Manufacturing Constraints46:18
It's like I'd argue Google probably understands models better than Nvidia. I don't think people would disagree. Um, and Nvidia understands hardware better than, than Google. And so you end up with like, Google's hardware is competitive, but like does Amazon understand models better than Nvidia?
I don't think so. And does Amazon better hardware-- understand hardware better than N-Nvidia? No. Right? Like, it's like-
What about Anthropic's, um, investment? Or the, the investment in Anthropic. We'll see.
Yeah. I a-
Chance of that
... I'm, I'm also of the opinion that the labs are, um, they're useful partners, they're convenient partners, uh, but they are not gonna like-
Buddy up
... they're not gonna buddy up as close as people think, right? I don't, I don't even think like-- You know, I, I, I expect in the next few years that the OpenAI Microsoft probably falls apart too.
That'd be huge.
Um, I mean, I mean, they'll still continue to use GPUs and stuff there, but like-
Mm-hmm
... I think that the level of closeness you see today-
Yeah
... is probably the closest they get, right? Like, I mean-
At some point they become competitive, if OpenAI becomes its own cloud.
Yeah, I mean, I think, I think OpenAI w-wants to not just become a trillion dollar company, ten trillion dollar-- I mean, not, not a company, right? But like the, the level of the value that they deliver to the world, if you talk to anyone there, they truly believe it'll be tens of trillions, if not hundreds of trillions of dollars, right?
In which case, obviously, you know... I know, I know weird corporate structure aside, like, you know, this is the same like, like playing field as companies like Microsoft and Google. Like, like, like-
It's a-
... Google wants to also deliver hundreds of trillions of dollars of value, and it's like obviously you're competing, and Microsoft wants to do the same and you're gonna compete. Um, and like, yeah, I, I think, I think in general, right, like these, these lab partnerships are gonna be nice, but they're probably incentivized to, uh, you know, "Hey, Nvidia, you should...
You know, can you, can you design the hardware in this way?" And Nvidia's like, "No, it doesn't work like that. It works like this." And they're like, "Oh, so this is the best compromise," right? Like I think, I think, uh, I think OpenAI would be stupid not to do that with Nvidia, but also with AMD and, but also, hey, like how much ti-- And, and Microsoft's internal silicon, but it's like, how much time do you actually have, right?
Like, you know, should I do that? Should I spend all my, you know, super, super smart people's time and limited, you know, this caliber of person's time doing that? Um, or should they focus on like, "Hey, can I get like asynchronous training to work?"
Or like, you know, figure out this next multimodal thing? Or I don't know, I don't know, right? Like it's probably better, you know, "Hey, can I eke out five percent more MFU and work on designing the next supercomputer?"
Right? Like these kind of things, how much more valuable is that, right? So it's like, you know, it's, it's tough to, tough to, tough to see, you know, even OpenAI helping Microsoft enough to get their, uh, knowledge of models so, so, so good, right?
Like Microsoft's gonna announce their chip soon. Um, it's worse performance than the H100, uh, but the cost effectiveness of it is, is better for Microsoft internally just because they don't have to pay the Nvidia tax. But again, like by the time they ramp it and, and all these sorts of things and, oh, hey, that only works on a certain size of models.
Once you exceed that, then it's actually, you know, again, better for Nvidia. So it's like, it's really tough for OpenAI to be like, "Yeah, we, we wanna bet on, on Microsoft," right? Like, and hey, we have, you know- I don't know.
What's their number of people they have now? Like, 700 people? You know, of which how many do low-level code? Do I wanna have separate code bases for this and this and this and this? And, you know, it's like, it's just, like, a, a big headache to ...
I, I, I don't know. I think it'd be very difficult to see anyone truly pivoting to anything besides a GPU and a T- uh, TPU, especially if you have- if you need that scale, right? And, and that scale that the la- at least the labs, right, require is absurd, right?
Google, Google says millions, right, of TPUs. OpenAI will say millions of GPUs, right? Like, I truly do believe they think that, that number of, of next generation GPUs, right? Like, the numbers that we're gonna get to are, like ...
I bet you ... I, I mean, I don't know, but I bet Sam Altman would say, "Yeah, we're gonna build a $100 billion supercomputer in three years or two years," right? Like-
Easy
... and, and, like, after GPT-5 releases, if he goes to the market and says, like, "Hey, I wanna raise $100 billion at $500 billion valuation," I'm sure the market would give it to him, right? Like and then they build that supercomputer, right?
Like, I mean, like, I think that's, like, truly the path we're on. Um, and so it's hard to, hard to imagine. Yeah, I don't know.
One, one point that you didn't touch on, and Taiwan companies are famously very chatty about the fruit company, um, should we take Apple seriously at all in this game, or they're just in a different world altogether?
Apple and AI50:24
I don't know. I think, I think, you know, just from my view of Apple, I, I, I don't personally use Apple products. But every- I mean, like, my, my mom, I buy her a new iPhone every year, right?
Just to be clear, right? Like, yeah, no, Mom, you, you ... You know, new Apple Watch every couple years, right? Like, uh, of course.
Yeah.
Right? Um, so I, I, I respect their products, but, like-
Just-
... I don't think Apple will ever release a model that you can get to say, you know, really bad things, right? Or, you know, racist things or whatever, right? I don't think they can ever do that. But, like, you know, frankly, like, I'm sure, I'm sure OpenAI releasing, you know, 3.5 and 4 has had people, like, you know, break- jailbreaks for the, you know, kind of, uh, old terminology from iPhone, jailbreak the model and get it to do bad things, right?
Teach me how to u- make anthrax, right? Like, or, like, say these, like, hateful things, like race rank. Rank the races of the world, right? Like, you know, like, crazy. I mean, I've seen it on Twitter. I've seen all these three of these things, right?
Uh-
Like, it's like-
My, my grandma, my grandma wants to know like the-
Yeah
... how you get-
Yeah, yeah.
My grandma's dying. Please help.
Yeah.
Rank.
She needs the queer ... She needs to know how to make anthrax to live, right? Like, yeah. Um, but, like, there's all these jailbreaks, but also, like, as soon as they happen, like, you know, it gets fed back into OpenAI's, like, platform, and it gets them ...
It's, it's like being public and open is, like, you know, I guess open, quote-unquote, open to use, um, is, is accelerating their, like, ability to make a better and better model, right? Like the RLHF and all this kinda stuff.
Um, I, I, I don't see how Apple can do that structurally, like, as a company. Like, the, the fruit company ships perfect products, or, like, or else, right? Like, that is, that is their mentality.
They'll tell, they'll build a car before you even see it.
Right, and that's why everyone loves iPhones, right? Like, I have a Samsung. I can tell you how many b- I, I buy a new Samsung every, every other year, right? Uh, maybe I'll buy a Pixel this year. The new one looks nice.
But, like, it's like, you know how many bugs are on these things?
Yeah.
Like, how many times, like, I just have to, like, restart my phone? Like, I mean, it's, it's not, like, often, but it's like, hey, if, like, once a week I need to, like, you know, a, an app just crashes, it's like, "Oh, what the heck?"
Yeah.
Right? It's like, it's like Bing was only ever, like, a few percent behind Google, truly, for the last decade. A few percent, but that few percent is enough to make people be like, "Bing sucks," right? Um, so I think, I think, I, I think that sort of, like, applies to Apple is, like, are you willing to deploy a model of 3.5 capabilities that can say really and do really bad things potentially?
What about 4, right? And the n- the possibility of it doing worse things is even higher. Well, what about 5, right? Like, you can't get on that, like, iteration cycle, right? Um, to build 4, you need to be able to build a 3.5.
Build 3.5, you need to be able to build 3, right, like, of quality, right? Um, and Meta's clearly doing that, right? Like, and, and all these, like, uh, uh, open source firms and, like, all these folks are doing exactly that, right?
Building a bigger and better model every, you know, every few months, and I don't know how Apple gets on that train. Um, but, you know, at the same time, there's no company that has more powerful distribution maybe, right?
Maybe Google does. Maybe, maybe Microsoft does. You can argue that, but, like, so obviously Apple will be deploying things, and Siri will always suck, but-
It's so-
... it'll, it'll, it'll, it'll-
It's embarrassing.
Hey, if, if, if I have a Siri which is GPT-3.5 level in two years, I think a lot of people still use Siri, right? People still use Siri to this day, right? Like, so it's like, you know, same thing's gonna happen, right?
So I don't know.
Tim Cook is not in the AI safety discussions. He doesn't wanna be, you know. He's just in the product, um, side. And I know you had some safety hot takes, and I, I think it's, like, an interesting dynamic because, you know, Anthropic came out of OpenAI, and then you can kinda make the case that, like, by having more labs, if you're really worried about safety, you're, like, accelerating the unsafe because you have more labs and more compute.
Uh, yeah, what, what's your thought on, like, this whole, this whole space?
AI Safety54:18
So I, obviously I think safety's probably important, but, like- ... I mean, it is important, right? Like, I mean, I read sci-fi novels, right? It's clearly important, right? Like, um, you know, I could easily see how an LLM could ...
I, I, I wrote about this the other day where it was like, it's like, hey, like, if you just look at the demographics across the world, there are, like, there's, like, 30 to 50 million more men than there are women, and they will never get married.
Obviously, obviously on population, that level dynamics, you know, there are LGBTQ, all that stuff happens, and it's great, but, like, you know, like, there are 30 to 50 million more men across the world. They'll always be single. Why can't an LLM, like, like, radicalize them, right, by being its AI girlfriend and then all of a sudden be, like, inciting, like, you know.
And also, yeah, I don't know. There's, like, all sorts of stuff, like, that can happen, of course, right? Or, like, you know, teach some person to create a manufacturer what w- they th- what they thought was a good thing, and it ends up wiping out humanity.
Like, all these sorts of stuff can happen. But, you know, at the end of the day, I think security through obscurity doesn't work, right? So that's, that's the approach that the labs take. I truly, I truly do believe it, right?
Like, you know, the, they, they're very open internally, at least, at least Anthropic and OpenAI are. I know Google's a lot more gated with Gemini information. Uh, but, like- You know, of these three, it's like security through obscurity, and it's like this doesn't ever work.
And two, like innovating in the open is, is gonna have more people like figuring out what doesn't work, also figuring out how to, how to like maybe try and, and align things, you know, better.
Maybe the SemiAnalysis, uh, analyst point of view is, is it feasible to build this capacity up in the US?
Uh, no.
No, right?
People don't understand how fragmented the semiconductor ch- supply chain really is-
Right
... and how many monopolies there are. The US could absolutely shut, shut down the Chinese semiconductor supply chain. They won't, but... And China could absolutely shut down the US one actually, by the way. Uh, but more, more relevantly, right, is like, you know Austria has two companies?
Like, the country of Austria in Europe has two companies that have, you know, you know, super high market share and very specific technologies that are required for every single like, like chip, period. Right? There is no chip that is less than seven nanometer that doesn't get touched by, uh, this one Austrian company's tool, right?
And, and there is no alternative really. And there's another Austrian company, likewise, everything two nanometer and beyond will be touched by their tool. And it's like b- both of these companies are, like doing well, less than a billion dollars in revenue, right?
So it's like you think it's so inconsequential. There's like three or four Japanese chemical companies, same, same idea, right? It's like the, the supply chain is so fragmented, right? Like, people only ever talk about where are the fabs, where, where they actually get produced.
But it's like, I mean, TSMC in Arizona, right? TSMC is building a fab in Arizona. It's, it's quite a bit smaller than the fabs in, in, in Taiwan, but even ignoring that, those fabs still have to ship everything to Taiwan back anyways.
And also they have to get what's called a mask from Taiwan and get sent to, get sent to Arizona. And by the way, there's these Japanese, uh, companies that make these chemicals that need to ship to, you know, uh, you know, like TOK and Shinetsu, and you know, it's like...
And, and hey, it needs this tool from Austria. No matter what. It's like, oh wow, wait, actually, like the entire supply chain is just way too fragmented. You can't like re-engineer and rebuild it, uh, on a snap, right?
It's just like that. It's just complex to do that. Semiconductors are more complex than any other thing that humans do. Uh, without a doubt. There's more people working in that supply chain with XYZ backgrounds and, uh, more money invested every year and R&D plus CapEx, you know?
It's like, it's just by far the most complex supply chain that humanity has, and to think that we could rebuild it in a few years is absurd.
Yeah. In an alternate universe, the US kept Morris Chang and people think Right? Like it, it was just one guy that-
Yeah, in an alternative universe, Texas Instruments, uh, communicated to Morris Chang-
Poured on the trillion dollars
... that he would become CEO, and so he never goes to Taiwan and, you know, blah, blah, blah, right? Yeah, no. But I, I, I... You know, that's just... Also, I think, I think the world would probably be further behind in terms of technology development if that didn't happen, right?
Like technology proliferation is how you accelerate, uh, you know, the pace of innovation, right? So the, the, you know, the dissemination to, "Oh wow, hey, it's not just a bunch of people in Oregon at Intel that are leading everything," right?
Or, you know, "Hey, a bunch of people in Samsung, Korea," right? Or Hsinchu, Taiwan, right? It's actually all three of those, plus all these tool companies across the country in the, the Netherlands and, and in Japan and the US and, you know, it's, it's millions of people innovating on a disseminated technology that's led us to get here, right?
I don't even think, you know, if, if Morris Chang didn't go to Taiwan, would we even be at five nanometer? Would we be at seven nanometer? Probably not, right? Like there's innovations that, that, you know, happened because of that, right?
Mm-hmm.
Let's get a quick lightning round done.
Yeah, sure.
Uh, SemiAnalysis branded one. So the first one is, what are, like foundational readings that people that are listening today should read to get up to speed on, like semis?
Recommended Readings58:46
Keep in mind our audience is a lot of software engineers.
Yeah. Yeah. So I think, I think the easiest one is like, uh, is the PyTorch 2.0 and Triton one that I did. Um, you know, there's the advanced packaging series. Um, there's the Google infrastructure supremacy piece. I think that one's really, uh, uh, critical because it explains Google's infrastructure, uh, quite a bit from networking through chips, through all that sort of history of the TPU a little bit and all this sort of stuff.
AMD's MI300 piece, it talks a lot about... The one that I, we did on that are very good. Chip Wars by Chris Miller, who doesn't recommend that book, right? Uh, it's a really good book, right? I mean, like, I would say, uh, Gordon Moore's book is freaking awesome because you gotta think about, right, like, you know, LLM scaling laws are like Moore's law on crack, right?
Uh, kind of like, you know, in a different sense. Like y- you know, if you think about all of human productivity gains since the '70s, is probably just off of the base of semiconductors and technology, right? Of course, of course, people across the world are getting, you know, access to oil and gas and all this sort of stuff, but like, at least in the Western world, since the '70s, everything has just been mostly innovated because of technology, right?
Oh, we're able to build better cars because semiconductors enable us to do that. Or we're able to build better software because, or we're able to connect everyone because semiconductors enabled that, right? So it's like, that is like... I think that's why it's the most important industry in the world, but, like seeing the frame of mind of what Gordon Moore has written, you know, he's got a couple, you know, papers, books, et cetera, right?
Um, only the paranoid survive, right? Like I think, I think like that philosophy and thought process really translates to the now modern times, except maybe, you know, humanity has been an exponential S-curve and this is like another exponential S-curve on top of that.
So I think that's probably a good, good readings to do. Um-
Has there been a equivalent pivot? So Gordon, like that classic tale was more of like his, the pivot to memory.
From memory to logic.
To logic. Yeah.
Yeah.
Yeah. And then was there, is there, has there been an equivalent pivot, um, in, in semis history of, of that magnitude?
I mean, I, I mean like, you know, some people would argue that like, you know, Jensen, you know, he, he basically didn't care about- He only cared about, you know, like gaming and, and 3D professional visualization and like rendering and things like that until like he started to learn about, uh, AI, and then all of a sudden he's going to like universities like, "You want some GPUs?
Here you go." Right? Like, I think there's even stories of like, like, you know, not so long ago, Nerips, when it used to have the more unfortunate name, he would go there and just give away GPUs to people, right?
Like there was like stuff like that. Like, you know, very grassroots, like pivoting the company. Now, like you, you look on gaming forums and it's like everybody's like, "Oh, Nvidia doesn't even care about us. They only care about AI."
And you know, it's like Yes, you're right. They only care-- They, they mostly only care about AI, and, and the gaming innovations are only because of, like, they're putting more AI into it, right? It's like... But also, like, hey, they're doing a lot of chip design stuff with AI, and you know, I think, I think that's a, like, a not a-- I don't know if, uh, it's equivalent pivot quite yet, but, you know, because the digital s- you know, logic is a pretty big innovation.
But I think that's a big one, and, you know, likewise, it's like, you know, what did, what did OpenAI do, right? What did they pivot? How did they pivot? And they left the, like, you know, a lot of-- most people left the culture of, like, Google Brain and DeepMind and, and decided to build this, like, company that's crazy cool, right?
Like, and does things in a very different way and, like, is innovating in a very different way. So you can, you consider that a pivot even though it's not inside Google. Um, I don't know.
They're on a very different path with like the Dota games and all that before they eventually found like GPTs as the, as the thing. So it, it was a full, like, started in twenty fifteen-
Yeah
... and then really pivoted twenty nineteen to be like, "All right, we're the GPT company."
Yeah. Yeah.
Uh, if I could classify them. I, I don't-- I, I'm sure there's OpenAI people who are yelling at me right now. Uh, uh, okay, so maybe, maybe, uh, just a general question of a, you know, I'm a fellow writer on, on Substack.
You are obviously managing your, your consulting business while you're also publishing these amazing posts. Uh, how do you-- What's your writing process? How do you source info? Like, when you sit down and go like, "Here's the theme for the week," do you, do you have a pipeline coming up?
Writing Process1:02:39
Just anything you can describe.
So, um, I'm thankful for my, uh, you know, my teammates 'cause they are actually awesome. Like, uh, and they're much more, um, you know, directed, focused to working on one thing. You know, or not one thing, but a number of things, right?
Like, you know, someone who's just expert on X and Y and Z in the semiconductor supply chain. So that really helps with the con- the, that side of the business. I most of the times only write when I'm very excited or, you know, it's like, hey, like we should work on this and we should write about this.
So, like, you know, one of the most recent posts we did was we explained the manufacturing process for 3D NAND, uh, you know, flash storage, uh, gate-all-around transistors and, um, 3D DRAM and all this sort of stuff 'cause there's a company in Japan that's going public, uh, Kokusai Electric, right?
And it's like, okay, well, we should do a post about this, and we should explain this. But like, it's like, okay, we, we, you know... And so Myron, um, he, he did all that work, Myron Shi, in, in most of the work and, and awesome.
But, like, usually it's like there's a few like very long in-depth backburner-type things, right? Like that took a long time. Took, you know, over a month of, of research, and Myron knew- knows this stuff already really well, right?
Like, but also furthermore, it's like, you know... So there's stuff like that that we do, um, and, and that, like, like builds up a body of work for our consulting and, and some of the reports that we sell that aren't, you know, newsletter posts.
But a lot of times the process is also just like... Well, like Gemini Eats the World is the culmination of reading that, um, having done a lot of work on the supply chain around the TPU ramp and CoWoS and HBM capacities and all this sort of stuff to be able to, you know, figure out how many units and that Google's ordering, all this sort of stuff.
And then like also like looking at like open source. It's like all just the, all that culminated in like I wrote that in four hours, right? Sent it to a couple people, and they were like, "No, change this, this, this."
"Oh, you know, add this 'cause that's really gonna piss off, you know, the open source community." I'm like, "Okay, sure." Um, and then posted it, right? And so it's like there's no like specific process. Unfortunately, like the most viral posts, especially in the AI community, are just like those kind of pieces rather than the, like, the really deep, deep like...
What was in the, uh, Gemini Eats the World post, you know, obviously like, hey, like we, we do deep work. There's a lot more like factual, not leaks, uh, you know, it's just factual research. Hey, we, you know, we go-- Across the team, we go to forty-plus conferences a year, right?
All the way from like a photoresist conference to a photomask conference to a lithography conference, all the way up to like AI conferences and, you know, all, everything in between, networking conferences and piecing everything across the supply chain.
So it's like that's like the true like work and like... Yeah, I don't know. It, it is sometimes bad to like have the inf- inf- inf-infamousness of, you know, only people caring about this or the GPT-4 leak or the, uh, Google has no moat leak, right?
It's like, but like, you know, that's just like stuff that comes along, right? You know, it's really focused on like understanding the supply chain and how it's pivoting and who's the winners, who's the losers, what technologies are inflecting, things like that.
Where is the best place to invest resources, you know, sort of like stuff like that, and, and accelerating or, uh, capturing value, et cetera.
Awesome. And to wrap, we're trying a new question. Uh, if you had a magic genie that could answer any question that would change your worldview, uh, what question would you ask?
Magic Genie1:05:44
That's a tough one.
Like you, you operate based on a set of facts about the world right now, and there's maybe someone knows where you're like, "Man, if I really knew the answer to this one, I would do so many things differently," or, "I would think about things differently."
Everything that we've seen so far is that large-scale training has to happen in an individual data center with very high-speed networking. Now, everything doesn't need to be all to all connected, but you need very high-speed networking between all of your, uh, your chips, right?
I would love to know, you know, hey, magic genie, how can we build artificial intelligence in a way that it can use multiple data centers of resources where there is a significantly lower bandwidth between pools of resources, right?
Because that would instantly... Like one of the big bottlenecks is how much power and how many chips you can get into a single data center. So like, A, Google and OpenAI and Anthropic are working on this, right? Um, and I don't know if they've solved it yet.
But if they haven't solved it yet, then what is the solution? Because that will like accelerate the scaling that can be done by not just like a factor of ten, but like orders of magnitude because there's so many different data centers, right?
Like if you, you know, uh, across the world and, you know, uh, oh, if I could pick up, you know, if I could effectively use two fifty-six GPUs in this little data center here and then with this big cluster here.
You know, how can you r- make an algorithm that can do that, right? Like I think that would be like the number one thing I'd be curious to know if, how, what, because that changes the world significantly in terms of how con- we continue to scale, uh, this amazing technology that people have invented, uh, over the last, you know, five years.
Awesome. Well, thank you so much for coming on, Dylan.
Thank you so much for having me. Hopefully my, uh, rambling, especially on AI safety, was not, uh, not, not, not poorly taken, 'cause I think it will be poorly taken.






