# Ep 18: Petaflops to the People — with George Hotz of tinycorp

Latent Space · 2023-06-20

<https://addtry.com/fd3ba6fb-c9e4-4fa6-a006-d753bbb06e93>

George Hotz of tinycorp argues that a simplified, open-source ML framework (tinygrad) can democratize AI compute and replace NVIDIA, Google, and AMD. Tinygrad uses 25 ops (vs. XLA's 250) and fuses kernels automatically, achieving 2x speed on Qualcomm GPUs over their library. Hotz reports AMD kernel panics fixed after emailing CEO Lisa Su, but laments AMD's open-source as 'dumped on GitHub.' For hardware, tinybox is a six-GPU desktop at 350W per GPU, aiming for 5x better price-performance than NVIDIA H100 for 90% of training. He criticizes OpenAI's secrecy, revealing GPT-4 is an 8-way mixture of 220B models, and cites the Bitter Lesson. He also outlines FLOPcoin using hardware identity to prevent cheating, and his goal of an AI girlfriend as the third company, merging humans via data rather than implants.

## Questions this episode answers

### Why does George Hotz describe tinygrad as following a RISC-like philosophy compared to PyTorch and XLA?

George argues that ML frameworks like PyTorch and XLA have about 250 operations, making them complex like CISC processors, while tinygrad uses only about 25 ops, similar to RISC. By radically simplifying the instruction set, tinygrad achieves about 10x less complexity, making it easier to write and optimize for any hardware, especially non‑Nvidia chips where developer efficiency is key.

[3:01](https://addtry.com/fd3ba6fb-c9e4-4fa6-a006-d753bbb06e93?t=181000)

### What does George Hotz claim is the real architecture of GPT‑4?

George states that GPT‑4 is not a trillion‑parameter monolithic model, but an eight‑way mixture of experts, with each expert having 220 billion parameters. He asserts OpenAI trains eight copies of the same model and then combines them, calling this a trick used when you can’t make a single model bigger, not an amazing breakthrough that justifies secrecy.

[49:48](https://addtry.com/fd3ba6fb-c9e4-4fa6-a006-d753bbb06e93?t=2988000)

### What problems did George Hotz have with AMD GPUs, and how did he resolve them?

George experienced kernel panics that crashed his entire system when running AMD’s own demo apps. Frustrated, he emailed AMD CEO Lisa Su, who responded and connected him with engineers. They sent a pre‑release driver that fixed the panics, and he now feels AMD is beginning to understand the need for a reliable, open‑source‑friendly approach to keep developers from moving to competitors like Intel.

[23:20](https://addtry.com/fd3ba6fb-c9e4-4fa6-a006-d753bbb06e93?t=1400000)

## Key moments

- **[0:00] Intro**
  - [0:26] George Hotz traded the first unlocked iPhone for a Nissan 350Z and three new iPhones.
  - [0:59] "What's the difference between a dev kit and not a dev kit? Nothing. Just the question of do you think it's for you?"
- **[2:08] Tiny Theses**
  - [2:18] George Hotz: "What are the odds they nationalize Nvidia?" motivated him to start tinycorp.
  - [2:39] George Hotz says difficult experiences with Nvidia and Qualcomm buying chips led him to start tinycorp.
  - [3:22] Tinygrad reduces ML operations from 250 to 25, akin to RISC vs CISC in processors, says George Hotz.
  - [4:15] George Hotz: "If you can write a fast ML framework for GPUs, you just cannot write one for your own chip."
  - [4:48] George Hotz claims TPUs are the only other successful training chip besides Nvidia, because Google wrote their own ML framework.
  - [6:09] George Hotz's thesis: neural networks' static branches mean Turing completeness is harmful and should be avoided in ML hardware.
- **[11:22] Tinygrad Design**
  - [11:27] Tinygrad started with a hard 1,000-line limit to force code efficiency, now 2,800 lines and still readable.
  - [12:05] George Hotz says PyTorch code has 'so much boilerplate' that tracking down an LU is a deep stack.
  - [13:32] Tinygrad automatically fuses operations like A*B+C to avoid extra memory trips, while PyTorch can't for arbitrary code.
  - [14:39] George Hotz: "Why is Relu a class?" criticizing PyTorch's unnecessary stateful design.
  - [16:06] Tinygrad uses laziness—a middle ground between eager and graph compilation—to enable operation fusion automatically.
  - [17:53] George Hotz: "The magic is when you can make your software do more without adding complexity." Complexity leads to collapse.
- **[23:19] AMD Saga**
  - [23:19] George Hotz tried for a day to compile PyTorch on AMD and got kernel panics from AMD's own demo apps.
  - [24:54] George Hotz considered Intel GPUs for tinycorp because of stable drivers and open documentation, but found performance per dollar poor.
  - [25:20] George Hotz emailed AMD CEO Lisa Su about kernel panics and received a response, leading to a fix.
  - [26:53] George Hotz: "It's not open source if you dump the source code on a GitHub repo and forget about it."
  - [27:29] George Hotz experienced Nvidia responding to an NCCL bug within an hour, contrasted with AMD's slow open-source responsiveness.
- **[28:18] GGML & Quantization**
  - [28:37] George Hotz says GGML focused on Apple M1s, while tinygrad aims to be a universal acceleration framework.
  - [29:07] George Hotz found a bug in PyTorch's M1 Metal backend where a matrix operation gave wrong results compared to CPU.
  - [30:17] George Hotz emphasizes CI is 'super important' for tinygrad: the green check must mean it's safe to merge.
  - [30:35] George Hotz on Mojo: "It's closed source. I'm not that interested. I'm interested when it's open."
  - [31:56] George Hotz's Tinybox is a silent, six-AMD-GPU personal compute cluster designed to fit under a desk and run inference at home.
  - [32:55] George Hotz is limited by off-the-shelf GPUs; for running FP16 Llama, memory bandwidth is key.
- **[37:39] Tinybox Hardware**
  - [37:50] George Hotz explains building a six-GPU machine is hard: PCIe extenders often fail at Gen4, and managing power and noise is complex.
  - [40:01] George Hotz's Tinybox design goals: silent at 45 dB, fits under a desk, plugs into one wall outlet.
  - [40:34] George Hotz envisions Tinybox as a home AI hub for running local inference for robotics, avoiding cloud latency and cost.
  - [41:33] George Hotz uses the term "compute cluster" to legally use Nvidia GPUs under their license agreement.
  - [42:13] George Hotz says PCIe 4.0's 60 GB/s interconnect limits Tinybox training to ~7B parameter models, vs 70B with NVLink.
  - [42:52] George Hotz predicts the best chatbots will arise from 1,000 training runs, not one big run, favoring many experiments over model size.
- **[45:17] FLOPcoin**
  - [45:17] George Hotz proposes FLOPcoin: a cryptocurrency mined by Tinyboxes during idle cycles, with Sybil resistance via hardware keys.
  - [46:01] "If you ever send wrong data, you're banned from the network for life": George Hotz on FLOPcoin cheating deterrent.
  - [47:38] George Hotz: "I wanna be on the absolute edge of FLOPS per dollar and FLOPS per watt" for enterprise training clusters.
  - [48:05] George Hotz defines one 'person of compute' as 20 petaFLOPS; Comma.ai's cluster is 30 petaFLOPS.
- **[49:38] GPT-4 & Bitter Lesson**
  - [49:46] George Hotz claims GPT-4 is a 220B parameter model with an 8-way mixture of experts, a trick used when out of ideas.
  - [50:37] "When a company is secretive, it's because they're hiding something that's not that cool": George Hotz on GPT-4 secrecy.
  - [52:53] George Hotz endorses Rich Sutton's 'The Bitter Lesson' that compute always beats hand-engineering.
- **[55:31] Hiring & AI Tools**
  - [55:31] George Hotz hires at tinycorp via open bounties on GitHub; only contributors who prove themselves get offers.
  - [56:56] George Hotz: "The problem is, if your loss function is categorical cross-entropy on the internet, your responses will always be mid."
  - [58:09] George Hotz suggests putting 10 LLMs in a room to debate answers, mimicking how humans don't code by typing straight through.
  - [1:00:02] George Hotz says coding is 'tool complete'—above the API line where humans direct machines, so AI tools will supercharge, not replace, developers.
  - [1:01:00] George Hotz explains the "API line": if your manager is a computer, you're below it; programmers are above it.
  - [1:01:47] George Hotz: "Did Photoshop replace artists? Like what are you talking about?" — on AI creative tools.
- **[1:07:29] AI Girlfriend**
  - [1:07:29] George Hotz's third company will build an AI girlfriend, which he sees as the ultimate way to merge with a machine and live forever.
  - [1:09:54] George Hotz estimates his brain's Kolmogorov complexity at just a couple gigabytes after compression, enabling digital immortality.
  - [1:10:37] "Don't believe everything you see on social media. Your life could depend on it." — George Hotz on training AI with filtered online personas.
- **[1:11:01] Philosophical Takes**
  - [1:11:08] George Hotz interprets 'the goddess of everything else' as a force that prevents paperclip-maximizer AI, analogous to how cancer doesn't win.
  - [1:12:20] George Hotz is grateful AI progress is open and accessible, not a closed Manhattan Project, and that it only requires high school math.
  - [1:15:23] George Hotz argues transformers work because of dynamic weight generation, not attention: they're continuous-weight-set selectors.
  - [1:17:52] George Hotz contrasts Elon Musk's physics-based ambitions (Mars) with his own information-based ambitions (AI girlfriend, digital immortality).
  - [1:18:55] George Hotz: "Only the left takes ideology seriously. e/acc is not serious" — on effective accelerationism.
  - [1:20:20] George Hotz rewrote Avatar 2's script, killing Jake Sully first, because his second arc couldn't top the first.

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **George Hotz** (guest)

## Topics

Hardware, Open Source Tools, Language Models

## Mentioned

AMD (company), Comma (company), Google (company), NVIDIA (company), tinycorp (company), CUDA (product), Comma Body (product), FLOPcoin (product), GGML (product), Mojo (product), NCCL (product), ONNX (product), ONNX Runtime (product), OpenPilot (product), PyTorch (product), TPU (product), TensorFlow (product), XLA (product), tinybox (product), tinygrad (product)

## Transcript

### Intro

**Swyx** [0:03]
Hey, everyone. Welcome to the Latent Space Podcast. This is swyx, writer and editor of Latent Space, and Alessio is taking over, uh, with the intros. Alessio's partner and CTO in residence at decimal partners.

**Alessio** [0:13]
Hey, everyone. Today, we have geohot on the podcast, AKA George Hotz, um, for the, the human name. Everybody knows George, so I'm not gonna do a big intro. A couple things that people might have missed. So you, you were the first to unlock the iPhone.

You traded the first ever unlocked iPhone for a Nissan 350Z and three new iPhones. Um, you were then one of the first people to break into the PS3 and run arbitrary code. Uh, you got sued by Sony. You wrote a rap song, uh, to fight against that, which is still live on, on YouTube, which we're gonna have on the show notes.

Um, then you did not go to Tesla to build vision, and instead you started Comma.ai, um, which was an amazing engineering feat in itself until you got a cease and desist, uh, from the, from the government to not put these things on the street.

Uh, turned that into a, a research-only project.

**George Hotz** [0:59]
Wait, you know they're out there.

**Alessio** [1:00]
Yeah, yeah. No, no-

**George Hotz** [1:01]
Yeah.

**Alessio** [1:01]
... no, no. They are out there. But like in a-- They're not a, you know, you market them as a research, kinda like no warranty.

**George Hotz** [1:07]
Because I use the word dev kit? That's not about the government. It has nothing to do with the government.

**Alessio** [1:11]
Mm-hmm.

**George Hotz** [1:11]
We offer a great one-year warranty. The truth about that is it's gatekeeping.

**Alessio** [1:17]
Mm-hmm.

**George Hotz** [1:18]
What's the difference between a dev kit and not a get dev kit? Nothing. Just the question of do you think it's for you? And if you think it's for you-

**Alessio** [1:25]
Hmm

**George Hotz** [1:25]
... buy it. It's a consumer product. We call it a dev kit. If you have a problem with that, it's not for you.

**Swyx** [1:30]
Good framing.

**Alessio** [1:31]
That's great insight. Um, and then I was going through your blog post to get to today. You, you've wrote this post about the, the hero's journey.

**George Hotz** [1:38]
Mm-hmm.

**Alessio** [1:38]
And you linked this thing called the portal story, which is kind of this set of stories and movies and books about people living this arbitrary life, and then they run into this magic portal. It kinda takes them into a new, very exciting life and dimension.

When you wrote that post, you talked about tinygrad, which is one of the projects you're working on today, and you mentioned this is more of a hobby, something that is not gonna change the course of history. Obviously, you're now going full speed into it.

So we would love to learn more about what was the portal that, that you run into to, to get here.

**George Hotz** [2:08]
Well, what you realize is... You know what made me realize that I absolutely had to do the company? Seeing Sam Altman go in front of Congress.

### Tiny Theses

**Alessio** [2:17]
Why?

**George Hotz** [2:18]
Uh, what are the odds they nationalize Nvidia?

**Alessio** [2:22]
Hmm.

**George Hotz** [2:22]
You know, what are the odds that large organizations in the government, but of course I repeat myself, um, d- decide to try to clamp down on, uh, accessibility of ML compute? Uh, I wanna make sure that can't happen structurally, so that's why-

**Alessio** [2:38]
Mm-hmm

**George Hotz** [2:39]
... uh, I realized that it's really important that I do this. And actually, from a more practical perspective, I'm working with Nvidia and Qualcomm to buy chips. Nvidia has the best training chips. Qualcomm has the best inference chips.

Uh, working with these companies is really difficult, uh, so I'd like to start another organization that, uh, eventually, in the limit, either works with people to make chips or makes chips itself and makes them, uh, available to anybody.

**Alessio** [3:01]
Yeah. You share kinda three core thesis to, to tinycorp. Maybe we can dive into each of them. So xla, Prime Torch, those are the complex instruction system. Tinygrad is the restrict, uh, restricted instruction system. So you're kind of focused on, again, tinygrad being small, not being overcomplicated, and trying to get as close to, like, the DSP as possible in a way where it's add more.

**George Hotz** [3:22]
Well, it's a very, it's a very, uh, clear analogy from how processors developed. So a lot of processors back in the day were CISC, complex instruction set. Um, System 360 and then x86.

**Alessio** [3:32]
Mm-hmm.

**George Hotz** [3:33]
Uh, then this isn't how things stayed. Uh, they went to now the most common processor is Arm, um, and people are excited about RISC-V, right? No one's excited about... RISC-V is even less complex than Arm.

**Alessio** [3:44]
Mm-hmm. Yeah.

**George Hotz** [3:45]
Um, no one is excited about CISC processors anymore. They're excited about redis- uh, RISC, reduced instruction set processors. So tinygrad is we're going to make a RISC, uh, op set for all ML models. And yeah, it can run all ML models with, with basically 25 instead of the 250 of XLA or Prime Torch.

So about 10X less complex.

**Alessio** [4:05]
Yeah. You talk a l- a lot about existing AI chips. You said if you can write a fast ML framework for GPUs, you just cannot write one for your own chip, so that's another one of your core insights.

I don't know if you wanna expand on, on that.

**George Hotz** [4:18]
Yeah. I mean, your chip is worse, right? There's no way the chip that you're gonna tape out, especially on the first try, is going to be easier to use than an AMD GPU.

**Alessio** [4:25]
Mm-hmm.

**George Hotz** [4:25]
Right? And yet there's no good stack for AMD GPUs. So why do you think you can make one for your chip? You can't, right? The, the only company-- There's one other company aside from Nvidia who's r- succeeded at all at making training chips.

What company?

**Alessio** [4:41]
Um.

**Swyx** [4:42]
AMD?

**Alessio** [4:43]
Intel?

**George Hotz** [4:43]
No. No. No. I've never trained... Who's trained a model on AMD or Intel?

**Alessio** [4:48]
No- nobody on AMD.

**Swyx** [4:49]
Cerebras.

**George Hotz** [4:50]
Cerebras, uh, I'm talking about you might know some startups who trained models on these chips. I'm surprised no one immediately gets this, because there is one other chip aside from Nvidia that normal people have actually used for training.

So-

**Alessio** [5:02]
Apple Neural Engine?

**George Hotz** [5:04]
No. Used for training.

**Alessio** [5:06]
Nobody. Nobody.

**George Hotz** [5:06]
You, you can only buy them in the cloud.

**Alessio** [5:09]
Oh, TPU.

**George Hotz** [5:10]
Exactly.

**Alessio** [5:10]
Yeah, yeah.

**Swyx** [5:10]
TPUs.

**George Hotz** [5:10]
Right? So, uh, uh, Midjourney is trained on TPU, right?

**Alessio** [5:14]
Yeah.

**George Hotz** [5:14]
Like, a lot of startups do actually train on TPUs. Uh, and they're the only other successful training chip aside from Nvidia. But what u- what's unique about Google is that they also wrote their own ML framework.

**Alessio** [5:25]
Mm-hmm.

**George Hotz** [5:26]
Right? And if you can't write your own ML framework that is performant on Nvidia, there's no way you're gonna make it performant on your...

**Alessio** [5:32]
Yeah. And they started from TensorFlow, and then they, they, they made the chip after.

**George Hotz** [5:36]
Yeah.

**Alessio** [5:36]
That's-

**George Hotz** [5:37]
Exactly.

**Alessio** [5:37]
Yeah.

**George Hotz** [5:37]
Exactly.

**Alessio** [5:37]
Yeah.

**George Hotz** [5:37]
And you, you have to, you have to do it in that direction. Otherwise, you're gonna end up, um, you know, a, a Cerebras? What? Those things are million... Uh, does anyone... I've never seen a Cerebras. No one's ever like, "Oh, I trained my model on a Cerebras."

Most people are like, "I trained my model on GPUs." Some people, 20%, are like, "I trained my model on TPUs."

**Alessio** [5:57]
Yeah. And then the third one, which is the one that surprised me the most, is, uh, true incompleteness is harmful-

**George Hotz** [6:02]
Mm-hmm

**Alessio** [6:02]
... should be avoided. It make, it made sense once I read it But maybe tell us a bit more about how you got there.

**George Hotz** [6:09]
Um, okay. So CPUs, uh, devote tons of their silicon and power to things like reorder buffers-

**Alessio** [6:18]
Mm-hmm

**George Hotz** [6:18]
... and speculative execution and branch predictors. And the reason that you need all these things is because at compile time, you can't understand how the code's going to run, right? This is, this is Rice's theorem. This is the halting-

**Alessio** [6:29]
Mm-hmm

**George Hotz** [6:29]
... problem and its limit.

**Alessio** [6:31]
Yep.

**George Hotz** [6:31]
Um, and this is not like, "Oh, the halting problem is, is theoretical." No, no, no, no. It's actually very real. Does this branch get taken or not? Well, it depends on X. Where does X come from? Yeah, forget it, right?

Um, but no branches depend on X in a neural net. Every branch is a static loop.

**Alessio** [6:46]
Mm-hmm.

**George Hotz** [6:46]
Like, if you're doing a matrix multiply, it's a static loop over the inner dimension. Um, and neural networks are even better. No loads even depend on X, right?

**Alessio** [6:54]
Mm-hmm.

**George Hotz** [6:54]
So with a GPU shader, right, you're like, your load might depend on which texture you're actually loading into RAM. But with a neural network, your load is, "Well, I load that way." Why? "Well, 'cause I load that way the other million times I ran the same net."

Every single time you run the net, you do the exact same set of loads, stores, and arithmetic. The only thing that changes is the data, and this gives you a very powerful ability to optimize that you can't do with, uh, CPU-style things which have branches, and even GPU-style things which have loads and stores.

**Alessio** [7:23]
Oh, that makes sense.

**George Hotz** [7:24]
Well, GPUs, if you want GPU-style stuff, you have, like, load-

**Alessio** [7:26]
Yeah

**George Hotz** [7:26]
... based on X, you now need a cache hierarchy.

**Alessio** [7:29]
Mm-hmm.

**George Hotz** [7:29]
And not a explicit cache hierarchy, an implicit cache hierarchy with, with, with eviction policies that are hard-coded into the CPU. Like, you start doing all this stuff, and you're never gonna get, like, theoretically good performance. Again, I don't think there's 100X.

You know, some startups will talk about 100X, and they'll talk about absolutely ridiculous things like clockless computing or analog computing.

**Alessio** [7:49]
Mm-hmm.

**George Hotz** [7:49]
Okay. Here, analog computing just won't work, and clockless computing- ... um, sure, it might work in theory, but your EDA tools are... Maybe, uh, AIs will be able to design ch- clockless chips, but not humans. Um, but what actually is practical is changing cache hierarchies-

**Alessio** [8:06]
Mm-hmm

**George Hotz** [8:06]
... and removing branch predictors and removing warp schedulers, right? GPUs spend tons of power on warp scheduling because we have to hide the latency from the memory. Well, hide the latency if everything's statically scheduled.

**Alessio** [8:15]
Yeah. Why do you think people are still hanging on to Turing complete?

**George Hotz** [8:21]
Well, 'cause it's really easy. Turing completeness is really easy, right? It's, it's really easy to just, "Ooh, you know, it'd just be so nice if I could do, like, a, like, an if statement here and actually branch the code," right?

**Alessio** [8:32]
Mm-hmm.

**George Hotz** [8:32]
So i- it requires a lot more thought to do it without Turing completeness.

**Swyx** [8:37]
And would this be qualitatively different than TPUs?

**George Hotz** [8:41]
Uh, so TPUs are a lot closer.

**Swyx** [8:42]
Yeah.

**George Hotz** [8:42]
TPUs are a lot closer to what I'm talking about-

**Swyx** [8:45]
Exactly

**George Hotz** [8:45]
... than, than, like, like CUDA. Okay, so what is CUDA? Well, CUDA's a C-like language which compiles to an LLVM-like IR, which compiles to PTX, which compiles to SAS-

**Swyx** [8:55]
Mm-hmm

**George Hotz** [8:55]
... which are all Turing complete. Uh, TPUs are much more like this, yeah. Their, their memory is pretty statically managed. They have a V... I did some reverse engineering on the TPU. Uh, it's published in, in tinygrad. Uh, it has, like, a VLIW instruction, uh, and it runs them.

So it's similar. I think the TPUs have a few problems. Uh, I think systolic arrays are the wrong choice. Um, systolic array, I think they have systolic arrays because-

**Swyx** [9:17]
Mm-hmm

**George Hotz** [9:17]
... that was the guy's PhD.

**Swyx** [9:18]
Right.

**George Hotz** [9:18]
And of course, Amazon makes-

**Swyx** [9:19]
Could you, could you summarize systolic arrays for us?

**George Hotz** [9:21]
Systolic arrays are just, um, okay, so basically you have, like... This is a way to do matrix multiplication. Uh, think of a grid of mul-adds, and then the grid can multiply and then shift, multiply then shift, multiply-

**Swyx** [9:31]
Mm-hmm

**George Hotz** [9:32]
... then shift. And they are very power efficient, but it becomes hard to schedule a lot of stuff on them if you're not doing, like, perfectly sized dense matrix multiplies, which you can argue, well, design your models to use perfectly sized dense matrix multiplies, sure.

But, um, it's just, it's just-

**Swyx** [9:52]
No, but thanks for indulging on, on these, uh, explanations. I think we need to keep our audience along with us-

**George Hotz** [9:58]
Yeah

**Swyx** [9:58]
... by s- pausing every now and then to explain key terms.

**George Hotz** [10:01]
Y- you know, the, when I say e- explain a systolic array, I, I just immediately get a picture in my head of, like, tilting a matrix and shifting it. It's, like, hard to kind of explain.

**Swyx** [10:09]
Yeah.

**Alessio** [10:09]
Yeah.

**Swyx** [10:10]
Well, there is a video, so you, you, you had it-

**Alessio** [10:11]
We'll have, we'll have show notes with the-

**Swyx** [10:12]
And, and we edit, we edit in visuals.

**George Hotz** [10:14]
Yeah, yeah, yeah. There's some great graphics that just show you, oh, so that's what a systolic array is. But it's a, it's a, it's a mul-add shift machine that looks kinda different from the tr- typical, like, APU sort of machine.

Uh, a, sorry, ALU sort of machine. I think the right answer is something that looks more like queues that feed into ALUs, and then you can, like, prefetch the loads from the memory, put in a bunch of queues, and then the queues just, like...

and feeds into another queue over here. Um, but yeah, uh, but that, that's not even the main problem with TPUs. The main problem with TPUs is that they're closed source.

**Alessio** [10:45]
Mm-hmm.

**George Hotz** [10:45]
Not only is the chip closed source, but all of... XLA is open source, but the XLA to TPU compiler is a 32-megabyte binary blob called Lib TPU on Google's Cloud instances.

**Alessio** [10:54]
Mm-hmm.

**George Hotz** [10:54]
Right? It's all closed source. It's all hidden stuff, and you know, well, there's a reason Google made it closed source. Amazon made a clone of the TPU. It's called Inferentia, um, or they have some other name for it, a training-

**Swyx** [11:03]
Trainium, yeah.

**George Hotz** [11:04]
Trainium, yeah, yeah, yeah. And you would look, it's a clone of the TPU. Uh, it, the software doesn't work though. The Google software at least kinda works.

**Alessio** [11:12]
Um, so those are kinda, like, the three core pieces. Um, the first thing you're working on, that you've been working on is tinygrad. Um, and one of the, your Twitch streams you said is the, the best thing you've ever written.

**George Hotz** [11:22]
Yeah.

### Tinygrad Design

**Alessio** [11:22]
Um, yeah, tell us a bit more about, um, that creation.

**George Hotz** [11:27]
For a long time, tinygrad had a hard limit at 1,000 lines of code, and what this would force you to do is really make sure you were not wasting lines. Um, I got rid of the restriction because it became a little code golfy at the end.

But once, like, the core framework of tinygrad was there in those 1,000 lines... It's, it's, it's not huge now. It's, like, 2,800 lines now. It's still very readable. Um, but, like, the core framework, the ideas are expressed with no boilerplate.

If you go read PyTorch... Uh, you know, PyTorch I think is actually pretty good code. I think Facebook's pretty good. Um, but there's so much boilerplate.

**Alessio** [12:04]
Mm-hmm.

**George Hotz** [12:05]
Go, go, go in PyTorch and try to track down how an LU actually works.

**Host** [12:10]
It's just a lot of distractions

**George Hotz** [12:11]
Oh, you're gonna be, you're gonna be, you're gonna be diving down a, a long stack from Python to C to custom libraries to dispatchers to... And then I don't even know how to read TensorFlow. Like, I don't even know where, where's the LU in TensorFlow?

**Host** [12:22]
Mm-hmm.

**George Hotz** [12:22]
No-nobody knows. Um, someone at Google knows maybe. Uh, Google as an organism knows. I don't know if anyone individual at Google knows.

**Host** [12:32]
What are, like, the important ergonomics, like, for a developer as you think about designing the tinygrad API?

**George Hotz** [12:37]
So the tinygrad front end looks very similar to PyTorch. Um, there's an even higher level front end you can use for tinygrad, which is just ONNX.

**Host** [12:44]
Mm-hmm.

**George Hotz** [12:44]
Uh, we support-- We have better support for ONNX than Core ML does. Um, and we're going to have... I think we're gonna pass ONNX Runtime soon too. And, like, people think ONNX Runtime, that's the gold standard for ONNX.

No, you can do better.

**Host** [12:54]
Pass them in what specifically?

**George Hotz** [12:55]
Uh, test, uh, compliance tests.

**Host** [12:57]
Okay.

**George Hotz** [12:57]
So ONNX has a big set of compliance tests that you can check out.

**Host** [13:00]
Okay.

**George Hotz** [13:01]
Um, and we have them running in tinygrad, and there are some failures. Um, we're below ONNX Runtime, but we're beyond Core ML. Uh, so, like, that's, like, where we are in ONNX support now. But we, we will pass, we will pass ONNX Runtime soon, uh, because it becomes very easy to add ops because of how, like, you don't need to do anything at the lower levels.

You just do it at this very high level, and tinygrad compiles it to something that's fast using these minimal ops with. Um, you can, like, write... I mean, mo-most concretely, what, what tinygrad can do that, like, PyTorch can't really do is if you have something like A times B plus C, right?

If you write that in naive PyTorch, what it's going to do on the GPU is, well, read A, read B in a kernel, and then store A times B in memory, and then launch another kernel to do A times B plus C.

Okay? Gotta do those loads from memory. And now I did a whole extra round trip to memory that I just didn't have to do.

**Host** [13:50]
Mm-hmm.

**George Hotz** [13:50]
And you're like, "Yeah, but you can use the Torchjit, and it corrects this." Yeah, for that one example. For that one example of mul-add, but, oh, now you did three multiplies, six multiplies. Right? It, it doesn't, uh... It won't compile arbitrary code.

It-

**Host** [14:04]
And have you looked into, like, the other approaches like PyTorch Lightning, um, it-- to accelerate PyTorch itself?

**George Hotz** [14:10]
Well, PyTorch Lightning, my understanding, is it's a mostly a, a, uh, framework around PyTorch, right? PyTorch Lightning is not gonna fix this fundamental problem of I multiply six tensors together. Why is it going to memory any more than a single read from each and a single write to the output?

**Host** [14:24]
Okay. Yeah.

**George Hotz** [14:24]
Um, yeah, there are, uh, there are lower level things in PyTorch that are... I'm not exactly sure what Dynamo does, um, but I know they're generating some Triton stuff, which is gonna generate the kernels on the fly. Um, but, you know, Py-PyTorch Lightning is as, as at a higher level of abstraction.

So tinygrad's front end stuff looks like PyTorch. I made a few tweaks. There's a few things I don't like about PyTorch. Why is Relu a class? No, really. Like, what, what, what is the state? As you, you make a class and there's a state.

Everything should just be torch.functional.nn.relu, but just .relu on the tensor. Also, like, there's things in Torch where you have to do tensor. and not a tensor., right?

**Host** [15:00]
Mm-hmm.

**George Hotz** [15:00]
Um, and, like, why, why are these things... Like, this just, it just shows an API that's, like, not perfectly refined. But when you're doing stuff tinygrad style where you don't have lines, well, it has to work this way.

**Host** [15:10]
Mm-hmm.

**George Hotz** [15:11]
Because even the lines to express the, well, you can't use the where operator unless... And the where operator in PyTorch. Why is it, uh, true case, condition, false case? You know, ugh, the worst. That's, like, how Python expresses ifs.

It's disgusting, right? Ternary operators are much nicer. It should be... I can do my, like, A less than zero.where A, comma, one, right?

**Host** [15:32]
Mm-hmm. The very pandas, uh, like APIs.

**George Hotz** [15:36]
Yeah, yeah, yeah, yeah, yeah, yeah. It's, it's, it's some-- It looks like Torch, NumPy, pandas. They're all very similar. I tried to take, like, the cleanest subset of them and express them. But like I said, you can also interact with it using ONNX.

**Host** [15:47]
Yeah. Mm-hmm.

**George Hotz** [15:49]
Um, but I've, I have a rewrite of Stable Diffusion, I have a rewrite of Llama, I have a rewrite of Whisper. You can look at them. They're shorter than the Torch versions, and I think they're cleaner.

**Host** [15:54]
And you stream them all?

**George Hotz** [15:55]
Yeah.

**Host** [15:56]
Very nice. Um, laziness is kind of the other important concept that you're leveraging to do operation fusing. Um, yeah, talk a bit more about that.

**George Hotz** [16:06]
So yeah, you have, you have basically, like, a, a few different, like, models for, uh, compute. The simplest one's eager, right?

**Host** [16:15]
Mm-hmm.

**George Hotz** [16:15]
The simplest one is eager, is as soon as the, the, uh, interpreter or... sees A times B, it actually dispatches A times B, right? Then you have graph, like, uh, TensorFlow-

**Host** [16:26]
Mm-hmm

**George Hotz** [16:26]
... which will put A times B into a graph, and then will do absolutely nothing until, uh, you actually compile the graph at the end. Um, I like this third choice, just somewhere in the middle, laziness. Laziness is you don't know when the ops are gonna dispatch, and don't worry about that.

You don't have to worry about this as a programmer. You just write out all your stuff, and then when you actually type .numpy, it'll be ready by the time you, you know, copy the thing back to CPU. Or you can do .realize, and it will actually, like, force that tensor to be allocated in RAM.

**Host** [16:56]
Mm-hmm.

**George Hotz** [16:57]
Um, but yeah, a lot of times, right, like... And if you think about it, PyTorch is kind of lazy in a way, but they didn't extend the paradigm far enough, right? When I do A times B in PyTorch, it's going to launch a CUDA kernel to do A times B, but it's not gonna wait for that CUDA kernel to complete.

So you're getting the worst possible world. You're getting the same laziness, but you also can't get fusion because PyTorch-

**Host** [17:18]
Mm-hmm

**George Hotz** [17:18]
... doesn't know that I'm then gonna do plus C. There's no way for it to be like, "Whoa, whoa, whoa, don't launch that CUDA kernel. Whoa." "Let's do this one too." Right? Um, you can kinda... Like, again, this stuff, PyTorch is working on this and, uh, you know, it's, it's a little bit harder.

Like, in Comma, I felt like I was competing against a lot of idiots. Um, here I'm competing against, you know-

**Host** [17:37]
Smart-

**George Hotz** [17:37]
... smart, smart, very smart people who have made-

**Host** [17:39]
Programming language compiler people

**George Hotz** [17:40]
... yeah, who, who've made some, I think, different trade-offs, right? Who've made some different trade-offs. Whereas if you're trying to build something that is just straight up good on Nvidia, and we have a lot of people and complexity to throw at it, yeah, PyTorch made a lot of the right choices.

I'm trying to build something that manages complexity. Like, you can always make your software do more. The magic is when you can make your software do more without adding complexity, right? Um, 'cause, you know, complex things eventually collapse under their own weight.

**Host** [18:06]
Mm-hmm.

**George Hotz** [18:06]
So it's kind of that.

**Host** [18:07]
Yeah. How does fusing actually work? Like how-

**George Hotz** [18:10]
Like Tensor, TensorFlow actually collapsed under its own right? That's kinda, that's kinda what happened, right? Uh, how does fusing actually work? Um, so yeah, uh, there's this thing called lazy.py. Uh, and when you do like A times B, that's...

Uh, it's put into a graph, but it's a very, uh, local graph.

**Swyx** [18:27]
Mm-hmm.

**George Hotz** [18:27]
There's no global graph optimizations. And even this can change, right? Again, like the programming model for tinygrad does not preclude eagerness, right?

**Swyx** [18:35]
Mm-hmm.

**George Hotz** [18:35]
Laziness is not guaranteed laziness.

**Swyx** [18:37]
Hmm.

**George Hotz** [18:37]
It's just gonna try its best. Um, so you put in A times B, and that's a binary op, right? And then you put in A times B, like that's a node in the graph. It's a virtual node 'cause it's not realized yet, plus C.

Okay, here's a new node, which then takes the C tensor in here and takes the output of A times B. It's like, whoa, wait, there's two binary ops. Okay, we'll just fuse those together.

**Swyx** [18:55]
Mm-hmm.

**George Hotz** [18:55]
Okay, here I have a kernel. This kernel has A, B and C as inputs. It does A times B plus C in the local registers, and then outputs that to memory. And you can, uh, graph.one in tinygrad. Another, another, like, amazing thing that tinygrad has that I've not seen in any other framework is two things.

Uh, graph.one, graph equals one, which is an environment variable. It will output a complete graph-

**Swyx** [19:17]
Mm-hmm

**George Hotz** [19:18]
... of all the operations. Um, people are like, "Oh, you can use PyTorch, export it to ONNX and use Netron." Yeah, you can, but like what? You-- That's not what's real, right? Graph.one will show you the actual kernels that were dispatched to the GPU.

You can also type debug equals two, which will print those kernels out, uh, in your, in your, in your command line. And it will sh- tell you the exact number of flops and the exact number of memory accesses in each kernel.

So you can immediately see, wait a second, okay, this kernel used this many flops. This was the gigaflops. This is how many bytes it read, and this was the gigabytes per second. And then you can profile without having to like...

Okay, I mean, in theory, in PyTorch, sure, use the NVIDIA Insight Profiler-

**Swyx** [19:58]
No one does that

**George Hotz** [19:59]
... which, uh, no one does-

**Swyx** [19:59]
No one does that

**George Hotz** [19:59]
... of course because it's so difficult, right? Like, like actually NVIDIA used to, uh, pre, pre... I think CUDA 9 was the last one that had it. They had a command line one, but now it's like, okay, I'm gonna generate this blob, use this NVIDIA GUI tool to convert it into a Chrome trace, and then load it in Chrome.

**Swyx** [20:15]
Mm-hmm.

**George Hotz** [20:15]
Yeah, no one does it, right? Um, just type debug equals two in any tinygrad model, and it will show you all the kernels that it launches and the efficiency of each kernel, basically.

**Swyx** [20:25]
Yeah. This is something that John Carmack has, uh, often, uh, commented about, is like when you code, you need to build in your instrumentation or observability-

**George Hotz** [20:33]
Yeah

**Swyx** [20:33]
... right into, into that. I wonder if whatever John is d- working on, he's adopting this style. And, uh, maybe we can sort of encourage it by, by, by like, I don't know, s- naming it and coin- coining a, a certain kind of debugging style.

**George Hotz** [20:46]
If he would, if he would like to start contributing to tinygrad, I'd be, uh-

**Swyx** [20:49]
You should hook up with him. I know-

**George Hotz** [20:50]
... I'd be so happy.

**Swyx** [20:51]
What's he-

**George Hotz** [20:51]
I've, I've, I've chatted, chatted with him a few times. I'm not really sure what his company's doing.

**Swyx** [20:54]
Yeah.

**George Hotz** [20:54]
Um, I think it's all... I think it's, it's pretty, uh, uh... But no, I mean, hopefully, like we get tinygrad to a point where people actually want to start using it. Um, so tinygrad right now is uncompetitive on, uh, it's uncompetitive on NVIDIA and it's uncompetitive on x86.

Yeah.

**Swyx** [21:11]
And specifically, what do you care about when you say uncompetitive?

**George Hotz** [21:13]
Uh, speed.

**Swyx** [21:14]
Okay.

**George Hotz** [21:14]
Straight up speed. Uh, it's correct. The correctness is there. The correctness for both forwards and backwards passes is there. But on NVIDIA, it's about 5x slower than PyTorch right now.

**Swyx** [21:22]
Mm-hmm.

**George Hotz** [21:23]
You're like, "5x, wow, this is, this is insurmountable."

**Swyx** [21:25]
Mm-hmm.

**George Hotz** [21:25]
No, there's reasons it's 5x slower, and I can go through how we're gonna make it faster. And it used to be, you know, 100x slower, so you know, we're making progress. But, um, there's one place where it actually is competitive, and that's Qualcomm GPUs.

Uh, so tinygrad is used to run the model in OpenPilot. Like right now, it's been live in production now for, for six months.

**Swyx** [21:41]
Mm-hmm.

**George Hotz** [21:42]
Um, and tinygrad is about 2x faster on the GPU than Qualcomm's library. Um, and-

**Swyx** [21:47]
And, and why specifically Qualcomm?

**George Hotz** [21:49]
Well, 'cause we have Qualcomm. We use Qualcomm in the, uh, comma devices.

**Swyx** [21:54]
Oh, I mean like what, what makes, what makes... What's, what about Qualcomm ar- architecture?

**George Hotz** [21:57]
Oh, what, what makes it doable?

**Swyx** [21:58]
Yeah, yeah.

**George Hotz** [21:59]
Well, because the world has spent how many millions of man-hours-

**Swyx** [22:01]
Sure. Okay

**George Hotz** [22:01]
... to make NVIDIA fast.

**Swyx** [22:02]
Yeah.

**George Hotz** [22:02]
And Qualcomm has a team of 10 Qualcomm engineers. Okay. Well, who can I beat here? Let's... Like, like what I propose with- What I propose with tinygrad is that developer efficiency is much higher.

**Swyx** [22:12]
Mm-hmm.

**George Hotz** [22:12]
But even if I have 10x higher developer efficiency, I still lose on NVIDIA, right?

**Swyx** [22:18]
Mm-hmm.

**George Hotz** [22:18]
You know, okay, I didn't put 100,000 man-hours into it, right?

**Swyx** [22:21]
Mm-hmm.

**George Hotz** [22:21]
If they put a million, like, like that's what I'm saying. But that's what I'm saying we can get. And we are gonna close this speed gap a lot. Like I don't support Tensor Cores yet. That's, that's, that's a big one-

**Swyx** [22:31]
Mm-hmm

**George Hotz** [22:31]
... that's just gonna, okay, massively close the gap. And then AMD, uh, I can't even get... I don't even have a benchmark for AMD because I couldn't get it compiled.

**Swyx** [22:40]
Mm-hmm.

**George Hotz** [22:40]
Oh, and I tried. Oh, I tried. I spent a day. Like, I spent actually a day trying to get PyTorch, and I got it built. I got it kind of working. Then when I tried to run a model, like there's all kinds of weird errors, and the, then the rabbit hole's just so deep on this.

I'm like... Um, so we... You know, you can compare the speed. Right now, you can run Llama, you can run anything you want on AMD. It already all works. Any OpenCL backend works, and it's not terribly slow. I mean, it's a lot faster than crashing.

So it's, uh, infinitely times faster- ... than PyTorch on AMD. Um, but pretty soon we're gonna start getting close to theoretical maximums, uh, on AMD. That's really where I'm pushing, and I wanna get AMD on MLPerf, uh, in a couple months hopefully.

**Swyx** [23:19]
Mm-hmm. Now that you bring up AMD.

### AMD Saga

**Alessio** [23:20]
Yeah, let's dive into that because when you announced the tinycorp fundraise-

**George Hotz** [23:24]
Yeah

**Alessio** [23:24]
... you mentioned one of your first goals is like build the framework runtime and driver for, for AMD. Uh, and then on, on June 3rd on Twitch, uh, you weren't as excited about AMD anymore. Maybe let's talk a bit about that and like, uh, you compared the quality of like commit messages from like the AMD kernel to like the Intel work that people are doing there.

What's important to know?

**George Hotz** [23:45]
So when I said I wanted to wr- I wanna write a framework, I did never intended on writing a kernel driver.

**Alessio** [23:49]
Mm-hmm.

**George Hotz** [23:50]
I mean, like I flirted with that idea briefly, but like realistically, I, I... Like there's three parts to it, right? There's like the ML framework, there's the driver, and then there's the user space runtime.

**Alessio** [24:01]
Mm.

**George Hotz** [24:01]
I was even down to rewrite the user space runtime. I have, I have a GitHub repo called CUDA IO Control Sniffer. It's terribly called, but you can actually launch a CUDA kernel without CUDA So you don't need CUDA installed.

Just the NVIDIA open source driver and this open source repo can launch a CUDA kernel.

**Alessio** [24:16]
That's cool.

**George Hotz** [24:17]
So rewriting the user space runtime is doable.

**Alessio** [24:19]
Mm-hmm.

**George Hotz** [24:19]
Rewriting the kernel driver-

**Alessio** [24:21]
Right

**George Hotz** [24:21]
... I don't even have docs. I don't have any docs for the GPU. Like, it would just be a massive reverse engineering project.

**Alessio** [24:26]
Mm-hmm.

**George Hotz** [24:26]
Um, so that is... When I saw that there... Like, it wasn't, like, I wasn't complaining about it being slow. I wasn't complaining about PyTorch not compiling. I was complaining about the thing crashing my entire computer. It panics my kernel, and I have to wait five minutes while it reboots 'cause it's a server motherboard and they take five minutes to reboot.

Um, so I was like, "Look, if you guys do not care enough to get me a decent kernel driver, there's no way I'm wasting my time on this, especially when I can use Intel GPUs." Intel GPUs have a stable kernel driver, and they have all their hardware documented.

You can go and you can find all the register docs on Intel GPUs. So I'm like, "Why don't I just use these?" Now, there's a downside to them. Uh, their GPU is $350, and you're like, "What a deal.

It's $350." You know, when you get about $350 worth of performance. And if you're paying about 400 for the PCIe slot to put it in, right? Like between the power and the-

**Alessio** [25:11]
Mm-hmm

**George Hotz** [25:11]
... and all the other stuff, you're like, "Okay, never mind." You gotta use Nvidia or AMD, um, from that perspective. But I sent an email to Lisa Su, and she responded.

**Alessio** [25:20]
Nice.

**Swyx** [25:20]
Oh. Because you published that email in, in a Discord and-

**George Hotz** [25:23]
I did. I did.

**Swyx** [25:23]
Yeah.

**George Hotz** [25:23]
And she responded. Um, and I've had a few calls since.

**Swyx** [25:28]
Ooh.

**George Hotz** [25:28]
And like, what, what I did was like... What I tried to do... Well, first off, like, thank you for responding. It shows me that like, like if you don't care about your kernel panicking, I, I can't. Like, like this-

**Swyx** [25:39]
Mm-hmm

**George Hotz** [25:40]
... is just a huge waste of my time, right? I'll find someone who will care. Like I do... I'm not asking for your seven by seven Winograd convolution when transposed to be fast. Like, I'm not asking for that.

I'm asking literally for-

**Swyx** [25:51]
The basics of getting it done. Yeah

**Alessio** [25:52]
To not die, yeah.

**George Hotz** [25:52]
Oh, and this isn't Tinygrad. This is your demo apps. I ran their demo apps in loops, and I got kernel panics.

**Alessio** [25:58]
Hmm.

**George Hotz** [25:58]
I'm like, you know, okay. There's a... Um, but no. Uh, uh, Lisa Su reached out, connected with a whole bunch of different people. Uh, they sent me a pre-release version of Rockm, uh-

**Swyx** [26:11]
Mm-hmm. Yeah

**George Hotz** [26:11]
... 5.6. They told me, "You can't release it," which I'm like, "Okay, so why do you, why do you care?" But, um, they say they're gonna release it by the end of the month, and it fixed the kernel panic.

The guy managed to reproduce it, uh, with the two GPUs in the computer. Uh, and yeah, sent me a driver and it works. So, um, yeah, I, I had, I had that experience. Uh, and then I had another experience where I had two calls with like AMD's, like, communication people, and just like...

I tried to explain to these people like open source culture.

**Swyx** [26:39]
Mm-hmm.

**George Hotz** [26:39]
Like, it's not open source if you dump the source code on a GitHub repo- ... and then forget about it until the next release. It's not open source if, you know, all your issues are from 2022. Like, like, like it's just no one's gonna contribute to that project, right?

Sure, it's open source in a very, like, technical sense. To be fair, it's better than nothing. It's better than nothing, but... Um, I fixed a bug in NCCL. I fixed a... There's a fun fact, by the way. Uh, if you have a consumer G- a consumer AMD GPU, they don't support peer-to-peer.

**Swyx** [27:07]
Hmm.

**George Hotz** [27:07]
Um, and their all-reduce bandwidth is horrendously slow because it's using CUDA kernels to do the copy between the GPUs, and it's putting so many transactions on the PCIe bus that it's really slow.

**Swyx** [27:17]
It is. Mm-hmm.

**George Hotz** [27:17]
But you can use cudaMemcpy, and there's a flag to use cudaMemcpy, but that flag had a bug. Um, so I, I posted, uh, the issue on NCCL. I expected nothing to happen. The Nvidia guy replied to me within an hour.

He's like, "Try this other flag." I'm like, "Okay, I tried the other flag. It still doesn't work, but here's a clean repro." And I spent like three hours writing a very clean repro. Um, I ended up tracking the issue down myself, but just the fact that somebody responded to me within an hour and cared about fixing-

**Swyx** [27:44]
Mm-hmm

**George Hotz** [27:45]
... the issue, okay, you've shown that it's worth my time.

**Swyx** [27:47]
Mm-hmm.

**George Hotz** [27:48]
And I will put my time in.

**Swyx** [27:49]
Yeah.

**George Hotz** [27:49]
Because, like, let's make this better. Like, I'm here to help. Um, but if you show me that, you know, you're like, "Yeah, the kernel panics, that's just, like, expected," okay.

**Swyx** [27:57]
Well, it sounds like AMD's getting the message.

**George Hotz** [27:59]
They are, and I just, I don't really think they've had someone explain to them like, like I s- I was like, "You gotta like build in public." And they're like, "What's an example of building in public?" I'm like, "Go look at PyTorch."

Go look at PyTorch, right? Like, you know, I have, I have, I have, I have two minor things merged into PyTorch because it's very responsive, you know. They're like minor bug fixes, but I feel like it's... You know.

**Alessio** [28:18]
Yeah. Um, so that's kinda like the lowest level of the stack, and then at a slightly higher, uh, level, obviously, there's Tinygrad, there's Mojo, uh, there's GGML. How are you thinking about breadth versus like depth and like where you decided to focus early on?

### GGML & Quantization

**George Hotz** [28:33]
Um, so GGML is very much like a, "Okay, everyone has M1s," right?

**Alessio** [28:37]
Mm-hmm.

**George Hotz** [28:37]
Actually, I was thinking, in the beginning, I was thinking of something more like GGML focused on the M1s, but GGML showed up and was just like, "We're actually just focusing on the M1s."

**Alessio** [28:46]
Mm-hmm.

**George Hotz** [28:47]
Um, so... And actually, M1 PyTorch is considerably better than AMD PyTorch. M1 PyTorch works. It only gives wrong answers sometimes, and it only crashes sometimes. But like some models kinda run. Um, when I was writing the, uh, Metal backend, I was comparing to MPS PyTorch, and I had like a, I had a, a discrepancy.

Like Tinygrad checks all its outputs compared to Torch, and I had one where it didn't match.

**Alessio** [29:11]
Mm-hmm.

**George Hotz** [29:11]
I'm like, I really, I, I, I checked the matrix by hand. It matches Tinygrad. I, I don't understand. And then I switched PyTorch back to CPU, and it matched. And I'm like, "Oh." Yeah. Well, there's like bugs, like if you like transpose the matrix because like I think this like has to do with like multi-views in PyTorch and like weird under-the-hood stuff that's not exposed to you.

Like there, there's bugs, and maybe they fixed them, but like, you know, it seems like there was a lot of momentum, again, because you're getting a huge vari- you're getting... How many engineers care about making PyTorch work on M1, right?

Thousands. Tens of thousands.

**Alessio** [29:41]
Mm-hmm. Yeah.

**George Hotz** [29:42]
And you have an open development process, and-

**Alessio** [29:44]
Yeah

**George Hotz** [29:44]
... guess what? It's gonna be good. How many engineers care about AMD working, PyTorch AMD working? Uh, you got 10 guys that work for AMD, and then like a couple hobbyists.

**Swyx** [29:54]
You revealed a, an interesting detail about how you debug, uh, which is you check, you, you hand check the matrix math?

**George Hotz** [30:00]
No, I don't hand check it.

**Swyx** [30:01]
Oh, okay.

**George Hotz** [30:01]
There's a, there's a... One of the best tests in Tinygrad is a file called test_ops.py.

**Swyx** [30:06]
Uh-huh.

**George Hotz** [30:06]
And it's just 100 small examples written in Tinygrad and PyTorch, and it checks-

**Swyx** [30:13]
Mm

**George Hotz** [30:13]
... both the forwards and backwards to make sure they match.

**Host** [30:15]
The test suite

**George Hotz** [30:16]
Yeah

**Host** [30:16]
Very important.

**George Hotz** [30:17]
That's, I mean, that's one of them where you like... I, I really, I put a lot of effort into CI for tinygrad. I think CI is super important. Like, I want that green check to mean I can merge this.

**Host** [30:25]
Yeah.

**George Hotz** [30:26]
Right? Like, I don't want my tests to... And if the green check, if you somehow manage to introduce a bug and get the green check, okay, we're fixing the test, top priority.

**Host** [30:32]
Yeah. Uh, Mojo?

**George Hotz** [30:35]
It's closed source. Uh, no, I'm not that interested, you know what I mean?

**Host** [30:39]
Yeah.

**George Hotz** [30:39]
Like, like, look, I, I, I like Chris Lattner. I, I think he's gonna do great things, and I understand the, uh, the like, kind of the wisdom even in keeping it closed source. But, uh, you know, I'm interested when it's open.

**Host** [30:49]
Yeah.

**George Hotz** [30:49]
Right?

**Host** [30:50]
It, it... You have an interesting design, uh, deviation from him, 'cause he's decided to be a, well, promised-

**George Hotz** [30:55]
Yes

**Host** [30:55]
... to be a superset of Python.

**George Hotz** [30:56]
Yes.

**Host** [30:57]
And you have decided to break, uh, with, with PyTorch APIs. Uh, and I think that's, that affects learnability and, and trans-transportability of code.

**George Hotz** [31:06]
You know, if the PyTorch thing ends up being like a, like a stumbling block, I, I could write a perfect PyTorch. Uh, like I'd, like I'd, like I'd... You know-

**Host** [31:17]
Mm-hmm

**George Hotz** [31:17]
... instead of import PyTorch, uh, instead of like, yeah, import torch, you type import tiny torch as torch. And if, if that really becomes the stumbling block-

**Host** [31:25]
Okay

**George Hotz** [31:25]
... I, I will do that. Um, no, Chris Lattner went much further than PyTorch.

**Host** [31:30]
Mm-hmm.

**George Hotz** [31:30]
Replicating the PyTorch API is something I can do with a couple, you know, like an engineer month or two.

**Host** [31:35]
Like a shim, yeah.

**George Hotz** [31:35]
Right, like a shim, yeah. Um, replicating Python? There's a, there's a, there's a big graveyard of those projects. How's, how's, uh, how's Piston going? How's, oh, Jython? How's-

**Host** [31:47]
Yeah, you can go back

**George Hotz** [31:48]
... PyPy?

**Host** [31:48]
Yeah.

**George Hotz** [31:49]
It's all... You can go way back. So-

**Host** [31:52]
Um, so tinygrad and Smalllayer, um, you announced TinyBox recently-

**George Hotz** [31:56]
Mm-hmm

**Host** [31:56]
... which is, um, you know, you made it... So your core mission is, uh, commoditizing the petaflop.

**George Hotz** [32:01]
Yeah.

**Host** [32:01]
Um, and then your business goal is to sell computers for more than it costs to make, which-

**George Hotz** [32:05]
Yeah

**Host** [32:06]
... seems super reasonable. Uh, what are... And you're gonna have three TinyBoxes? Red-

**George Hotz** [32:10]
No, no, no

**Host** [32:11]
... green, blue?

**George Hotz** [32:11]
No, no, no, no, no, no, no. That was my... Look, you know, a lot of people... Like, I love, you know, leaning into like saying I'm giving up, right? It's great to give up, right? Giving up is this wonderful thing.

It's so liberating.

**Host** [32:21]
Mm-hmm.

**George Hotz** [32:22]
And then, like, you can decide afterward if you really give up or not. There's very little harm in saying you give up, except like, you know, great, Twitter haters have something to talk about, and all press is good press, kids, so...

**Host** [32:31]
Um, so uh, obviously-

**George Hotz** [32:33]
Just red. Only red.

**Host** [32:35]
Yeah.

**George Hotz** [32:35]
TinyBox red.

**Host** [32:36]
TinyBox red.

**George Hotz** [32:37]
Unless AMD, you know, upsets me again, and then we're, you know, we're back to-

**Host** [32:41]
Yeah

**George Hotz** [32:41]
... we're back to other colors. We have other colors to choose from.

**Host** [32:44]
When you think about hardware design, what are some of the numbers you look for? So teraflops, it's per second is one, uh, but like memory bandwidth is-

**George Hotz** [32:52]
Mm

**Host** [32:52]
... another big limiter. Like, h-how do you make those trade-offs?

**George Hotz** [32:55]
Well, I mean, fundamentally I'm limited to what GPUs I can buy. But, uh, yeah, for, for something that I think a lot of people are going to wanna reasonably do with, um... Uh, a, a coworker of mine described them as luxury AI computers, right?

Like, luxury AI computers-

**Host** [33:10]
Mm-hmm

**George Hotz** [33:11]
... for, for people, and that's like what we're building, and I think a common thing people are gonna wanna do is run like Large LLAMA, right? Or large like Falcon or whatever.

**Host** [33:17]
FP16 LLAMA.

**George Hotz** [33:18]
FP16, exactly. Exactly. Um, you know, I... Int8 I think can work. I think that like what GGML is doing to go to like Int4, like this doesn't work. Like, have you done... And maybe they have, but like I, I read what it was, and I was like, "This isn't from any paper."

This is just some, like you, you made-

**Host** [33:34]
Squeezing as much as possible.

**George Hotz** [33:35]
Yeah, you made up some quantization standard to make it run fast. And like, like maybe it works, but okay, where's like the HellaSwag number, right? Where's your, where's your, where's your, uh, you know, all your-

**Host** [33:45]
The, the thesis-

**George Hotz** [33:46]
... big bench

**Host** [33:46]
... is right that-

**George Hotz** [33:46]
Yeah

**Host** [33:46]
... like if you have billions, hundreds of billions of parameters, that the individual quantization doesn't actually matter that much.

**George Hotz** [33:52]
Well, the, the real way to look at all of that is to just say you wanna compress the weights, right? It's a form of weight compression. Quantization is a form of weight compression, right? Now, this is obviously not lossless.

It's not a lossless compressor, right? If it's a lossless compressor, and you can show that it's correct, then okay, we don't have to have any other conversation. But it's a lossy compressor.

**Host** [34:07]
Yes.

**George Hotz** [34:07]
And how do you know that your loss isn't actually losing the power of the model?

**Host** [34:11]
Interesting.

**George Hotz** [34:11]
Maybe, maybe Int4 65B LLAMA is actually the same as FP16 7B LLAMA, right? Uh, we don't know. Uh, maybe someone has done this yet, but I looked for it when it like first came out-

**Host** [34:22]
Yeah

**George Hotz** [34:22]
... and people were talking about it, and I'm like, I just have... Like, it's not from a paper, right? The Int8 stuff is from a paper where they... Like, some of the Int8 stuff is from a paper.

**Host** [34:30]
Mm-hmm.

**George Hotz** [34:30]
There's one paper, I think it's like Int8, llm.int8, where they actually, uh, you know, do all the tests, and they didn't go fully Int8. They, they made like 90% of it Int8 and kept like 10% of it in FP16 for what they called like the, like outliers or whatever.

Um, so I think that this is not quite so easy, and I think being able... Well, so first off, if you're training, no one's gotten training to work with Int8 yet. There's a few papers that vaguely show it.

But if you're training, you're gonna need, uh, BF16 or, or float16. Um, so this is why I target that. Now, the thing that you're gonna wanna do is run these large language models out of the box on your hardware in FP16, and that's memory bandwidth.

So-

**Host** [35:08]
Mm

**George Hotz** [35:09]
... you, you, you need, you need large amounts of memory bandwidth too. Uh, so ask how I trade off memory bandwidth and flops, well, what GPUs can I buy?

**Host** [35:15]
Mm-hmm.

**George Hotz** [35:16]
But, um-

**Host** [35:18]
And I, I saw one of your bou- So first of all, you have this, um, hiring process, which is you gotta solve one of the bounties, um, that are open on tinygrad. There's no, uh, technical interview. One of them is Int8 support.

Do you already have some things you wanna test on?

**George Hotz** [35:32]
Uh, we have Int8 support. Um, what I'd like to see somebody do is just load the GGML Int8 LLAMA into tinygrad and then benchmark it against the FP16 one. Uh, Int8 already works in, in tinygrad. It doesn't actually do the math in Int8, which is even a, which is even a stronger...

Like, it does all the math still in FP32.

**Host** [35:51]
Mm-hmm.

**George Hotz** [35:52]
So Int8 can mean you just have your weights in Int8, or Int8 can mean you actually do your math in Int8. And doing your math in Int8, the, the big like gain that people care about is actually, uh, having your weights in Int8, because weights in Int8 mean less memory and less memory bandwidth.

Uh, whereas the math, keep it in FP32. With, with, with, on, on M1s, it doesn't even matter if you're doing... It doesn't matter what data type you're doing in the, for the, in the GPO. I, I'm not even sure it can do Int8, but FP16 and FP32 is the same.

It's the same teraflops.

Um, so yeah, no, that's one of the bounties. One of the bounties is get, get int8 Llama running with the int8 weights. And then-

**Swyx** [36:30]
Mm-hmm

**George Hotz** [36:30]
... actually, you don't even need to... What you could even do, if you really wanna test this, just take the FP16 weights, convert them to int8, then convert them back to FP16, then compare the unconverted and converted.

**Swyx** [36:41]
Oh, that's a nice hack.

**George Hotz** [36:42]
Right?

**Swyx** [36:42]
Oh, yeah.

**George Hotz** [36:43]
Right? Like, like, like I was-

**Swyx** [36:43]
This should be lossless f- in the other direction.

**George Hotz** [36:46]
Yeah. It, it will. Yeah, it's, uh, uh, yeah, I think FP16, it should be lossless in the other direction. I'm actually not 100% about that.

**Swyx** [36:53]
Why not?

**George Hotz** [36:54]
Uh, oh, 'cause like, you ever try to like, like if you wanna represent... If it was like int16, it's not lossless.

**Swyx** [37:00]
Sure.

**George Hotz** [37:00]
I think, I think all of int8 can be represented in FP16, but I'm not 100% about that.

**Swyx** [37:06]
Okay, so-

**George Hotz** [37:06]
Uh, actually, I think it, we-

**Swyx** [37:08]
Just draw out the bytes and

**George Hotz** [37:09]
We just have to do it, right? Just literally do it. There, there's only 256 to check. Like, um, but yeah, ei-either way, or I mean int4 definitely. So do your int4, convert it back, and now see even with int4 weights and FP32 math, like okay, how much does your performance degrade of this model?

**Swyx** [37:26]
Yeah.

**George Hotz** [37:26]
Yeah.

**Swyx** [37:27]
So, so can, can we s-s, uh... I, I, I'm about to zoom out a little bit from the, from the details. I don't know if you, you had more to, to-

**Alessio** [37:33]
No, I think like the-- you're planning to release the first tinybox, ship them in like two to six, eight months, something like that. Uh, what's top of mind for you in terms of building the team? Who should, who are you calling for?

### Tinybox Hardware

**George Hotz** [37:46]
Yeah. Uh, well, like to, to, to stay on the tinybox for, for, for, for one minute.

**Alessio** [37:50]
Yeah, exactly.

**Swyx** [37:50]
Yeah.

**George Hotz** [37:50]
Um, so if the GPU's picked out and you're like, "Well, I could make that computer with the GPUs," and my answer is, "Can you?" Do you know how to put... Do you know how hard it is to put six GPUs in a computer?

And people think it's really easy, and it's really easy to put one GPU in a computer.

**Alessio** [38:06]
Mm-hmm.

**Swyx** [38:06]
It does.

**George Hotz** [38:06]
It's really easy to put two GPUs in a computer, but now you wanna put in eight. Okay, so I'll tell you a few things about these GPUs. They take up four slots. What kind of computer... You, you can buy the nicest Supermicro.

You can't put eight of those in there. You need two-slot blowers.

**Alessio** [38:20]
Mm.

**George Hotz** [38:20]
If you wanna use one of those, those 4U Supermicros, you need two-slot blowers, right? Or water cooling, right? If, if you're trying to get the four-slot cards in there, you're gonna need some form of water cooling. Uh, or you're gonna need...

There are some like Chinese 4090s that are blowers, right? You either need blowers or water cooling if you're trying to get it in those things, right? Um-

**Swyx** [38:36]
So are you, are you doing water?

**George Hotz** [38:38]
No, I'm not using that chassis.

**Swyx** [38:40]
Okay.

**Alessio** [38:40]
Mm.

**George Hotz** [38:41]
Um, then the other thing that... Okay, so now you wanna get six GPUs in a computer, so that's a big challenge. And you're like, "Oh, I'll just use, uh, PCIe extenders. I saw it online as Tech Tips. It works great."

No, it doesn't. Try PCIe extenders that work at PCIe 4.0, and interconnect bandwidth's super important.

**Swyx** [38:56]
Yes.

**George Hotz** [38:57]
They won't work at 3.0. No PCIe extender I've tested, and I've bought 20 of them, uh, works at PCIe 4.0. So you're gonna need PCIe redrivers. Now, okay, how much is that adding cost, right? Like these things all get really hard.

And then tinybox, I've even added another constraint to it. I want this thing to be silent.

**Swyx** [39:16]
Yeah.

**George Hotz** [39:16]
Not totally silent, but my limit is like 45, maybe 50 dB. But not Supermicro machine, 60 dB . We have a small, we have a compute cluster at Comma.

**Alessio** [39:26]
Mm-hmm.

**George Hotz** [39:27]
You gotta wear, you gotta wear ear protection to go in there. Like it's-

**Swyx** [39:29]
Yeah. I've seen some videos where you give a tour.

**George Hotz** [39:31]
Oh, yeah.

**Swyx** [39:31]
It's, yeah.

**George Hotz** [39:32]
It's super-

**Swyx** [39:32]
Noisy

**George Hotz** [39:33]
... super loud, or you got all these-

**Alessio** [39:34]
Loud

**George Hotz** [39:34]
... just screaming-

**Alessio** [39:35]
Yeah

**George Hotz** [39:36]
... from like 10,000 RPM just screaming.

**Alessio** [39:39]
Mm-hmm.

**George Hotz** [39:39]
Like I wanna be able to use the normal big GPU fans and make this thing so it can sit under your desk, plug into one outlet of power, right? Six GPUs. But those, your, your, your GPUs are 350 watts each.

Can't plug that into a wall outlet. Okay, so how are you gonna deal with that? Mm, good questions, right? Um-

**Swyx** [39:59]
And you're not sharing them.

**George Hotz** [40:01]
Well, that one, I mean, that one is pretty obvious. You have to limit the power on the GPUs, right?

**Swyx** [40:05]
Yeah.

**George Hotz** [40:05]
Um, you have to limit the power on the GPUs. Now, you can limit power on GPUs and still get... You can, you can use like half the power and get 80% of the performance. Um, this is a known fact about GPUs, but like that's one of my design constraints.

**Swyx** [40:15]
Mm.

**George Hotz** [40:15]
So when you start to add all these design constraints, good luck building a tinybox yourself.

**Swyx** [40:20]
Mm-hmm.

**George Hotz** [40:21]
Um, you know, uh, obviously it can be done, but you, you need something that has actually quite a bit of scale and resources to do it.

**Alessio** [40:26]
Mm-hmm. And, and you see like the under the, the desk as like one of the main use cases, kinda like individual developer use or?

**George Hotz** [40:34]
Yeah. What I also see is more of a, like an AI hub for your home.

**Alessio** [40:37]
Mm.

**George Hotz** [40:38]
Right? As we start to get like home robotics kinda stuff, you don't wanna put the inference on the robot. But you also don't wanna put the inference on the cloud. Uh, well, you don't wanna put it on the robot because, okay, it's 1,500 watts, tinybox.

You'll put batteries? You gotta charge them. Bad idea. I mean, just, just, just wireless.

**Alessio** [40:56]
Mm-hmm.

**George Hotz** [40:56]
Wireless is, is .5 milliseconds.

**Alessio** [40:58]
Yeah.

**George Hotz** [40:58]
Right? This is super fast. Um, you don't wanna go to the cloud for two reasons. One, uh, cloud's far away.

**Alessio** [41:04]
Okay.

**George Hotz** [41:04]
It's not that far away. You can kind of address this. Uh, but two, cloud's also mad expensive.

**Alessio** [41:10]
Yeah.

**George Hotz** [41:10]
Like cloud GPUs are way more expensive-

**Alessio** [41:12]
Mm

**George Hotz** [41:12]
... than running that GPU at your house. At least any rates you're gonna get, right? Maybe if you commit to buy, well, yeah, I'm gonna buy 10,000 GPUs for three years, then maybe-

**Alessio** [41:20]
Yeah

**George Hotz** [41:20]
... the cloud will give you a good rate. But like you wanna buy, you wanna buy one GPU in the cloud? Whew. I mean, okay, you can go to like Vast, but like if you're talk- going on Azure or AWS, oh, that's expensive.

**Alessio** [41:29]
Yeah.

**Swyx** [41:29]
This is like a, like a personal data center, you know, instead of a-

**George Hotz** [41:33]
Oh

**Swyx** [41:33]
... cloud data center.

**George Hotz** [41:33]
We like the term compute cluster-

**Swyx** [41:35]
Compute cluster

**George Hotz** [41:35]
... so we can use Nvidia GPUs.

**Swyx** [41:37]
Yeah. Data center's maybe a little bit dated.

**George Hotz** [41:39]
It's a compute cluster-

**Swyx** [41:40]
Yeah

**George Hotz** [41:40]
... which is totally legal under the CUDA license agreement.

**Swyx** [41:43]
Uh, you, you talk a lot about the PCIe, uh, connection. Do you think there's any fat there to, to trim?

**George Hotz** [41:48]
What do you mean?

**Swyx** [41:49]
Uh, just y- you're limited by bandwidth, right?

**George Hotz** [41:52]
You... Okay, for some things, yes. Um, so the bandwidth is the, is roughly 10X less than what you can get with NVLink to A100s.

**Swyx** [42:01]
Yeah.

**George Hotz** [42:01]
Right? NVLink to A100s, you're gonna have... And then you can even get like full fabric, and then Nvidia really pushes on that stuff. Um, 600 gigabytes per second, right? And PCIe 4, you're gonna get 60, right? So you're getting 10X less.

**Swyx** [42:12]
Yeah.

**George Hotz** [42:13]
Um, that said, why do you need the bandwidth, right? And the answer is you need it for training huge models. If you're training on a tinybox, your limit's gonna be about 7 billion Right? If you're li- if you're training on big stuff, your limits could be like 70 billion, right?

Okay, you can hack it to get a bit higher. You can hack it like GPT hacked it to get a bit higher. But like that 65 billion in Llama, like there's a reason they chose 65 billion, right? And that's what can reasonably fit model parallel on, on-

**Host** [42:40]
Mm-hmm

**George Hotz** [42:40]
... on AGPUs, right? So, um, yes, you, you are going to end up training models. The cap's gonna be like seven billion. But I actually heard this on your podcast. I don't think that the best chatbot models are going to be the big ones.

I think the best chatbot models are gonna be the ones where you had 1,000 training runs instead of one. And I don't think that the interconnect bandwidth is going to matter that much.

**Host** [43:01]
So what are we optimizing for instead of compute optimal?

**George Hotz** [43:04]
Uh, what do you mean compute optimal?

**Host** [43:05]
Uh, so the, this- You're, you're talking about this, um, the Llama-style models-

**George Hotz** [43:09]
Yeah

**Host** [43:09]
... where you, you train for like 200 X.

**George Hotz** [43:11]
You train longer, yeah.

**Host** [43:12]
Yeah, yeah.

**George Hotz** [43:12]
Yeah, so okay. You can always make your model better by doing one of two things, right? And at Comma, we just have a strict limit on it. Um, you can always make your model better by training longer, and you can always make your model better by making it bigger.

**Host** [43:23]
Mm-hmm.

**George Hotz** [43:23]
But these aren't the interesting ones, right?

**Host** [43:26]
Mm.

**George Hotz** [43:26]
Particularly the making it bigger.

**Host** [43:27]
Mm.

**George Hotz** [43:27]
Because training it longer, fine. You know, you're getting a better set of weights. The inference is the same. The inference is the same whether I trained it for a day or a week.

**Host** [43:33]
Yeah.

**George Hotz** [43:34]
But the... Okay if I trained to 1 billion versus 10 billion, well, I 10x my inference too, right? So I think that these big models are kind of, uh, sure, they're great if you're research labs and you're trying to like max out this-

**Host** [43:45]
High DPI

**George Hotz** [43:45]
... hypothetical thing.

**Host** [43:46]
Which we can talk about later.

**George Hotz** [43:47]
Yeah, yeah, yeah. But if you're, but if you're like a startup or you're like an individual or you're trying to deploy this to the edge anywhere, eh, you don't, you don't need that many weights.

**Host** [43:56]
Yeah, yeah. Yeah.

**George Hotz** [43:57]
You actually don't want that many weights.

**Host** [43:58]
Optimizing for inference rather than capabilities-

**George Hotz** [44:00]
Yes

**Host** [44:01]
... doing benchmarks.

**George Hotz** [44:02]
Yes. Yes. Um, and I think the, the inference thing, right? There's gonna be so much more. Right now, the ratio between like training and inference on clouds, I think it's only still like, I think it's like two or three X, right?

It's two or three X more inference, which doesn't make any sense, right? There should be way more inference.

**Host** [44:15]
Yeah.

**George Hotz** [44:16]
There should be, uh, 10 to 100X more inference in the world than, than training. Um, but then also, like, what is training, right? You start to see these things like Laura, like, uh, you're get- you're getting kinda-

**Host** [44:26]
Mm-hmm

**George Hotz** [44:26]
... it's kinda blurring- ... the lines between inference and training, and I think that that blurred line is actually really good. I'd like to see much more like on-device training or on-device fine-tuning of the final layer.

**Host** [44:36]
Yeah.

**George Hotz** [44:36]
Um, where we're pushing toward this stuff at Comma, right? Like, why am I shipping a fixed model? I totally want this model to fine-tune based on like how, you know, your left tire is flat, right?

**Host** [44:45]
Mm-hmm.

**George Hotz** [44:45]
Like every time you cut the same turn because your left tire is flat, well, it should learn that, right? That's ML.

**Host** [44:50]
So would Comma pursue parameter efficient fine-tuning?

**George Hotz** [44:54]
Yeah. Yeah, yeah.

**Host** [44:55]
Right?

**George Hotz** [44:55]
We're, we're, we're-

**Host** [44:55]
Seems like a-

**George Hotz** [44:56]
We're, we're looking into stuff like that. I mean, Comma's already very parameter efficient because we have to like run this thing in a car, and you have to like cool it and power it.

**Host** [45:03]
Mm-hmm.

**George Hotz** [45:03]
Yeah.

**Host** [45:04]
Yeah. Yeah. And so this kind of like intelligence cluster you have in your home, you see when the person is using third-party model, they load them locally-

**George Hotz** [45:14]
Mm

**Host** [45:14]
... and kinda do the final fine-tuning. It kinda stays within the box.

**George Hotz** [45:17]
Yeah. I think that that's one thing, that's one version of it for the privacy conscious. Um, I also see a world where, uh, you can have your tinybox in its down cycles, um, mine FLOPcoin, right? You know, not all- Turns out not all crypto is a scam.

### FLOPcoin

**George Hotz** [45:32]
There, there's one way to tell if crypto is a scam. If they're selling the coin before they make the product, it's a scam.

**Host** [45:36]
Mm-hmm.

**George Hotz** [45:37]
If they have the product and then they sell the coin-

**Host** [45:39]
If they pre-mine, yeah

**George Hotz** [45:39]
... it's maybe not a scam, right? So yeah, my thought is like each tinybox would let you-- would have a private key on it. Uh, and you have to do it this way. You can't just let anyone join because of Sybil attacks, right?

**Host** [45:48]
Mm-hmm.

**George Hotz** [45:48]
There's a real problem of like-

**Host** [45:48]
Yeah

**George Hotz** [45:48]
... how do I, uh, how do I ensure your data's correct? And the way that I ensure your data is correct on the TinyNAT is if you ever send wrong data, you're banned from the network for life.

**Host** [45:57]
You're out. Oh, wow.

**George Hotz** [45:57]
Yeah, yeah. Your, your $15,000 hardware box is banned. So, you know, don't cheat.

**Host** [46:01]
Mm.

**George Hotz** [46:01]
Um, obviously, if it messes up, we'll forgive you. But, um, I'm saying like-

**Host** [46:04]
Somebody's gonna try to jailbreak your devices.

**George Hotz** [46:07]
There's no jailbreak. There's no jailbreak.

**Host** [46:08]
There's no jailbreak. There's just a different network.

**George Hotz** [46:10]
Well, there's just a private key on each device, right?

**Host** [46:11]
Yeah, yeah. Exactly.

**George Hotz** [46:11]
Like if you buy a tinybox from the tinycorp-

**Host** [46:13]
Makes sense

**George Hotz** [46:13]
... I give you a private key. It's in my backend server, right? You wanna hack my server, that's illegal.

**Host** [46:17]
Yeah.

**George Hotz** [46:17]
Anything you wanna do on the device, the device is yours.

**Host** [46:18]
Yeah.

**George Hotz** [46:18]
My server's mine, right? Like...

**Host** [46:21]
Yeah. Yeah. Uh, have you looked into like, uh, federated training at all?

**George Hotz** [46:26]
Yeah. So I mean, okay, you're, you're now... There's-- Okay, there's order of, of magnitude. Federated training, you mean like, uh, over the cloud and stuff?

**Host** [46:32]
Um-

**George Hotz** [46:32]
Over the internet?

**Host** [46:33]
Yeah. Over the internet, but also distributed on a bunch of devices, right?

**George Hotz** [46:36]
Yeah. I, I'm, I'm-

**Host** [46:38]
Which some people are experimenting

**George Hotz** [46:39]
... I'm very bearish on this stuff.

**Host** [46:40]
Yeah.

**George Hotz** [46:40]
Um, because your interconnect bandwidth, right? So, okay, at the high end, you have your interconnect bandwidth of NVLink, which is 600 gigabytes per second, right?

**Host** [46:47]
Yeah.

**George Hotz** [46:47]
The tinybox has 60 gigabytes per second, and then your internet has 125 megabytes per second.

**Host** [46:54]
Mm-hmm.

**George Hotz** [46:54]
Right? Not gigabits, 125 megabytes, right? So okay, that's, um-

**Host** [46:59]
Orders of magnitude.

**George Hotz** [47:00]
That's, that's how-

**Host** [47:00]
Three, four

**George Hotz** [47:01]
... that's how many orders of magnitude we're talking here, like from 60 down to 125.

**Host** [47:05]
Yeah.

**George Hotz** [47:05]
Like, all right, that's over a hund- that's over 100X.

**Host** [47:07]
Yeah.

**George Hotz** [47:07]
That's, that's 400X, right?

**Host** [47:08]
Yeah.

**George Hotz** [47:08]
So like, no. Uh, what-- but what you can do is inference, right? Like there's-- for inference, you don't care.

**Host** [47:13]
Mm-hmm.

**George Hotz** [47:14]
Right? For inference, I, I-- there's so little bandwidth at the top and the bottom of the model, um, that like, yeah, you can do federated inference, right? And that's kinda what I'm talking about. Um, th- there's also interesting things to push into, like you're like, but okay, what if you wanna run closed source models?

This stuff gets kind of interesting, like using TPMs on the boxes and stuff. Um-

**Host** [47:34]
Yeah.

**George Hotz** [47:34]
But then someone might jailbreak my device, so you know, maybe we don't try to do that.

**Host** [47:38]
Yeah. What's like the enterprise use case? Do you see companies buying a bunch of these and like stacking them together?

**George Hotz** [47:43]
Um, so the tinybox is like the first version of what we're build- building, but what I really wanna do is be on the absolute edge of FLOPS per dollar and FLOPS per watt.

**Host** [47:52]
Mm.

**George Hotz** [47:53]
Uh, these are the two numbers that matter. Uh, so the enterprise use case is you wanna train like, like Comma, right?

**Host** [47:58]
Mm-hmm.

**George Hotz** [47:58]
So Comma just built out a new compute cluster. It's about, uh, it's about a person and a half. Uh, so you know, it's, it's decent size. A person and a half.

**Host** [48:05]
A person being 20, uh, petaFLOPS.

**George Hotz** [48:07]
A person is, a person is 20 petaFLOPS.

**Host** [48:08]
Yeah.

**George Hotz** [48:08]
It's about 30 petaFLOPS. Um, we built, we built out a little, uh, a little compute cluster and, you know, we, we paid double what you theoretically could per FLOP, right? You theoretically could pay half per FLOP if you designed a bunch of custom stuff.

And yeah, I mean, I could see that being, you know, tinycorp, Comma's gonna be the first customer. I'm gonna build a box for Comma-

**Host** [48:28]
Mm-hmm

**George Hotz** [48:28]
... and then I'm gonna show off the box I built for Comma and be like, "Okay, like, do you wanna build... I sell $250,000 training computers." Or how much is one H100 box? Uh, it's, uh, it's 400 grand?

Okay, I'll build you a 400 grand training computer, and it'll be 10X better than that H100 box. For, again, not for every use case. For some, you need the interconnect bandwidth. But for 90% of most companies' model training use cases, the tinybox will be 5X faster for the same price.

**Alessio** [48:53]
Yeah. Awesome. You mentioned the person of compute. Um, how do we build a human for $20 million?

**George Hotz** [49:00]
Well, it's a lot cheaper now. It's a lot cheaper now. Uh, so like I said, we... Comma, Comma spent about, uh, about half a million on our, on our person and a half, so, you know.

**Alessio** [49:10]
Yeah. What are some of the numbers people should think of when they compare compute to, like, people? So, GPT-4 was 100 person years of training. That's more like on, on the timescale. Um, 20 petaFLOPS is one person. I think you, um...

Right now, the math was that for the price of the most expensive thing we've built, which is the International Space Station-

**George Hotz** [49:30]
Mm-hmm

**Alessio** [49:30]
... we could build, uh, one Tampa of, uh-

**George Hotz** [49:32]
Yeah, yeah. One, one Tampa of compute.

**Alessio** [49:34]
Which is 400,000 people.

**Swyx** [49:35]
Yeah, it's the ultimate currency of com-

**George Hotz** [49:36]
Yeah

**Swyx** [49:36]
... of, uh, measurement.

**George Hotz** [49:38]
Um, yeah, yeah. We could build... So, like, the biggest training clusters today, I know less about how GPT-4 was trained. I know some rough numbers on the weights and stuff, but, uh, Llama-

### GPT-4 & Bitter Lesson

**Alessio** [49:46]
A trillion parameters?

**George Hotz** [49:48]
Well, okay, so GPT-4 is 220 billion in each head, and then it's an eight-way mixture model. So mixture models are what you do when you're out of ideas. Um-

**Alessio** [49:57]
Yeah

**George Hotz** [49:57]
... so, you know, it's a, it's a mixture model. Uh, they just trained the same model eight times, and then they have some little trick. They actually do 16 inferences. But, um, no, it's not like-

**Alessio** [50:04]
So the multimodality is just a vision model kinda glom- glomed on.

**George Hotz** [50:09]
I mean, the multimodality is, like, obvious what it is too. You just put the vision model in the same token space as your language model.

**Alessio** [50:14]
Yeah, yeah, yeah.

**Swyx** [50:14]
Mm-hmm.

**George Hotz** [50:14]
Oh, did people think it was something else? No, no, the mixture has nothing to do with the vision or language aspect of it. It just has to do with, well, okay, we can't really make models bigger than 220 billion parameters.

Uh, we want it to be better. Well, how can we make it better? Well, we can train it longer, and okay, we're actually, we've actually already maxed that out. Uh, getting diminishing returns there. Okay.

**Alessio** [50:33]
A mixture of experts.

**George Hotz** [50:34]
Yeah, a mixture of experts. We'll train eight of them, right?

**Alessio** [50:36]
All right.

**George Hotz** [50:37]
So, all right. So, you know, you know... You know, the, the real truth is whenever a start-- whenever a company is secretive, with the exception of Apple, Apple's the only exception. Whenever a company is secretive, it's because they're hiding something that's not that cool.

**Alessio** [50:48]
Yeah.

**George Hotz** [50:48]
And people have this wrong idea over and over again that they think they're hiding it 'cause it's really cool. It must be amazing. It's a trillion parameters. No, it's a little bigger than GPT-3, and they did an eight-way mixture of experts.

Like, all right, dude, anyone can spend eight times the money and get that. All right. Um, but

yeah, so, uh, coming back to what I think is actually gonna happen is, yeah, people are gonna train smaller models for longer and fine-tune them and find all these tricks, right? Like, I, I, you know, I think, uh, OpenAI used to publish stuff on this, you know, uh, when they would publish stuff, uh, about how much better the training has gotten given the same...

Holding compute constant. And it's gotten a lot better, right? Think compare, like, batch norm to no batch norm.

**Alessio** [51:32]
Yeah.

**George Hotz** [51:33]
Right? And now we have, like-

**Alessio** [51:34]
'Cause you're finding algorithms like FlashAttention and-

**George Hotz** [51:36]
Yeah. Well, FlashAttention, yeah.

**Alessio** [51:38]
Yeah.

**George Hotz** [51:39]
Um, and FlashAttention's the same compute. FlashAttention's an interesting factor where it's actually the identical compute. It's just a more efficient way to do the compute. But I'm even talking about, like, like, um, look at the new, look at the new, uh, embeddings people are using, right?

They used to use these, like, boring old embeddings. Now, like, Llama uses that complex one, and now there's, like, Alibi. I'm not up to date on-

**Alessio** [51:57]
Mm-hmm

**George Hotz** [51:57]
... all the latest stuff, but, uh, those tricks give you so much.

**Alessio** [52:01]
There's been a whole round trip with positional embeddings. I don't know if you've, uh, seen this discussion.

**George Hotz** [52:06]
I haven't followed-

**Alessio** [52:06]
Like, you need them, you need rotational, and then you don't need them.

**George Hotz** [52:09]
I haven't followed exactly. I mean, you quickly run into the obvious problem with positional embeddings, which is you have to invalidate your KV cache if you run off the context.

**Alessio** [52:17]
Mm-hmm.

**George Hotz** [52:17]
So that's why I think these new ones, they're playing with them. But, uh, I, I'm not that, I'm not that up... I, I'm not an expert on, like, the latest up-to-date language model stuff.

**Alessio** [52:26]
Yeah.

**George Hotz** [52:26]
Um, I mean, we have what we do at Comma, and I know how that works, but, like...

**Alessio** [52:32]
Um, what are some of the things, I mean, that people are getting wrong? So back to autonomous driving, there was, like, the whole, like, LIDAR versus vision thing. You know, it's like people don't get into accidents because they cannot see well.

They get into accidents because they get distracted and, and all these things. What are... Do you see similarities today on, like, the path to AGI? Like, are there people... Like, what are, like, the-

**George Hotz** [52:53]
Nothing, nothing I say about this is ever gonna compete with how Rich Sutton stated it. Rich Sutton is writer of-

**Alessio** [52:59]
The Bitter Lesson. Yeah

**George Hotz** [52:59]
... Reinforcement Learning: The Bitter Lesson. Nothing I say is ever gonna compete with... The Bitter Lesson's way better than any way I'm going to phrase this. Just go read that and then, like, I'm sorry it's bitter, but you actually just have to believe it.

Like, you, over and over again, people make this mistake. They're like, "Oh, we're gonna hand engineer this thing. We're gonna hand..." No, like, stop wasting time.

**Alessio** [53:17]
Which is, I mean, OpenAI is not taking the bitter lesson.

**George Hotz** [53:21]
No. OpenAI-

**Alessio** [53:23]
They, they were, they were leaders in deep learning-

**George Hotz** [53:26]
Yes

**Alessio** [53:26]
... for a long, long, long time.

**George Hotz** [53:27]
OpenAI-

**Alessio** [53:27]
But you're telling me that G-GPT-4 is not. Yeah.

**George Hotz** [53:29]
Well, OpenAI was the absolute leader to the thesis that compute is all you need.

**Alessio** [53:33]
Yes.

**George Hotz** [53:33]
Right? Uh, and there's a question of how long this thesis is going to continue for, right? It's a cool thesis, and look, I think, um, I would be lying along with everybody else. I was into language models, like, way back in the day for the Hutter Prize.

I got into AI through the Hutter Prize. Like, 2014, I'm trying to build compressive models of Wikipedia, and I'm like, "Okay, why is this so hard?"

**Alessio** [53:52]
Mm.

**George Hotz** [53:52]
Like, what this is, is a, is a language model, right?

**Alessio** [53:54]
Mm-hmm.

**George Hotz** [53:54]
And I'm playing with these, like, like, Bayesian things, and I'm just like, "Oh, but, like, I get it. Like, it needs to be like, like it's like I have two data points, and they're, like, almost the same, but how do I measure that almost," right?

I just, like, you know, wrap my head around... I couldn't, like, like, wrap my head around this. And this was around the time Karpathy released the first, like, uh, RNN that generated the Shakespeare stuff. And I'm like, "Okay, I get it," right?

It's, it's, it's neural networks that are compressors. Now, this isn't actually... You can't actually win the Hutter Prize with these things because the Hutter Prize is, is, is MDL. It's the model, size of the model plus the size of the e-encodings, embeddings.

So yeah, you can't... I mean, probably now you can-

**Alessio** [54:31]
Mm-hmm

**George Hotz** [54:32]
... 'cause it's gotten so good. But, uh, yeah, back in the day, you kinda couldn't. So I was like, "Okay, cool. Like, this is what it is. I kinda get it." Um, yeah, I mean, I think I didn't expect that it would continue to work this well.

I thought there'd be real limits to how good autocomplete could get. That's fancy autocomplete.

**Alessio** [54:48]
Mm-hmm.

**George Hotz** [54:48]
Um, but yeah, no, like, it, it, it works, uh- It works well. So like, yeah, what is OpenAI getting wrong? Technically, not that much. I don't know. Like, if I was a researcher, why would I go work there?

Like-

**Alessio** [55:03]
Yes.

**George Hotz** [55:03]
Go-

**Alessio** [55:03]
So why, why is OpenAI like the Miami Heat?

**George Hotz** [55:07]
No, I... Look, I don't, I don't, I don't... This is, this is my technical stuff. I don't really wanna harp on this. But like- ... why go work at OpenAI when you could go work at Facebook, right? As a researcher.

Like, OpenAI can keep ideologues who, you know, believe ideological stuff, and Facebook can keep every researcher who's like, "Dude, I just wanna build AI and publish it."

**Alessio** [55:23]
Yeah.

**George Hotz** [55:24]
Yeah.

**Alessio** [55:25]
Awesome. Um, yeah. Any other thoughts? tinycorp, bounties.

### Hiring & AI Tools

**George Hotz** [55:33]
Um, yeah. So we have... You know, I've been thinking a lot about, like, what it means to hire in today's world.

**Alessio** [55:40]
Mm-hmm.

**George Hotz** [55:41]
W- what actually is the, like, core... Okay. Look, I'm a believer that machines are gonna replace everything in about 20 years. Uh, so okay. What is that, what is that thing that people can still do that computers can't, right?

Um, and this, this is a narrowing list. But like, you know, back in the day, like, imagine I was starting a company in 1960, right? "Oh, we're gonna have to hire a whole bunch of calculators in the basement to do all the, you know, math to support the calcu-"

**Alessio** [56:10]
Mm-hmm.

**George Hotz** [56:10]
"Dude, have you heard about computers?" "Dude, why don't we just buy a few of those?" "Oh. Oh, wow, man. You're right." Um, so like I feel like that's kinda happening again, and I'm thinking about... I will post to my Discord.

I'll be like, "Okay, who wants to, like..." Okay, I just changed my unary op. Used to be log and exp in like E. Um, I changed them to be log2 and exp2-

**Alessio** [56:31]
Mm-hmm

**George Hotz** [56:31]
... because hardware has log2 and exp2 accelerators.

**Alessio** [56:35]
Built-in. Yeah.

**George Hotz** [56:35]
Yeah, and of course, you can just change the base. It's one multiply to, to-

**Alessio** [56:37]
Yeah

**George Hotz** [56:37]
... get it back to E. But, like, I made the primitives log2 and exp2, right? And this is the kind of-- I just posted in the Discord. I'm like, "Can someone put this pull request up?" Right? And someone eventually did, and I merged it.

But I'm like, this is almost to the level where models can do it. Right? We're almost to the point where I can say that to a model, and the model can do it. Um-

**Alessio** [56:55]
Have you tried?

**George Hotz** [56:56]
Yeah. I'm... I don't know. I, I'm, like, I'm... I think it went further. I think autocomplete went further than I thought it would.

**Alessio** [57:06]
Mm-hmm.

**George Hotz** [57:06]
But I'm also relatively unimpressed with these chatbots-

**Alessio** [57:10]
Mm-hmm

**George Hotz** [57:10]
... uh, with what I've seen from the language models. Like, they're...

The problem is, if your loss function is categorical cross-entropy on the internet, your responses will always be mid.

**Alessio** [57:21]
Yes.

**George Hotz** [57:22]
Right?

**Alessio** [57:22]
Mode, mode collapse is what I call it. I don't know.

**George Hotz** [57:25]
Maybe... And I'm not even talking about mode collapse. You're actually trying to predict the... Like, like look, I rap. I'm a, I'm a hobbyist rapper, and, like, when I try to get these things to write rap, the raps sound like the kind of raps you read in the YouTube comments.

**Alessio** [57:35]
Yeah.

**George Hotz** [57:35]
Mm-hmm.

**Alessio** [57:35]
Nursery school.

**George Hotz** [57:35]
Yeah. It's like, all right, great. You rhyme box with fox. Sick rhyme, bro. Uh- You know. Uh, you know, and ri- or Drake is rhyming give it up for me with napkins and cutlery, right? Like-

**Alessio** [57:47]
Yeah.

**George Hotz** [57:47]
Like, all right, come on. What's-

**Alessio** [57:48]
He's, he's got, like, this thing about orange. They have-- Like orange is famous that you can't rap rhyme.

**George Hotz** [57:51]
Yeah, yeah, yeah, yeah. But now, of course, you know, four-inch screws and orange juice is in, is in-

**Alessio** [57:55]
In Jordan's

**George Hotz** [57:56]
... GPT's training corpus. Um, but, uh, yeah. So I, I think it went further than, like, everyone kinda thought it would. But the thing that I really wanna see is, like, somebody put 10 LLMs in a room and have them discuss the answer before they give it to me, right?

**Alessio** [58:09]
Mm-hmm.

**George Hotz** [58:09]
Like, you can actually, like, do this, right? Um, and I think the coding things have to be the same way. There is no coder alive, no matter how good you are, that sits down, "Well, I'm going to start at cell A1 and type my program, and then I'm going to press run and it's going to work."

No one programs like that.

**Alessio** [58:23]
Yeah.

**George Hotz** [58:24]
Um, so why do we expect the models to, right? So, so there's, there's a lot that, like, still needs to be done. But, you know, at the tinycorp, I wanna be on the cutting edge of this too. I wanna be like program generation.

I mean, what is tinygrad? It's a compiler. Generates programs. Generate the fastest program that meets the spec, right?

**Alessio** [58:38]
Mm-hmm.

**George Hotz** [58:38]
Why am I not just having ML do that? So, you know, it's kind of a... You have to exist fluidly with the machines. And I come around on a lot of stuff. I'm like, wait, tinygrad, tinycorp should be a remote company, right?

Why, why, I can't do this in person.

**Alessio** [58:53]
Really?

**George Hotz** [58:54]
Yeah. Like-

**Alessio** [58:54]
Oh

**George Hotz** [58:54]
... like, Comma makes sense to be in person. Like Comma, sure, yeah, we're getting an office in San Diego. Like, like, but that was a six-year-old company, right? And it works, and it works for a certain type of people and a certain type of culture, but what's gonna be different this time?

Okay, remote. But now it's remote. And now I'm getting these, like, people who apply, and I'm like, I literally have 1,000 applications. I'm not calling you to do a technical screen. I can't really tell anything from a technical screen.

**Alessio** [59:16]
Mm-hmm.

**George Hotz** [59:16]
What am I gonna do? Make you code on a whiteboard? Like, bring up, bring up a shared notebook document so we could... No, like, that's not gonna work. Um, okay, so then I move to the next thing. We do this at Comma with good success, programming challenges.

I've also found them to be, like, completely non-predictive. I found one thing to actually be predictive, and it's wait a second, just write code in tinygrad.

**Alessio** [59:36]
Yeah.

**George Hotz** [59:36]
It's open source.

**Alessio** [59:37]
Mm-hmm.

**George Hotz** [59:38]
Right? And yeah. Um, so you know, I'm, I'm talking to, to a few people who've been contributing. And like, contribute or, you know, the job's not for you. Um, but you can do it remote, and it's, look, it's a chill job.

Like, you're not... You're like, "Oh, yeah, well, I work for the tinycorp." Like, well, you're writing MIT licensed software. Like, you see what it's doing, right? Like, we'll just... Let's... I think, think of it as maybe more of like a stipend than a salary, and then also some equity.

Like if, you know, I get rich, we all get rich.

**Alessio** [1:00:00]
Yeah.

**George Hotz** [1:00:00]
Um, yeah.

**Alessio** [1:00:02]
How do you think about agents w- and kinda like thinking of them as people versus, like, job to be done? Um, Sian built this thing called Small Developer, um, and then-

**George Hotz** [1:00:12]
It's in the same vein.

**Alessio** [1:00:13]
Or-

**George Hotz** [1:00:13]
Like the, the human in the loop with the language model and just, uh, iter- iterating while you write code. Um, I think, I think that's, that's absolutely where it goes.

**Alessio** [1:00:20]
And there's like a... It's not, like, one thing. It's like there's small interpreter. There's, like, small debugger. It's kinda like all these different jobs to be done.

**George Hotz** [1:00:27]
It's a small world.

**Alessio** [1:00:28]
Yeah. It's a... I know. This is like-

**George Hotz** [1:00:30]
Tiny, tiny, small world

**Alessio** [1:00:30]
... the small pockets. It's like small AI meet tinycorp. So we're all on the same wavelength. How do you think about that? Do you think people will have a human-like interaction where it's like, "Oh, this is like the AI developer," or like is it, "I'm the human being supercharged by the AI tools"?

**George Hotz** [1:00:46]
Oh, I think it's, yeah, much more like I'm the human supercharged by the AI tools. I think that like coding is tool complete, right? Like driving is not tool complete, right?

**Swyx** [1:00:54]
Mm.

**George Hotz** [1:00:54]
Like driving is just like, like we, we hire people to drive who are like below the API line.

**Swyx** [1:00:57]
Mm-hmm.

**George Hotz** [1:00:58]
Right? There's an API line in the world, right?

**Swyx** [1:00:59]
Love that. Yes.

**George Hotz** [1:01:00]
Yeah, yeah. There's an API line in the world and like you can think-- like Uber is a really clear example, right?

**Swyx** [1:01:03]
Yes.

**George Hotz** [1:01:03]
There's the people below the API line and the people above the API line.

**Swyx** [1:01:06]
Yeah.

**George Hotz** [1:01:06]
And the way you can tell if you're below or above, by the way, is, is your, uh, manager a computer, right? Who's the manager of the Uber driver?

**Swyx** [1:01:12]
Yeah.

**George Hotz** [1:01:12]
Or computer, right?

**Swyx** [1:01:12]
Does the machine tell you what to do or do you tell machines what to do?

**George Hotz** [1:01:14]
Exactly. Exactly. Um, so coding is tool complete, right? Coding is tool complete. Coding is above the API line. So it will always be, uh, tools supercharging your coding workflow, and it will never be you performing some like task like, "Okay, well, I can do everything except for actually starting a Docker container."

Like it just doesn't make any sense, right? Um, yeah, so it will always be sort of tools. And, you know, look, we see the same stuff with all the... Like people are like, "Stable Diffusion's gonna replace artists," or whatever.

It's like, dude, like-

**Swyx** [1:01:46]
It's gonna create new artists.

**George Hotz** [1:01:47]
What-- Did Photoshop replace artists? Like what are you talking about, right? Like, oh, you know, a real artist finger paint. They can't use brushes. Brushes are, you know, brushes are gonna replace all the... Okay. Like I j- I just can't.

Like it's, it's all just tools, and the tools are gonna get better and better and better.

**Swyx** [1:02:03]
Yeah.

**George Hotz** [1:02:03]
And then eventually, yes, the tools are going to replace us. But you know, that's still 20 years away, so you know, I got a company to run in the meantime.

**Swyx** [1:02:10]
Yeah. So I, I've written about the API line before, and I-

**George Hotz** [1:02:12]
Yeah

**Swyx** [1:02:12]
... I, I think that's from Venkatesh. I don't know if you've, if actually-

**George Hotz** [1:02:15]
I don't know. I definitely took it from someone. It's definitely not mine.

**Swyx** [1:02:16]
It's VGR. Uh-

**George Hotz** [1:02:17]
Yeah

**Swyx** [1:02:17]
... yeah, but uh, I also have speculated a higher line than that, which is the Kanban board. Like who tells the, the programmers what to do.

**George Hotz** [1:02:24]
Hmm.

**Swyx** [1:02:25]
Right? So are you above or below the Kanban board? At d- does-- has that evolved your, your management thinking?

**George Hotz** [1:02:30]
Yeah. Like that's sort of what I mean. Like it's like a, like I'm gonna just gonna describe the pull request in two sentences and then like, yeah, you know.

**Swyx** [1:02:37]
Yeah. So you are running the Kanban board or the bounties or, you know?

**George Hotz** [1:02:39]
Yes.

**Swyx** [1:02:40]
Yes.

**George Hotz** [1:02:40]
Yeah. The, the bounties are the Kanban board.

**Swyx** [1:02:41]
Yes.

**George Hotz** [1:02:41]
Exactly. And, and that is kind of the, the high level and then like, yeah, we'll get AIs to fill in some, and we'll get people to fill in others, and yeah. And that, that's also what it means to be like full-time at tinycorp, right?

Would you start... And I, I wrote this up pretty concretely. I'm like, okay, step one is you do bounties for the company.

**Swyx** [1:02:56]
Mm.

**George Hotz** [1:02:57]
Step two is you propose bounties for the company.

**Swyx** [1:02:59]
Mm. Mm-hmm.

**George Hotz** [1:02:59]
Right? You don't obviously pay them. We pay them.

**Swyx** [1:03:01]
Mm.

**George Hotz** [1:03:01]
But you propose them, and I'm like, "Yeah, that's a good bounty." That like helps with the main workflow of the, the company. And step three is you get hired full-time, you get equity, we all, you know, maybe get rich.

Um-

**Swyx** [1:03:11]
What, what else are you des- designing differently about the employee experience?

**George Hotz** [1:03:16]
I mean, I'm very much a like, you know, some people really like to like, like keep a separation, right? Some people really like to keep a separation between like employees and management or customers and employees. Like at Comma, you know, the reason I do the dev kit thing, it's like, dude, you buy a Comma thing, you're an employee of the company.

Like you're just-

**Swyx** [1:03:33]
Mm

**George Hotz** [1:03:33]
... part of the company. It's a, it's all the same thing. There's no like secrets. There's no dividing lines. There's no like... It's all a spectrum for like, you know, down here at the spectrum, like you pay, and then up here at the spectrum, you get paid.

You understand this is the same spectrum of college, right? Like for undergrad, you pay, and then you get up here to like, you know, doing a PhD program, you get paid. Okay. Well, cool. Welcome to the, you know.

**Swyx** [1:03:53]
Mm-hmm.

What about, um, Comma Bodies? You know, you mentioned a lot of this stuff is clearly virtual, but then there's below the API line you actually need, uh-

**George Hotz** [1:04:05]
Yeah.

**Swyx** [1:04:05]
Wait, this is a thing that's been announced, Comma Bodies?

**George Hotz** [1:04:07]
We sell them.

**Swyx** [1:04:08]
Oh, okay.

**George Hotz** [1:04:08]
You can buy them.

**Swyx** [1:04:08]
And they're-

**George Hotz** [1:04:09]
They're a thousand bucks on our website.

**Swyx** [1:04:10]
Oh, okay. No, no, no. I, I'm thinking about like the, what Tesla announced with like the humanoid robots.

**George Hotz** [1:04:14]
It's the same thing.

**Swyx** [1:04:15]
Yeah, yeah, yeah. Okay.

**George Hotz** [1:04:15]
Except of course, we made the Comma version of it.

**Swyx** [1:04:17]
Yeah.

**George Hotz** [1:04:17]
Tesla uses 20 actuators. We use two, right?

**Swyx** [1:04:20]
Yeah.

**George Hotz** [1:04:20]
Like how do you, how do you build the simplest possible thing that can like turn the robotics problem into entirely a software problem?

**Swyx** [1:04:27]
Yeah.

**George Hotz** [1:04:27]
So right now it is literally just a Comma-3 on a pole with two wheels. Um, it balances, keeps the Comma-3 up there, and like there's so much you could do with that already, right? Like this should replace... How many security guards could this replace?

**Swyx** [1:04:42]
Mm-hmm.

**George Hotz** [1:04:42]
Right? If this thing could just competently wander around a space and take pictures and, you know, focus in on things, send you a text message when someone's trying to break into your building, you know, like, like this could already do so much.

Of course, but the software's not there yet.

**Swyx** [1:04:56]
Mm-hmm.

**George Hotz** [1:04:57]
Right? So how do we turn robotics into a thing where it's very clearly a software problem? You know, the people don't accept that self-driving cars are a software problem. Like I, I don't, I don't know what to tell you, man.

Like literally just watch the video yourself and then drive with a joystick, right?

**Swyx** [1:05:10]
Yeah.

**George Hotz** [1:05:11]
Can you drive? And we've actually done this test.

**Swyx** [1:05:13]
Mm-hmm.

**George Hotz** [1:05:13]
We've actually done this test where we've had someone, "Okay, you just watch this video and here's a joystick and you gotta drive the car." And of course, they can drive the car.

**Swyx** [1:05:19]
Yeah.

**George Hotz** [1:05:19]
It takes a little bit of practice to get used to the joystick, but, um, the problem is all the model, right? So, okay, now make the model better.

**Swyx** [1:05:26]
Yeah. Uh, and specifically anything in computer vision that you think, uh... Our, our second most popular episode ever was about Segment Anything coming out of, uh, Facebook, which is, as far as I understand, is state-of-the-art in computer vision.

Um, what are you hoping for there that, that you need for Comma?

**George Hotz** [1:05:42]
I haven't used Segment Anything. Like the large lo- large YOLOs or not. I've used like large YOLOs, and I'm super impressed by them.

**Swyx** [1:05:48]
Yeah.

**George Hotz** [1:05:49]
Um, I gotta check-

**Swyx** [1:05:50]
Do you think it's solved?

**George Hotz** [1:05:50]
I gotta check out Segment Anything. I don't think it's a distinct problem, right? Okay, here's something that I'm interested in. All right, we have great LLMs, uh, we have great text-to-speech models.

**Swyx** [1:05:59]
Yeah.

**George Hotz** [1:05:59]
And we have great speech-to-text models.

**Swyx** [1:06:00]
Yeah.

**George Hotz** [1:06:01]
Okay. So why can I not, why can I not talk to an LLM? Like I'd have a normal conversation with it.

**Swyx** [1:06:05]
You can with a latency of like two seconds every time.

**George Hotz** [1:06:07]
Right. Um, why, why isn't this... And then it feels so unnatural. It's this like staccato like... I don't like the RLHF models. I don't like the tuned versions of them.

**Swyx** [1:06:17]
Mm-hmm.

**George Hotz** [1:06:17]
I think that they become-- You take on the personality of a customer support agent, right?

**Swyx** [1:06:22]
Yeah.

**George Hotz** [1:06:22]
Like, oh, come on.

**Swyx** [1:06:23]
Yeah.

**George Hotz** [1:06:23]
You know? I, I sort of, I, I like, I like Llama more than ChatGPT.

**Swyx** [1:06:26]
Mm-hmm.

**George Hotz** [1:06:27]
ChatGPT's personality just grated on me. Whereas Llama like, "Cool. I write, I write a little bit of pretext paragraph. I can put you in any scenario I want," right? Like that's interesting to me. I don't want some like, you know.

Yeah. So, um, yeah, I think there is really no like distinction, uh, between computer vision and language and any of this stuff. Like- It's all eventually gonna be fused into one massive... So to say computer vision is solved, well, it doesn't make any sense, 'cause what's the output of a computer vision model?

Segmentation? Like what a weird task. Right? Who cares?

**Swyx** [1:07:00]
OCR.

**George Hotz** [1:07:01]
Who cares? I don't care if you can segment which pixels make up that laptop, I care if you can pick it up.

**Swyx** [1:07:06]
Yeah.

**Alessio** [1:07:06]
Mm-hmm.

**George Hotz** [1:07:07]
Right? Like-

**Swyx** [1:07:07]
Yeah, interact with the real world.

**George Hotz** [1:07:08]
Yeah.

**Alessio** [1:07:10]
And you're gonna have the local cluster, you're gonna have the, the body.

**George Hotz** [1:07:14]
Yeah. Yeah, yeah, I think, I think that's kinda where that goes.

**Swyx** [1:07:17]
So, so the, you have vision-- Like you- maybe we can paint the future of like the year is 2050.

**George Hotz** [1:07:22]
Yeah.

**Swyx** [1:07:23]
You've achieved all you wanted at tinycorp. What, what is, what is the AI-enabled future like?

**George Hotz** [1:07:29]
Well, tinycorp's the second company. Comma was the first. Comma builds the hardware infrastructure. tinycorp builds the software infrastructure. The third company's the first one that's gonna build a real product, and that product is, uh, a AI girlfriend.

### AI Girlfriend

**Swyx** [1:07:42]
Okay.

**George Hotz** [1:07:42]
No, like I'm dead serious, right?

**Swyx** [1:07:43]
Yeah.

**George Hotz** [1:07:43]
Like this is the dream product, right? This is the absolute dream product. Girlfriend is just the like-

**Swyx** [1:07:49]
Stand-in.

**George Hotz** [1:07:50]
Well, no, it's not a stand-in. No, no, no, no, I actually mean it, right? So I've been wanting to merge with a machine ever since I was like mad little. Like, you know, just like how do I merge with a machine, right?

And like you can look at like in, like a maybe the Elon style way of thinking about it is Neuralink, right?

**Swyx** [1:08:03]
Yeah.

**George Hotz** [1:08:03]
And I'm like, "I don't think we need any of this," right? You ever-- Some of your friends maybe, they get into relationships, and you start thinking of, you know, them and their partner as the same person. You start thinking of them as like one person.

**Swyx** [1:08:15]
Mm-hmm.

**George Hotz** [1:08:15]
Like they are kinda like merged, right? Like humans can just kinda do this. It's so cool. It's this ability that we already have, right? So I don't need to put, you know, electrodes in my brain to merge with a machine.

I need an AI girlfriend, right? So that's what I mean. Like this is, this is the third product. This is the third company. And yeah, in 2050, I mean like... Ah, it's, it's so hard. I j- like maybe I can imagine like 2035.

I don't even know 2050. But like, yeah, 2035, like yeah, that'd be really great. And like I have this like kind of, you know.

**Swyx** [1:08:48]
So in terms of merging, like isn't it-- Shouldn't you work on brain upload rather than AI, AI girlfriend?

**George Hotz** [1:08:54]
But I don't need brain upload, right? I, I don't need brain upload either. Like there's, there's thousands of hours of me on YouTube, right?

**Swyx** [1:09:00]
Yes.

**George Hotz** [1:09:00]
If you-- My-- How much of my brain's already uploaded?

**Swyx** [1:09:03]
That's only the stuff that you voice.

**George Hotz** [1:09:04]
Yeah, it's not that different. It's not that different, right? You really think a powerful-- You really think a, a model with, you know, an exaFLOP of compute couldn't extract everything that's really going on in my brain? I'm a pretty open person, right?

Like I'm not running a complex filter. Humans can't run that complex of a filter.

**Swyx** [1:09:20]
Yeah.

**George Hotz** [1:09:20]
Like humans just can't. Like this is actually a cool quirk of, of, of, uh, biology. It's like, well, humans like can't lie that well.

**Swyx** [1:09:27]
Yeah.

**Alessio** [1:09:27]
Yeah. So is it good or bad to put all of your stream of consciousness out there?

**George Hotz** [1:09:34]
I mean, I think it's good.

**Swyx** [1:09:35]
Yeah. I mean-

**George Hotz** [1:09:36]
I don't know. I'll-

**Swyx** [1:09:37]
... he's streaming every day.

**George Hotz** [1:09:38]
I wanna, I wanna live, I wanna live forever. Like that's-

**Swyx** [1:09:41]
Yeah. We, we said off, off mic that we, we may be the first immortals, right?

**George Hotz** [1:09:44]
Yeah. Yeah. Like this is how you, this is how you live forever. It's a question of, okay, how many weights do I have, right? Okay, let's say I have a trillion weights, all right? So talking about a terabyte, uh, 100 terabytes here.

I mean, but it's not really 100 terabytes, right? Because it's Kolmogorov complexity. How much redundancy is there in those weights? So like maximally compressed, how big is the weight file for my brain? Um, quantize it whatever you want.

Quantization is, is a poor man's compression. Um, I think we're only talking really here about like maybe a couple gigabytes, right? And then if you have like a couple gigabytes of true information of yourself up there, cool, man.

Like what does it mean for me to live forever? Like that's me.

**Alessio** [1:10:24]
Yeah. No, I think that's good. And I think like the-- there's a bit of like a professionalization of social media, where like a lot of people only have what's like PC out there, you know? And I feel like you're gonna get-- Going back to the ChatGPT thing, right?

**George Hotz** [1:10:37]
Wow.

**Alessio** [1:10:37]
You're gonna train a model on like everything that's public about a lot of people, and it's like-

**George Hotz** [1:10:41]
Then no one's gonna run their model and they're gonna die.

**Alessio** [1:10:46]
I know.

**George Hotz** [1:10:46]
Don't believe everything you see on social media. Your life could depend on it.

**Swyx** [1:10:51]
Um, we have a segment-- Uh, so, uh, we're, we're moving on to a, what, what would normally be called the lightning round, but just, uh, just general takes 'cause you're a generally interesting person with many other interests.

**George Hotz** [1:11:01]
Sure.

### Philosophical Takes

**Swyx** [1:11:01]
Um, uh, what does the goddess of everything else mean to you?

**George Hotz** [1:11:08]
Oh, it means that AI's not really gonna kill us.

**Swyx** [1:11:10]
Really?

**George Hotz** [1:11:11]
Of course.

**Swyx** [1:11:13]
Tell us more.

**George Hotz** [1:11:14]
Look, uh, L- Lex asked me this, like, "Is AI gonna kill us all?" And I was quick to say yes, but I don't actually really believe it. I think there's a decent chance that AI- I think there's a decent chance that AI kills 95% of us.

Okay.

**Alessio** [1:11:29]
But they saw on your Twitch streams that you're with them, so they're not gonna-

**George Hotz** [1:11:32]
No, I don't think... I actually, I don't also think it's AI. Like I think the AI alignment problem is so misstated. I think it's actually not a question of whether the computer is aligned with the company who owns the computer.

**Alessio** [1:11:42]
Mm-hmm.

**George Hotz** [1:11:42]
It's a question of whether that company's aligned with you or that government's aligned with you, and the answer is no, and that's how you end up dead. But, um, so what, what the goddess of everything else means to me is like the complexity will continue.

Paperclippers don't exist. You know, there are forces-- The paperclipper is cancer, right? The paperclipper is really just a perfect form of cancer, and the goddess of everything else says, "Uh, yeah, but cancer doesn't win," you know?

**Swyx** [1:12:07]
Yeah. It's a, it's a beautiful story for those who haven't heard it.

**George Hotz** [1:12:09]
Yeah.

**Swyx** [1:12:09]
Uh, and you, you read it out and I, I listened to it.

**George Hotz** [1:12:12]
It does. Yeah.

**Swyx** [1:12:12]
Um, yeah. Good.

**Alessio** [1:12:13]
Uh, what else we have here?

**Swyx** [1:12:14]
Pick a, pick a question. So many.

**Alessio** [1:12:16]
Yeah. What are you grateful for today?

**George Hotz** [1:12:20]
Uh, oh, man. I mean, it's all just like-- I haven't, I haven't thinking about this stuff forever. Like that it's actually like happening, and it's happening in an accessible way too. I guess that's what I'm really grateful for.

It's not like... Like AI is not some Manhattan Project style, you don't know anything about it-

**Swyx** [1:12:38]
Mm-hmm. Closed doors.

**George Hotz** [1:12:39]
Closed doors.

**Swyx** [1:12:39]
Yeah.

**George Hotz** [1:12:39]
And, you know, I'll fight really hard to keep it that way. Uh, you know, uh, that's, that's just-- I'm grateful for just, just how much is released out there and how much I can just learn and stay up to date.

And I guess I'm grateful to the true fabric of reality that, you know, I, I didn't need differential equations to understand it. Like I don't need-

**Swyx** [1:12:58]
Mm-hmm. Yeah, yeah

**George Hotz** [1:12:58]
... you don't need, you don't need some like- Like, like there's, there's... I've tried to do-- There's a limit to my, to my math abilities. I can do most undergrad math, but I took some grad math classes and okay, now we're getting to the end of what I can do.

**Swyx** [1:13:09]
Mm-hmm.

**George Hotz** [1:13:09]
And it's just the actual, like, end of what I can do. Like, I'm limited by my brain, but, you know,

ML stuff, like, you need high school math.

**Swyx** [1:13:18]
Yeah.

**George Hotz** [1:13:19]
Like, I could do all of... Uh, nothing like... Do you know what I mean? When I learned to multiply a matrix, seventh grade. Like-

**Swyx** [1:13:23]
Mm-hmm

**George Hotz** [1:13:23]
... it's, it's all easy.

**Swyx** [1:13:23]
You need more electrical engineering than you need high school math, really. Yeah.

**George Hotz** [1:13:27]
Y- yeah. Well, you need electrical engineering to, like, build the machines, but even that, like, these machines are simpler than the machines that have existed before. Like, the compute stack looks really nice. So, you know, yeah, I just, uh, I'm grateful that it's all happening and I get to understand it and be here, so.

**Swyx** [1:13:40]
Yeah. Yeah. Um, John Carmack mentioned there's about six insights we have left. Do you have an intuition for what some of the paths people should be taking? Obviously, you're working on one. Um, what are some of the other branches of the tree that people should go under?

**George Hotz** [1:13:55]
I don't think I'm working on one of the six insights. I don't think tinygrad's any one of the six insights. Um, something I, I really like that Elon does, and I try to take it from-- uh, try to, uh, be inspired by it, is, um

look at the boring tunnel machine and ask how you can build a 10X cheaper one, right? Look at the rocket. How can I build a 10X cheaper one, right? Look at the electric car and say, "How can I build a 10X cheaper-"

**Swyx** [1:14:17]
Mm-hmm

**George Hotz** [1:14:17]
... like, or cheaper or, you know, can go further or whatever, whatever, whatever, right? And you just do the straight up physics math, right? Like, I'm trying to do the same thing with, with, uh, ML frameworks. Right? And in, and in, in doing so, making sure that this stuff remains accessible, right?

You could imagine a world where if Google TPUs were actually the ultimate-

**Swyx** [1:14:36]
Mm-hmm

**George Hotz** [1:14:36]
... if Google TPUs were actually the best training things. I mean, actually, you know, I'm kinda grateful for Nvidia, right? Like, because if Google TPUs were the ultimate, now you have this huge closed source compiler in between, uh-

**Swyx** [1:14:45]
Yeah

**George Hotz** [1:14:45]
... XLA-

**Swyx** [1:14:46]
Yeah

**George Hotz** [1:14:46]
... and, and the hardware and yeah, that's, uh, just a really bad thing. So, I mean, something that is somewhat upsetting about the tinycorp is it, is it, is that it is trying to prevent downside. But, uh, it's not all trying to prevent downside.

Like, we're also building computers-

**Swyx** [1:14:59]
Mm-hmm

**George Hotz** [1:14:59]
... and we're gonna build some awesome, powerful, cheap computers, uh, along the way. Uh, so no, I'm not really working directly on any of the six tricks. I also think the six tricks are kinda gonna be like luck.

**Swyx** [1:15:09]
Yeah.

**George Hotz** [1:15:09]
I think it's just gonna be like, you know, please tell me more about what covariate shift is and how that inspired you to come up with batch normalization. Please tell me more about why it's a transformer and it has a query, a key, and a value, right?

Like, Schmidt Huber described it better in Fast Weights.

**Swyx** [1:15:22]
Mm-hmm.

**George Hotz** [1:15:23]
You know?

**Swyx** [1:15:23]
Yeah.

**George Hotz** [1:15:23]
Like, like, I mean, uh, my theory about why transformers work have nothing to do with this attention mechanism and just the fact that, like, it's semi weight sharing, right? Like, the... Because the weight matrix is being generated on the fly, you can, you can, like, compress the weight matrix, right?

Like, this is what that... There's a, there's an operation in the, in the transformer which, uh, like... And by the way, this is like Qualcomm's SNPE can't run transformers for this reason.

**Swyx** [1:15:46]
Mm-hmm.

**George Hotz** [1:15:46]
So most matrix multipliers in neural networks are weight times values.

**Swyx** [1:15:50]
Yeah. Mm-hmm.

**George Hotz** [1:15:50]
Right? Whereas, um, you know, when you get to the, the, the, the outer product in, uh, in, uh, transformers, well, it's weight times weight. It's, uh, it's values times values.

**Swyx** [1:15:58]
Yeah.

**George Hotz** [1:15:58]
Right? Um, so SNPE, like, doesn't even support that operation, right? So it's like that operation that gives the transformer its power. It has nothing to do with the fact that it's attention, right?

**Swyx** [1:16:07]
Mm-hmm.

**George Hotz** [1:16:07]
And this just is a funny like... But that is one of the six tricks, right? Like, batch, like, these norms are a trick. Transformers are a trick. Okay. Six more.

**Swyx** [1:16:17]
I- is there a reason why... So you c- you talk, you talk about, uh, attention as weight compression. Um-

**George Hotz** [1:16:24]
I think compression is not exactly the right word. What I mean is that the weights can change dynamically based on the context.

**Swyx** [1:16:29]
Dynamic weights. Yeah.

**George Hotz** [1:16:30]
So there was this thing in PAC8, in the Hutter Prize, that I absolutely loved, and I've never seen it again in neural networks, and it's a really good trick. Okay. Imagine you have 256 weight sets for a layer, right?

And then you choose which of the weight sets you're loading in-

**Swyx** [1:16:42]
Mm-hmm

**George Hotz** [1:16:42]
... based on some context, and that context can come from another neural net, right? So I have another neural net which protect, projects, you know, a 256 wide, one hot, do a softmax, predict it, and then I actually load the weights in.

And I can do this operation at both test time and train time, right? I can do this operation at both training and inference, and I load in the weights given the context, right? Like, that is what transformers do.

But transformers, instead of having 256 discrete ones-

**Swyx** [1:17:06]
Yeah

**George Hotz** [1:17:06]
... it's actually just that, but continuous.

**Swyx** [1:17:07]
Yeah.

**George Hotz** [1:17:08]
Um, which is funny that that was in language models, and I just like... When I understood that about transformers, I'm like, "Oh, this is a real trick, and why are they using the word attention?"

**Swyx** [1:17:15]
Yeah. And today is actually the anniversary of, uh, "Attention Is All You Need."

**George Hotz** [1:17:19]
What? Oh.

**Swyx** [1:17:20]
Yeah.

**George Hotz** [1:17:20]
That's so cool.

**Swyx** [1:17:21]
Today, today, six years ago. Six years.

**George Hotz** [1:17:22]
Six years.

**Swyx** [1:17:24]
Changed the world.

**George Hotz** [1:17:24]
Wow. Well, there's one of your envelope tricks, right? And you could easily-

**Swyx** [1:17:27]
Yeah

**George Hotz** [1:17:27]
... write it on an envelope, you know? Think about how if you write out the... How many times have you written that? Because it's not in any libraries, 'cause it's, like, all used a little differently each time.

**Swyx** [1:17:35]
Yeah.

**George Hotz** [1:17:35]
Like, you just write out that exact same, you know...

**Swyx** [1:17:38]
Yeah.

**George Hotz** [1:17:39]
Yeah.

**Swyx** [1:17:39]
Uh, you've name-checked, uh, Elon a few times.

**George Hotz** [1:17:42]
Yeah.

**Swyx** [1:17:42]
Um, I think about both of you as systems thinkers. Input, output, thing, something in between.

**George Hotz** [1:17:48]
Sure.

**Swyx** [1:17:48]
Um, what i- what's different about your style versus his?

**George Hotz** [1:17:52]
Um, Elon's fundamental science for the world is physics, mine is information theory.

**Swyx** [1:17:57]
Huh. But you, you do a lot of physics as well. I mean, like, you, you base a lot-

**George Hotz** [1:18:00]
And Elon does, and Elon does a lot of information theory as well-

**Swyx** [1:18:02]
Yeah

**George Hotz** [1:18:02]
... too. But if the question is fundamentally... The, the difference maybe is expressed in what your ambitions are, right? Elon's ambitions may be, like, just-

**Swyx** [1:18:12]
Go to Mars

**George Hotz** [1:18:12]
... go to Mars, right? Go to Mars is the ultimate modern, modernist physics ambition.

**Swyx** [1:18:17]
Mm-hmm.

**George Hotz** [1:18:17]
Right? It's a physics problem getting to Mars, right? Well, what are electric cars? It's a physics problem, right? Okay, now he's, like, pushing on the autonomy stuff, and you push a little on information theory, but fundamentally, his dreams are physics-based dreams, right?

My dreams are information-based dreams. I wanna live forever in virtual reality with my AI girlfriend, right? Those are, those are the aspirations-

**Swyx** [1:18:36]
Mm-hmm

**George Hotz** [1:18:36]
... of someone who, who, who accepts information theory as a core science. So I think that's the main difference between me and him. He has physics-based aspirations, and I have information-based aspirations.

**Swyx** [1:18:44]
Hmm. Very, very neat. Uh, Mark Andreessen, uh, he is a... Uh, hi, Mark. He's a listener. Um, he is heavily... He's a big proponent of effective accelerationism. You've been a bit-

**George Hotz** [1:18:55]
Sure

**Swyx** [1:18:55]
... bit more critical. Why do you say that e/acc is not taken seriously by its adherents?

**George Hotz** [1:19:00]
Oh, um, well- Only the left takes ideology seriously.

**Swyx** [1:19:06]
Why is that?

**George Hotz** [1:19:07]
Well, just as a fact. It's just like, it's just like a fact, right?

**Swyx** [1:19:10]
Is the right more cynical? Is that, is that what it is?

**George Hotz** [1:19:12]
I don't know. It's like, it's like the left actually manages to get energy around the ideologies, right? Like, like, like, th-there's a lot more... Look, here you have, you have two effective altruists named Sam going in front of Congress, right?

Only one of them is in jail. Um, you know, it's, it's interesting. Uh, they're both calling for regulation in their respective spaces, right?

**Swyx** [1:19:30]
So SBF is definitely like kind of a ch- wolf in sheep's clothing kind of, right? Like he, he only adopted e/acc or EA, uh, to market.

**George Hotz** [1:19:37]
Oh, and Sam Altman is a, is a genuinely good guy who is not interested in power-seeking for himself. All right.

**Swyx** [1:19:42]
Okay. Okay.

**George Hotz** [1:19:42]
Um, we don't, we don't have to go there.

**Swyx** [1:19:44]
Fair enough. Fair enough.

**George Hotz** [1:19:45]
Um, but uh, no, e/acc is not like-- Like you are not serious, right? Um, you are not actually a, a, a serious, uh, ideology. You know, uh, Mark Andreessen, I like Mark Andreessen, but I, I think that like some of his Twitter things are like, dude, you like just dis- like, it's like, it's like someone who's in like 2019 whose like eyes were opened about like the political world being not exact.

You mean all the people on the news were lying to me? Yeah, bro, they were lying to you. Like, okay, we all figured this out five years ago. Now what are you gonna do about it? I'm gonna complain about it on Twitter.

Right. And that's what e/acc is.

**Alessio** [1:20:20]
Um, last and maybe most important, uh, why was Avatar 2 bad?

**George Hotz** [1:20:26]
Oh. Well, I have a whole... You can go on my blog. I rewrote the script of Avatar 2. I wrote a script that actually might make you feel something for the characters. I killed Jake Sully in the first scene, like you had to, right?

Do you really think his second story arc-

**Swyx** [1:20:40]
Yeah

**George Hotz** [1:20:40]
... topped his first one? No, of course not. You had to kill the guy and make the movie about the brothers, right? And just that alone and realizing that, right? Like you could've kept the Titanic scene. It would've been fine.

**Swyx** [1:20:49]
Yeah.

**George Hotz** [1:20:49]
I didn't even take it out. I left your Titanic scene, James Cameron. But I wrote you a story that... So you know, just, just, just...

**Swyx** [1:20:55]
He needs ships to sink in water.

**George Hotz** [1:20:57]
He needs-- Well, I-- Look, it's a, it's a great scene, but like the movie was just like, like the, the roman-- Uh, never mind.

**Swyx** [1:21:03]
Great CGI, you know, let down by the writing maybe. Yeah.

**George Hotz** [1:21:06]
Yeah, no, but like the CGI, like it was, it was a, it's a beautiful world, and that's why like I care so much, right? Like you don't hear me ranting about Pirates of the Caribbean 2-

**Swyx** [1:21:13]
Right

**George Hotz** [1:21:13]
... being a terrible story, 'cause come on, what do you expect, man? Like Johnny Depp's like, "Wow, I had a movie that made me rich? I love this." Like...

**Alessio** [1:21:21]
Yeah. But this goes back to like the midpoint, you know. I think you wrote like, feels like ChatGPT wrote the movie, and I-- That, that's my worry a little bit. It's like kinda converging towards that.

**George Hotz** [1:21:30]
Oh, I, I-

**Swyx** [1:21:30]
Malek, Malek wrote the movie.

Sorry, I didn't wanna interrupt you.

**George Hotz** [1:21:35]
No, no, I, I closed a f- I closed a pull request two days ago. I was like, "Was this written by ChatGPT?" And I just closed it.

**Swyx** [1:21:41]
Yeah.

**George Hotz** [1:21:41]
And like, you know what? I honestly feel bad if you were a human who wrote this. Like I, I-

**Swyx** [1:21:45]
You're incapable of be- of being more pathex. Uh, yeah.

**George Hotz** [1:21:48]
But, but, but, but if you-- If I have a classifier running in my head that asks-

**Swyx** [1:21:51]
Yeah

**George Hotz** [1:21:51]
... you know, is this a AI or is this a human? Like, you know. The, the only way to deal with all this like, like, like-- Oh, God, it's like the worst possible. Like, you know, people are like...

Like, like how are you mad about like these chatbots and you're not mad about like Tesla? Well, because if, if I don't wanna buy a Tesla, I don't have to buy a Tesla, and it won't really impact my life negatively.

But if I don't wanna use a chatbot, it's still gonna impact my life negatively. All the amount of like personalized spam that now makes me spend more cycles on my classifier- ... to tell if it's spam or not, because you can now use AIs and generate this so cheaply.

**Alessio** [1:22:25]
Oh, God.

**George Hotz** [1:22:25]
Like, no, I mean, we have to move to a model where everything's just a dollar, right? Like you wanna send me an email, it's a dollar. Like you guys wouldn't care. None of my friends would care. No one would care, except the spammers, right?

Like we just gotta move to those sort of models.

**Alessio** [1:22:36]
Yeah.

**George Hotz** [1:22:36]
Yeah.

**Alessio** [1:22:36]
Awesome. Um, one last message you want everyone to remember?

**George Hotz** [1:22:42]
Uh, look, go, uh, go try tinygrad. Uh, I hope that we're a serious competitor to, to what's out there. And then I wanna, you know, I wanna take it all the way. We'll start with just building something for GPUs, and then we'll start building chips, and then we'll start building fabs, and then we'll start building silicon mines, and then we'll have the first self-reproducing robot using...

Yeah, okay.

**Alessio** [1:23:06]
All right, George. Thank you so much for coming on. You did a big inspiration.

**Swyx** [1:23:08]
Thank you for having me. Thank you.

**Alessio** [1:23:09]
Thanks.

**Swyx** [1:23:10]
All right.

How was that?

**Alessio** [1:23:14]
Awesome.

**Swyx** [1:23:14]
We, uh, not, not quite Lex Fridman, but we hope to do something different than him.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
