LALatent SpaceDec 7, 2024· 43:54

[Paper Club] Weight Streaming on Wafer-Scale Clusters (w/ Sarah Chieng of Cerebras)

Sarah Chieng of Cerebras presents the weight streaming technique for training giant neural networks on Cerebras wafer-scale clusters, arguing it enables near-linear scaling by separating parameter storage from primary compute. The system uses the Wafer Scale Engine (WSE-2/3) with 900,000 cores and 44 GB on-chip SRAM, avoiding off-chip memory bottlenecks that limit NVIDIA GPUs. MemoryX provides external storage for weights and optimizer states (up to 2.4 petabytes), streaming weights to compute units via the SwarmX interconnect fabric, which aggregates gradients. This design allows training models with up to 120 trillion parameters without complex hybrid parallelism. Cerebras also leverages unstructured weight sparsity to prune 90% of data during transmission and skip zero-value computations on-chip, reducing bandwidth and improving efficiency.

  1. 0:00Cerebras Introduction
  2. 5:47Hardware Architecture
  3. 10:32Paper Overview
  4. 15:52Model Training
  5. 17:41Scaling Training
  6. 22:06Weight Streaming
  7. 24:49Components
  8. 34:58Weight Sparsity
  9. 38:44Data Layout
  10. 41:51Competitors
  11. 43:19Q&A

Powered by PodHood

Transcript

Cerebras Introduction0:00

Host0:00

Okay, there we go. Let's go.

Sarah Chieng0:04

Okay. Well, hi, everyone. I'm Sarah Chieng, and today I'm gonna be presenting the weight streaming paper by Cerebras. So I dropped the paper in the chat, but the paper is called Training Giant Neural Networks Using Weight Streaming on Cerebras Wafer-Scale Clusters.

And so this paper is a paper introduced by Cerebras about c- weight streaming, and-

Host 20:40

Did we lose her?

Host0:41

Uh, she seems to have just paused. Yeah. Oh my God, that's a really awkward, awkward time to pause. Uh, let me just see.

I'll text her. It's funny 'cause she's upstairs. Hey, you paused.

Oh my God, what an awkward time to come down. Um-

Host 21:06

Let's, let's do some of her background. Cerebras, chip company. They have big chips.

Sarah Chieng1:10

Is it working now?

Host1:11

Okay, you're back. You're back, you're back.

Sarah Chieng1:12

Okay. I'm super glad about that.

Eugene1:13

Big fat chips.

Host 21:17

Not potato chip, GPU chip.

Eugene1:20

Wafer chips this size.

Sarah Chieng1:22

Anyways, sorry. I don't know where it froze. Um, if I freeze again, let me know. Hopefully, it doesn't happen again. I switched Wi-Fi networks. Um, but basically this paper is called Training Giant Neural Networks Using Weight Streaming on Cerebras Wafer-Scale Clusters, and this paper covers weight streaming, which is a training technique pub, um, training execution flow by Cerebras that separates parameter storage from primary compute.

And by doing so, you're able to efficiently and scale the training of large language models. And so usually you can imagine, you know, your weights are stored directly on your compute units, but this paper talks about how you actually store externally, and you stream those weights to your compute units during training, um, to allow your cluster to scale very quickly and easily.

And so this paper was published by Cerebras, and Cerebras is a AI processor company, and their biggest thing is that they're building the world's largest and fastest AI processor. So their chip is called the Wafer Scale Engine 3.

So the Wafer Scale Engine 3 is the latest iteration of their chip, although this paper talks about the Wafer Scale Engine 2, um, which is just, you know, a previous iteration. So I'll talk a little bit about the differences as well.

And so for a little bit of context on, you know, what Cerebras is doing and what this hardware is doing, um, a few weeks ago Cerebras... So Cerebras, you know, you can do inference and you can do training, and, you know, as Swix mentioned in the chat, a few weeks ago Cerebras came out that the Wafer Scale Engine 3 can serve Llama 70B at 2.1000, um, sorry, 200- 2,100 tokens per second, and serves Llama 405B at nearly 1,000 tokens per second.

So this, you know, to give you an understanding, like this is about 70 times faster than anything you can do with an NVIDIA GPU, or like running inference on an NVIDIA GPU. Um, and so they actually have a way to kind of experience this very easily.

So let me... Now I'll screen share. Um,

okay.

Let's see. Um, let me screen share.

Host3:59

Eugene, what is typical ser- um, inference for you guys?

Sarah Chieng4:04

It's gonna kick me out, so I'll be right back.

Host4:07

Okay. Oh, boy.

Eugene4:08

Um, so typical GPU inference typically is within the hundreds, uh, of tokens per second.

Host4:15

For 70B.

Eugene4:15

You, yeah, for, uh, and-

Host4:17

You always have to say which, which size you're talking about.

Eugene4:19

For which size, yes. Uh, for s- for 70B scale. If you want to do a tensor parallel throw four H100s at it, then you can, or eight H100, you could potentially start approaching 1000, but that's like tr- burning a lot of money into just one batch size.

Host4:37

Yeah. And you use all the speculative decoding and all that. Yeah.

Eugene4:41

Yep. It's, uh, I- as far as I know, no one is doing that other than in some specialized benchmarks because it's just too expensive to burn an entire H100 node just for one request.

Host4:52

Yeah. Uh, Sarah, I think you're back, but you're not speaking. Um-

Sarah Chieng5:01

Can everyone see the screen?

Host5:02

Yeah. All right, you're back.

Sarah Chieng5:03

Okay, awesome. So, you know, this is, you know, a public facing chat that Cerebras has. Um, it runs Llama 70B, um, and, you know, it's like a ChatGPT. It runs Llama 70B, and then it's powered by, you know, the Cerebras hardware.

So you can ask a question like, you know, "Tell me about World War II in great detail."

And you... 2,500 tokens per second. So it's, you know, generates all of that in the blink of an eye. So, you know, that's just a little demo, and I'll put the link to this in the chat as well if anyone wants to, like, play around with it.

Um-

Hardware Architecture5:47

Host 35:51

I got you.

Sarah Chieng5:53

Okay. Thank you. Okay, and then, yeah, so that is, you know, Cerebras hardware. So anyways, going back to the paper, though, um, so the paper's talking about how you combine each of these different, you know... So Cerebras' main hardware is the Wafer Scale Engine 3, and the Wafer Scale Engine 3 sits inside the CSX, which is the entire system.

And this paper, and that is one compute unit, and this paper talks about how do you combine multiple of these compute units to, you know, at, work together and achieve that near linear scaling. And so before we jump into the paper, I thought it would be some useful context to talk a little bit about the Wafer Scale Engine that's, like, kind of the heart of this system, and really br- and break down a little bit about the architectural design that is allowing Cerebras hardware to, you know, serve inference so quickly, training so quickly, et cetera.

You know? So I have a couple of slides for this introduction portion. Um, let me know if people can't see, if it's not working how I think it's working, but hopefully everyone can see the slide right now. Um, so this on the left is, you know, a picture of the Cerebras hardware, the Wafer Scale Engine 3.

It is pretty large, as you can see. This is, you know, the heart of, you know, this paper, this heart of, you know, what Cerebras is doing. It has about four trillion transistors, 900,000 cores, and 44 gigabytes of on-chip memory.

And so for context and, you know, as a side note, the paper does talk about the Wafer Scale Engine 2, which is a previous iteration, has 850,000 cores, 40 gigabytes of on-chip memory, just, you know, a little less powerful.

And so, you know, I'm sure everyone's wondering, everyone's much more familiar with the NVIDIA H100, and so now to kind of break down, you know, to understand why Cerebras is able to perform at the speed it does, I want to compare it to the NVIDIA H100, and, you know, one of the biggest bottlenecks and limitations that the NVIDIA H100 has.

And so I'm going to focus a little bit on, you know, the NVIDIA GPU. So in this picture, on the left you have the Cerebras hardware, on the right you have the NVIDIA H100. You can tell significantly less transistors, 80 billion versus four, versus 4 trillion.

And, you know, this size is pretty to scale as well in terms of, you know, that is how small the H100 is compared to the Wafer Scale Engine 3. And so to kind of, as I mentioned, understand, you know, why, expl- to understand the architectural decisions behind the Wafer Scale Engine 3, um, I want to talk about the H100 and one of its biggest limitations.

And so if you look at this picture of the H100, each of these, you know, GPUs comp- um, is made up of cores. And if you look at the red rec- red rectangle, that is one core. And these cores is what is actually making up, is actu- is what's actually doing all of the mathematical computations needed for, you know, inference and training.

And on the NVIDIA H100, there's about 17,000 cores on that H100. And so each of these cores, you know, it needs access to things like the weights to actually do all of the computations. And on the H100, all of these weights is stored off-chip.

So there's an, uh, memory is located off-chip, and this is what's storing the weights, you know, during inference like KB cache, all those values. And so now you can imagine you have these 17,000 cores, and each of them are constantly needing to load and unload values from this memor- from this off-chip memory, and now you have this very, very significant memory channel bottleneck.

So now in comparison to the Wafer Scale Engine 3. So what Cerebras has done for the Wafer Scale Engine 3 is that instead of storing all these weights and values, weights, um, and values off-chip, Cerebras stores everything on-chip in SRAM.

So every single one of the cores on the Wafer Scale Engine 3 has its own SRAM. So each core has direct access to each of the values that it needs to do its computations. And this design eliminates the need to repeatedly fetch weights from a central memory, which then significantly reduces the memory bandwidth requirements and power consumption of the Wafer Scale Engine.

Paper Overview10:32

Sarah Chieng10:32

And so that is one of, one example of a very key architectural difference between, you know, what's out there, like the NVIDIA H100, versus the Wafer Scale Engine 3 that's really, you know, now that these weights are not having to, these weights and values are not having to travel as far, you're not having these memory bottleneck constraints, it's able to power things like inference much faster.

So now that is kind of a little bit about the hardware. And so now kind of shifting gears and talking about this paper. So this is one system, and this paper is focusing on how do you combine multiple of these c- of these wafers, these wafers sitting inside of the CSX system, how do you combine multiple unit, of these units together to train very large models very fast and very efficiently and very easily?

And so this system that Cerebras introduces actually has three different components. It has... And, okay, so, sorry, to take a step back. So how I'm going to structure this part of actually going through the paper is I actually have some notes.

Um, I have a Notion doc of all the notes that I t- um, wrote up, and I'll present the notes next to the paper. But before we deep dive into the paper itself, I want to, um, you know, give a, a really quick overview of what this system is looking for, um- Or, sorry, what this system looks like.

So the paper presents this Weight Streaming solution that has a couple of different components. The first is a CS-3 CSX system. Those are the compute units. And instead of storing all of the weights that the compute units need on, on the compute unit, it's storing it externally in a external memory service.

In this case, it's called MemoryX. And during training, these weights are streamed from MemoryX to the compute units, once during the forward pass, once during the backwards pass. And then the MemoryX is also what is computing the updated weights.

And then this middle layer, this SwarmX that you see, is basically that interconnect fabric that is connecting each of the compute units to the mem- external memory service. So that's kind of a high level of what the different components in Weight Streaming are.

And so now... Okay, so, and so this also, you know, better illustrates that a little bit. You have your memory service, um, all of your memory, all of your... sorry, all of your weights. These weights are streamed through the SwarmX, which is your interconnect fabric, to each of the different systems.

And so now to k- let's, now let's go through the paper. And so what I actually want to do, so here are some notes. It's a pretty long paper. It's around 35 pages of technical content. Um, so to kind of make it easier, I've created these kind of like condensed notes, um, that pretty similarly follows the paper section by section.

Um, but you know, it's in notes form, it's a lot easier to digest. And then I also added a couple of, you know, like, side notes, you know, like what are other people doing? What is the updated version of this since this paper was published a couple of years ago?

Um, and on that note, this paper was published a couple of years ago, I think in 2021, but the techniques in this paper are still very relevant. They're still used in production today, um, still very applicable. So what I wanna do is, actually, let me

put it in the chat.

Let me stop the share really quick. I'm gonna put these notes in the chat. Thank you, Kyle. Yes. Um, let me know if this does not let you access it. Um,

okay. So e- everyone...

Let me know if people are... Okay, perfect. Thank you, Kyle, for confirming that you can access the notes. Um, and so, yeah, so basically I'm gonna be sharing the notes. Um, sorry. Let me re-share my screen.

Okay, cool. So on the left we have our paper, and on the right we have these notes. We're kind of... I'm gonna kind of be going through them together, um, a- and then obviously everyone has access to the paper and the notes as well, so you can look at it at your own pace as well.

And one thing that I do want to say about these notes is that if you have any questions or comments or clarifications or things that you are like, "This is unclear," or things that you think are missing, just please add a comment on the notes.

Um, and if there are ever any questions during this, the ones... and, like, maybe I'm not able to answer it, like, just leave it as a comment and I promise I will get to it, and I will answer it in the notes and update the notes.

Um, but yeah, I mean, bear with me if there are things that I am not able to answer. Um, but yeah, so now let's go through the paper. Okay, so it's just an abstract, and then... So yeah, so the first thing that this paper talks about is, you know, just an overview of model training.

Model Training15:52

Sarah Chieng15:52

I don't know the context of, or, like, the context of what people have, but we'll just go through it anyways. It's pretty, um, straightforward. So this first section is just, you know, how do you train a large neural network model?

And what that involves is you have all of your training data, um, you process it in batches, and for every single batch you do forward propagation, you compute the loss, backwards propagation, and you update the model weights. So for forward propagation, your model is, you know, composed of different layers.

Each of the layers you do different computations. You... for each layer you take the, um, ac- the, like, computation output and pass it to the next layer. And so in the first step, forward propagation, you're passing your training data through each model layer sequentially, and at each layer computing the activations, and that's just the, you know, the new computational output of, you know, that layer.

And then af- so you go through every single layer, you have the model's final prediction. And then you then compare your model's prediction to the actual value, the true label, um, from your training data, and use that to measure the prediction error.

And now, once again, you go through your model layer by layer during back propagation, but now you're going from the last layer to the first layer to calculate the loss gradient with respect to the output activations. And then from here, the main thing is that you're computing the activation gradients, so determining how changes in activations influence the loss layer by layer.

So now that you have the activation gradient, in the final step you use the activation gradient to compute the weight gradient, and this is how you're actually updating each of the weights during each training iteration. And so now you update your weights, and you continuously do this until you have an error that you are happy with.

And so that is ki- kind of like, you know, the concepts in NLP model training that the first section is talking about. So the next section that they're talking about is scaling training with stored weights. And so this is, you know, talking about different types of parallelism that you can employ to, you know- Scaling training involves distributing computations across multiple compute units.

Scaling Training17:41

Sarah Chieng18:06

And so there's two main types of scaling training. We have data parallelism and model parallelism. Oh, I have...

Sorry, wait, let me read the comments. Okay, the comments are chill. Or the chat, sorry. I just wanted to read the chat in case I was missing anything. Um, so there's two main methods of scaling training. The first is data parallelism, and the second is model parallelism.

And so... The weight updates are not on the compute units. That is correct, Eugene. The weight updates are not done on the compute units, but the memory component, MemoryX, is actually, is what's actually storing the optimizer. Um, so this is jumping ahead a little bit, but that is what's storing the optimizer.

And so all the weight gradients are streamed from the compute units to MemoryX, and then MemoryX is what's actually, you know, has the optimizer computing the updated weights and then streaming the updated weights back. So you're not doing any of that on the compute units itself.

Um, but anyways, back to data parallelism and model parallelism. So there's data parallelism, where you're splitting the batch of training samples over compute units, and then model parallelism, you're instead, instead of splitting up the batch and the training data, you're splitting up the actual model.

And there's two types of model parallelism that the paper talks about. We have pipeline model parallelism and tensor model parallelism. And so here's a diagram, figure one, pulled straight from the paper that kind of illustrates this concept. And now going into data parallelism.

So in data parallelism, you're storing a copy of the model on every single compute unit and sharding the training batch across the compute units. So each compute unit is calculating a partial weight gradient that's then summed up to give you that final weight gradient that's used to update all of the weights.

The main issue with data parallelism is that every single compute unit must store the entire model, and this is a very large limitation for large models 'cause this takes up a lot of memory. And so there's a variation of data parallelism called fully sharded data, data parallelism...

Here, let me scroll to it in the paper right here. Um, that kind of is meant to address this. And so in fully sharded data parallelism, model weights are distributed across compute units so that only one copy of each weight is persistently stored across the entire set of compute units.

So if you look at the steps here, each compute unit is responsible for a portion of the model. And during the forward and backward computations, the u- compute unit that o- owns that, you know, subsect- subs- subset of weights broadcasts those weights to all of the other compute units.

So, but then once those computations are done, each compute unit, you know, erases all the other com- other weights and only, like, persistently stores the weights that it's responsible for. And so that is, you know, allows... So now the model just has to fit in the aggregate memory of the compute cluster.

And then with model parallelism, there's two types of model parallelism, tensor model parallelism, pipeline model parallelism. Tensor model parallelism, you're splitting each layer across compute units. Pipeline model parallelism, you're distributing subsets of the models layer across compute units.

But the biggest thing to know is, you know, there's a lot of communication overhead with model parallelism. You have to share activation tensors. And that is why in this paper and, you know, in production, Cerebras is fo- Cerebras focus on data parallelism.

So all of this is mentioned in the paper as well. But model parallelism demands significant interserver commu- communication bandwidth, um, which, you know, as you increase cluster size just gets increasingly complicated. And there's, now you have communication bottlenecks, and it just overall you now have a very, very complicated system and implementation.

Weight Streaming22:06

Sarah Chieng22:06

Um, so now, now that we have that, um, background out of the way, we can kind of jump into weight streaming. And so weight streaming is where you separate model weight storage from compute. And so as I mentioned, what this means is that you're not storing your weights directly on the compute units, but instead on an external memory service called the MemoryX, which is an interconnect, um, MemoryX, which is the external memory service.

And so, you know, this is one thing that the paper mentions. Um, so let me scroll to the paper to match where we are. So here. This is actually a new solution is needed is kind of what this diagram is covering.

So you know you don't want to com- impose constraints on model size based on the memory available on the individual compute unit. So you know, that's... You want to be able to scale throughput, um, compute throughput with the computational requirements of the model, and you want to achieve scaling without complicated hybrid approaches to parallelism.

So that's kind of this column right here. And then here we're kind of introducing this weight streaming solution. So here. So instead of CPU/GPU processing, we're looking at the Cerebras CSX system, which as I mentioned, is what's, you know, the system that has that Wafer-Scale Engine inside.

And you know, we have significantly more on-chip memory, significantly more cores and transistors. So we're able to reduce the number of compute units to achieve an acceptable compute speed. Throughput, we're disaggre- aggregating compute from model storage, as I mentioned, that external memory service, MemoryX.

And then we also have a new me- uh, model for scheduling and coordinating the training work across the compute units. And so that's kind of... And so, you know, memory capacity, we have to hold... You know, when we're training models, we have to have the weights, the gradients, the optimizer.

And then Interconnect, you know, we're streaming the weights from the memory service to compute units, and then after those compute units, you do forward pass, backward pass. We're then transmitting the weight gradients back t- from the compute unit to the memory service.

I think as Eugene mentioned in the chat, it is the memory service that is actually, you know, that stores the optimizer and is actually, uh, taking those weight gradients and updating the weights, and then then streaming the updated weight ba- weights back.

So none of that is happening on the actual compute units itself. And so now to break down, as I've mentioned several times at this point, there are three main components of this Weight Streaming implementation. The first, as I mentioned, is that Wafer Scale Engine, the AI processor inside of the system.

Components24:49

Sarah Chieng24:49

The second is the MemoryX, which is the external memory service that stores model parameters and then, as we've mentioned, the optimizer states. And then the last is SwarmX. So this is the interconnect fabric that is transmitting, you know, the weights and the weight gradients between MemoryX, the memory service, and our actual compute units.

And this kind of abstraction and also having SwarmX as the middleman actually allows, you know... It, it black boxes a lot of things in terms of what the MemoryX sees and what each of the compute units sees, and we'll talk about it more, but it's, it's really cool, um, when we get to the SwarmX section.

And so as I mentioned, this was a little bit covered in kind of my introduction in the very beginning. But going back, you know, we have our AI processor, the Wafer Scale Engine, much larger than anything we've seen with the NVIDIA GPU, and it has three layers.

And so on this wafer itself, we have the processing layer, which is all of our cores. We have our on-chip memory layer, which is, you know, all that SRAM that's giving each of our cores direct access to the weights and values, you know, that it would need.

And then we have that interconnect fabric. So similar... So we have, you know, two fabrics. We have one fabric, which is the SwarmX between MemoryX and the compute units, and then we have an on-chip interconnect fabric that is linking all of these cores together for very efficient communication.

And so the Wafer Scale Engine 3, as I mentioned, 900,000 cores, 44 GB of SRAM, 4 trillion transistors. And I do want to note here that the fo- paper focuses on Wafer Scale Engine 2, and so the Wafer Scale Engine 3 is, you know, just an upgraded version of that.

And so now looking at the MemoryX service. So this is the external memory service that is providing persistent storage for the model parameters and optimizer states. And just so in this entire system, you have one MemoryX and, you know, it scales from 4 terabytes to 2.4 petabytes, supports models with up to 120 trillion parameters, um, and then it utilizes DRAM and flash storage.

But MemoryX, you know, it, it can ha- it's very, very large memory. And so in this system you have one MemoryX, you have that warrent one, like SwarmX interconnect layer, and then you can have all of your CSX systems.

So, you know, you can scale the number of CSX systems up, and you just need that one MemoryX memory service. Um, so you can kind of see that ratio here. You know, one MemoryX, one SwarmX, and then all of your compute units.

And so the MemoryX has two primary functions. One, it streams the weights to the compute units. So it streams the weights twice. It streams the weights once when, during the forward pass, um, to calculate the activations, and then it streams the weights to the compute units a second time to compute the activation gradients.

Um,

sorry, I'm catching up on the chat really quickly.

MemoryX also uses wafer chip to update values. Uses wafer chip to update values. I don't know exactly what you're trying to ask, but the MemoryX is using the weight gradients from the wafer chip to update the weight value.

So the wafer chip is having to stream the weight gradients. The wafer chip inside of, um, the CSX, which is our compute unit, does calculate, you know, the activation gradient, weight gradient, streams that weight gradient back to MemoryX to update the weight values.

So the... But then the MemoryX is what's actually updating the weight values. I think that's a clarification there. Um, so okay. So going back to this, MemoryX, two primary functions. One, it streams weights to the compute units twice, and then two, it also calculates the updated weight.

So it supports various optimizers, as I mentioned, like SGD and Adam, and then uses that to calculate the updated weights. And one kind of side note, I don't know if it's really talked about too much in the paper, but optimizers actually require a lot more memory than the actual parameter.

So optimizers require about 20 bytes per parameter compared to the parameter that's around, like 2 bytes. And so even for a smaller model like Llama 8B, um, it will require more memory than what is just available on the 44 from, on the chip itself, like that 44 gigabytes.

Um, and then as I mentioned, you know, this, this MemoryX can scale from 4 terabytes to 2.4 petabytes. So very, very, lots of, um, lots of memory. And then the next component is the interconnect. The next and last component y- is the interconnect fabric called SwarmX.

And SwarmX is what's connecting MemoryX to the different CSX systems, and it's what's, you know, is passing along the weights and the weight gradients. And so during the forward pass, it's broadcasting the weights from MemoryX to compute units.

During the backwards pass, it's releasing, receiving- The partial gradients from each CSX system reduces the gradients into a single tensor and sends the aggregated gradient tensors back to MemoryX. And one thing that's really interesting, and I don't know if I...

Oh, I, I did add a diagram, so I'll talk about this together. Um, but nodes, right, and so s- quickly the specifications, nodes with 100 gigabyte per second network interfaces, supports broadcast reduce operations one to four, one to two.

What this means is that when you have... Each broadcast reduce will then, um, communicate to either two or four systems. And so if you have four compute units, you have this one broadcast reduce going to four compute units.

But if you have more than that, you can, you know, see this kind of like this tree, um, architecture where you have one broadcast reduce, you know, communicates to two nodes, and then each of the broadcast reduce, um, operations here are going to four nodes.

So it's kind of a way to branch it. And also I realize that I need to scroll up in the paper or the paper is very behind on the left.

Okay, here we go. And then finally, as we've... Can MemoryX be used for inference weight broadcast functionality?

So none of this is actually used for inference. So this MemoryX and SwarmX... Sorry, I'm answering the questions, um, in the chat really quick. The MemoryX and SwarmX, this whole Weight Streaming system is just used for, um, is just used for training.

So for inference, you're just using the SRAM, you know, 44 gigabytes on the chip, and then you can network multiple chips together to support larger models. Yeah, Kyle is exactly right. This is... This, everything we're talking about right now is just, Weight Streaming is just for training.

Um, okay, so how do these three components work together? Do-

Host 232:07

Sarah?

Sarah Chieng32:07

Yes.

Host 232:07

Um, there's another question, which is, what type of compute does MemoryX use? Is it like conventional CPU or something else? Would you know?

Sarah Chieng32:15

I am not sure. Let me see.

Host 232:19

I guess the question is, what kind of compute does it use to update the weights? Um, I guess it's, yeah, a question people have.

Sarah Chieng32:26

Yeah. Let me see the...

Let's see. In the paper. And everything's going so harder.

Mm, I'm not sure, but I can look at it afterwards.

Host 232:58

Yeah, we can move on and then-

Sarah Chieng33:00

Yeah

Host 233:00

... we can just address this in the Discord channel. Thank you. Thank you, Sarah.

Sarah Chieng33:04

But yeah, if you want to add that as a comment, I'll definitely address it or respond to your comment afterwards. If you just want to comment on the document, everyone has comment access. Um, okay, continuing on, or Daniel, if you want to jump in, if you know, Daniel is from Cerebras and is also here.

Um, but yeah, so these three components, as I mentioned, the forward pass, MemoryX is streaming weights to the systems, backward pass, the core stores the activation gradients in SRAM, in SRAM, which are used to compute the weight gradients.

These weight gradients are sent back to MemoryX for weight updates. And then SwarmX is just communicating between the weights and, um, the comm... Taking care of the communications between weights and gradients. And so one thing that's really, really cool that this cartoon kind of illustrates...

This cartoon I think is actually not in the paper, um, I got it externally. But what the SwarmX does is basically this MemoryX only has to send out one set of all of these weights, and then the SwarmX then takes care of, you know, distributing these weights to all of the compute units.

So from this MemoryX perspective, it's looking, it's almost as if it's only working with one compute unit. Like everything else is abstracted. And then from the compute unit perspective, it's sending all of its, you know, like partial weight gradients, whatever it's calculated, back to SwarmX.

And then SwarmX is aggregating everything so that it's only passing, you know, one set of weight gradients back to the MemoryX. And so, you know, like kind of with what you see here, it's like, okay, when there's one compute unit, one Wafer-Scale Engine 3, people are only seeing...

Like, all of this whole system from the user perspective, from the compute units, the memory units, everyone is almost like seeing one device, if that makes sense. Like the person is seeing one device. Um, MemoryX is seeing, you know, it's like sends one copy.

Weight Sparsity34:58

Sarah Chieng34:58

It doesn't have to worry about everything. And so there's a lot of like very beautiful abstractions that take place that really simplify, you know, the understanding and like how all of this works. And then one other thing that I wanted to talk about, so this is a little out of order from what the paper has done it, and the paper talks about weight sparsity actually as part of, um, the section on the Wafer-Scale Engine, but I wanted to take it out or move it after just so that we can just like, you know, talk through the three components of Weight Streaming and then, you know, have this little distraction about weight sparsity.

Um, but the paper does talk a lot about weight sparsity as well. And so weight sparsity refers to the reduction of parameters by pruning insignificant parameters. And so a lot of the times in these, all of these matrix computations, there's a lot of zero and near zero values.

And so in a dense representation, we're including all of these zero values. Um, and in a sparse representation, we're removing all these zero and near zero values, especially when they result in like z- computations equal to zero. Um- Um, and so Cerebras leverages weight sparsity in two main ways during training.

And so one is during that transmission by MemoryX through SwarmX, and once it's just, like, on the wafer itself. And so memory... And so as MemoryX streams weights through SwarmX, it eliminates zero and near zero values. And so this reduces bandwidth requirements significantly, pruning up to 90% of the data while, um, maintaining accuracy.

So the paper actually includes a lot of research on this as well, um, looking at GPT-3, creating a sparse representation of GPT-3 and showing that, you know, the performance is still and accuracy is still, like, not... is still the same.

And then also mentions, like, the lottery ticket hypothe- uh, I think it's called the lottery ticket hypothesis, um, just also showing that model parameters can be pruned by 90% without a reduction in accuracy. And so this w- so basically Cerebras is able to really take advantage of this by, one, eliminating these values between MemoryX and then the actual compute units, but then on the actual compute units on the wafer itself, each core in the wafer scale engine also it skips the computations for these zero and near zero values.

So the core can recognize which values to skip, which values it did not receive, um, and knows to skip these. And so, you know, it optimizes compute efficiency by only focusing on meaningful operations. And this next part is not, I don't think, really talked about in the paper, but I'm sure pe- but, you know, to just add a little more,

uh, context, it's like why can't GPUs fully utilize weight sparsity? So there's two types of sparsity. One is fine-grained sparsity, fine-grained or unstructured sparsity, where there are randomly distributed zero values, and there's structured sparsity, where there are groups of zeros.

For example, entire rows in a matrix. So both Cerebras and GPUs can handle structured sparsity, but GPUs are not designed to handle unstructured sparsity, whereas what I've just mentioned before is able to handle this unstructured spars- sparsity. So that's a fun...

That's, like, an additional thing that the paper talks a lot about. This very last section is the principles of operation, and so let me scroll down in the paper. And I believe this is the last section. And there's also a section kind of going over, um, you know, what all this looks like in, you know, some of the actual work that's been done, you know.

Data Layout38:44

Sarah Chieng39:08

Um, but, like, you know, of training models on the full COVID genome sequence, but I'm not... That's not included, um, in the notes. And so going back up here to the

principles of operation. So this is the last section that we're really gonna talk about. Um, so... Oops. Bad. Delete that. Um, okay. So the Cerebras Graph Compiler, um, the CGC right here integrates with machine learning frameworks such as, such as TensorFlow and PyTorch.

So this kind of, you know, like, what is all of... You know, all of this, like, what does it actually look like? And so we have our compiler, and then what does it actually look like, the data on the chip, everything like that.

And so this is what's compiling user models into binaries that can actually execute on this compute unit system, and then it auto- natively supports the weight streaming that we've been talking about in this paper. And then how is this data actually distributed across the chip once it's streamed onto the chip?

And so activation tensors in models, um, like GPT-3 have three dimensions, batch, sequence, and hidden feature. Um, and so the batch, you know, that's the input samples. Sequence is, um, or the process, the tokenized input, and then the hidden feature is when we embed the tokenized input, the sequence, to form a 2D matrix S by H.

And so this paper talks about, you know, two different ways that you can take all these activation tensors and actually lay it out onto the Cerebras wafer. Um, we have... This is diagram 15 and 16 inside of the paper.

And so we have S and H are sequence and hidden feature compiling each layer, and then B is for batch. And so depending on the relative dimensions of S and H, you can... You know, if S and H are very similar dimensions, you can lay it out on the wafer something like this on the right.

If S and H have very different dimensions, you know, um, you can lay it out something like this on the left. But both are valid configurations. And then the last little part is just talking about, you know, the matrix multiplications.

These four, uh, all the phases, you know, all the calculations is really just matrix multiplications. There's two types of matrix multiplication, one that supports sparse input and... one sparse input and one dense input, so this is when we're calculating the activation and activation gradient, and then one that supports two dense inputs and a sparse output, which is the weight gradients.

Um, so that's, that's it for the paper. And then there's this optional appendix section that Swix recommend I add. Um, our MemoryX also uses 10G RDMA. Okay. Reading the chat. Um, so I actually didn't dive too deep into these techniques.

Competitors41:51

Sarah Chieng42:12

Um, so I just kind of gave a, a couple of, uh, techniques and what different companies are doing, um, just so you're aware of it, and feel free to deep dive further. And so when we're looking at the different companies, we have NVIDIA Megatron, Microsoft DeepSpeed, and Facebook or, like, Meta Faisser.

And then we also, I think, I don't know if they've previously presented on Paper Club, Swix mentioned Strong Compute's data streaming. But Strong Compute's data streaming is pretty different. It's more about, you know, cluster management and, you know, taking and, like, fully utilizing your cluster and, like, cluster orchestration versus, you know, an actual, you know, implementation where you're disaggregating model storage.

Like, basically not... no one is doing anything close to where you're disaggregating model storage from compute, and none of these examples above, um, do that either. And so, yeah. That's, that's all I have prepared in these notes. Um, I've tried to keep track with the chat, but yeah, feel free to add comments, add things in the chat.

Q&A43:19

Sarah Chieng43:19

I think it's can MemoryX be used for inference?

Yes. I see some hands raised. I think you can just drop it in the chat if you have any-

Host43:33

Well, we can also, uh, just open it up to Q&A. But first of all, um, rounds of applause for great presentation, and also drop the notion. Um, I think we traditionally do Zoom applause, so let's just stream stuff in here.

Okay. Awesome. Uh, but yeah, Eugene, go ahead.

Eugene43:51

Yeah. I wanted to ask how, how much of this is