LALatent SpaceJul 29, 2024· 1:23:46

[LLM Paper Club] Llama 3.1 Paper: The Llama Family of Models

This episode examines Meta's Llama 3.1 paper, detailing the 405B dense model, its scaling laws grounded on the ARC reasoning benchmark rather than perplexity, and the decision to train on 15 trillion tokens. Vibhu explains the training infrastructure: 16,000 H100s over 54 days with 419 interruptions, 78% from GPU hardware failures. Eugene Yan walks through the synthetic data pipeline—using Llama 2 for filtering, stepwise reward models, and Monte Carlo tree search to improve reasoning traces. Hassan shares building LlamaTutor.com with Together API, serving 4,000 visitors and 5,900 requests for about $12. The group also discusses quantization trade-offs (larger models degrade less), inference provider variability (Groq's non-deterministic temperature zero), and compares Llama 405B's performance to GPT-4o and Claude.

  1. 0:00Intro
  2. 4:38Paper Overview
  3. 15:00Pipeline Parallelism
  4. 23:00Code & Benchmarks
  5. 30:41Synthetic Data
  6. 41:15Model Comparisons
  7. 53:22Building Apps
  8. 58:27Serving & Quantization
  9. 1:11:23Live Demos

Powered by PodHood

Transcript

Intro0:00

Vibhu0:01

And I kicked off like a what do people want to work on section with I'm gonna do a deep dive in the paper because I need to make slides on this, and that very much overtook the hackathon. We had like a solid crew of like 20, 30 people that were just discussing the paper with me.

And slides didn't really get made, but also, I don't know, it was weird. Hackathon project winner was just a deep dive into the paper. But, uh, we had a in-person paper club session that I led yesterday, and a lot of people from there are trying to join in, so it should be vibes.

Um, I- I'm liking in-person hybrid format. I might start running those. We'll, we'll see how they go, but it was good. Everyone had good discussions.

Guest0:40

Amazing. Amazing. Um, yeah, I would be happy to join that once I get back to SF this Friday. Um-

Vibhu0:47

Ooh, Friday. Exciting.

Guest0:49

Yeah. So, uh, as, as you know, um, we also -- we interviewed Thomas, who was one of the paper co-authors. Uh, he did not give us the paper beforehand, which is annoying because after reading the paper, I have so much better questions than we actually end up asking in the podcast, but whatever.

Um, yeah, yeah, I think, you know, a bunch of us have read it. Um, I, I feel like, Vibhu, you're probably best situated to, um, to take over the screen if you want, uh, if you have, if you have stuff.

Vibhu1:16

Sure.

Guest1:16

Um-

Vibhu1:17

I have very basic stuff, but yeah. Oh, oh, sure.

Guest1:20

Perhaps-

Vibhu1:20

We also got someone at the hackathon that worked on a hackathon project that's paper to video. So someone's cooking up a video explainer of this. It's like literally doing inference right now. We'll share it once it's ready.

Guest1:32

Yeah.

Vibhu1:32

But yeah.

Guest1:34

This is, I mean, this is, uh... I, I'm, yeah, I'm excited about it, but also I, I'm wondering how to do it justice. Uh, I feel like we can, uh, pose questions in here, and then, uh, you know, people can just kind of discuss in, in the, the Zoom chat.

Um, and, and yeah, I mean, like, I, we classically have a lot of side discussions in the Zoom chat anyway, so I'm, I'm not worried about that. Um, yeah, I mean, yeah. While, while Vibhu, uh, Vibhu, Vibhu, Vibhu go ahead, you can get started.

Um, but like, what do, what do people think? What do people want to talk about? Um, you know, per- personally, I, uh, I, I, I call this the synthetic data paper. Um, so I'm, I, I have a lot of like interesting insights or at, at least questions about the synthetic data stuff, but we can talk about everything.

Like there's, there's just so much in here.

Vibhu2:20

The format that worked well yesterday was like, we're not getting through 100 pages. I'll give the high level, we'll go through the tweet overviews, and then let's just dig into whatever anyone found interesting and had like, you know, something that someone dove into.

So like part of it was they had a bunch of scaling laws for pre-training. They had scaling laws for how they picked 405b and 15 trillion tokens. So whatever someone chose to dive deep into is what like I was like, "Okay, we'll, we'll dig into that."

Also, other part of this, I'm probably gonna give like a longer one-hour breakdown of like I'll go through the whole paper, have like an hour talk at some point. So I started slides. They're very not ready. These are like 20 minutes of slides, just became discussions.

But basically, um, we'll spend like two minutes on overview. Everyone knows Llama. Um, interesting stuff. The -- so like they dropped 3 to 3.1. 3.1 was a pretty big update. The 8B got a lot better. The 70B got a lot better.

A lot of this is just for other talk, but, um, yeah, they dropped some sizes. The context been getting bigger. We thought their scaling laws were just overtrain and pray, but, um, no, they're, they're actually pretty grounded in real, um, scaling laws.

Uh, they're all dense models. After reading the paper, their whole justification for this was like, "We wanna see stuff that scales." It's the first actual research paper where they talk about everything, hardware inference, hardware failures, what happened, how they fixed it.

Um, so real research paper. There's a lot on it. It's like basically everything pre-training, post-training, their scaling laws, how they ran experiments. It's a great recipe on how to build it. That's what this talk would be later when I do it.

Um, they cooked. Model's really good. It's solid open source. It's like GPT-4 level. They talk about how they bring up performance. Um, at some point we'll probably, you know, discuss performance. So-

Guest4:07

Right

Vibhu4:07

... everyone has their thoughts on benchmarks, if anyone wants to pop in for a sec, and then we'll just go straight to paper. Um, also for anyone that finds any time to cut us off, cut us off. It's all vibes.

Um, but yeah, we can see the jumps for the 8B, basically from 3 to 3.1. It got better all around. Um, overview of the paper. There's a bunch of Twitter threads, so instead of me making slides, we'll go over like the main one shared in Discord.

For everyone that hasn't seen... Also, is my whole screen sharing?

Guest4:38

Yes.

Vibhu4:38

Or is it just... Okay, let me share my desktop real quick. So for people that are new and not in Discord, um, we have a very active-

Paper Overview4:38

Guest4:47

Shush

Vibhu4:47

... running Llama 3 section. I'll, I'll share the paper. So, um, if we go to this little like -- You have to find it, so you gotta go through the news and find the Llama. There's like 60 links that we've been posting of everything popular on Twitter.

So we'll go through these at some point. But paper overview was basically like they have multiple phases of training. Um, I'm very not ready. I do have other notes on this that I'll share through. So I started screenshotting stuff.

Um, they have like three aspects to a good foundation model. Basically data, scale, complexity. This is why they didn't go into an MoE. They wanted training stability. Basic two-phase training. There's like pre-training, post-training. Um, they, they do a lot of scaling law work, so in their pre-trained data set, how they d- how do they determine what the pre-training mix is?

In the post-train-- i- in the pre-training, they start doing most of the training at like low context, then they continue pre-training at long context. Um, a lot of what they said in their complexity section was like, "We wanna do basic stuff that will like scale up."

This is like the foundation for how we can like redefine scaling laws and train this stuff up. So no crazy complex RL, just, you know, SFT for chat tuning and then- They do a lot of models for rejection sampling to do dataset stuff.

DPO, um, a lot of that is in the post-training. That's where they start to see capabilities added. They have little niche sections, like they have this, um, post-train on, like, dataset stuff, so like post-train on benchmarks, and that normally helps small models, but this was the first time someone got to do it at a 400B scale, right?

So, like they post-trained on stuff similar to GSM 8K, and the small models had a big improvement. The big one, it kind of didn't. So kind of talks about how like the scale benchmarks and like their held-out test sets, why maybe the big Geminis and stuff don't do it.

They added their safety bullshit at the end. Um, other interesting stuff that didn't make it to Twitter was like they had multimodal experiments. They trained in adapters. They have a vision adapter to, uh, audio adapter stuff. Um, there was cool sections on their pre-training mix.

So basically, they, they used a lot of like traditional filtering techniques. So they have like RoBERTa-based, um, filters for high quality. They have this for their synthetic data distribution, like how do we extract out high-quality data. Then they have a lot of traditional NLP for like PII, text extraction.

They had a whole section on like how they scrape the web and how they train their own parsing HTML. They compared it to what's out there, and their stuff's better. There's a lot in this paper. Uh, data mix was a really interesting section as well.

So they basically go into, um, here's basically what they did, deduplication, all this stuff that you would expect. Uh, model-based filtering was pretty cool. They used a lot of, like they trained classifiers on Llama 2 outputs. On the synthetic data side, Eugene has a great tweet thread.

We'll probably go through it at some point. Um, this was an interesting section that we haven't seen before. So when you have like a base model that you're pre-training and you have like fifteen trillion tokens, how do you determine what the right mix of that is?

So their finding was like half the tokens are general knowledge, twenty-five percent math and reasoning, seventeen percent code, all this stuff. But they're like, this is the first research paper that actually like breaks this stuff down. They actually did like scaling law experiments.

So they trained small models that were like a couple billion parameters. They started testing different data mixes, and then they trained a large model to see what actually works and what's the right data mix, and then they're like, "Here's the answer for this stuff."

Um, model architecture was pretty similar. They like did a few little changes. They, they did some better attention masking, group query attention, here's architecture. All this stuff is like on Twitter, so not as interesting. Um, from the podcast that Sean had, the vocab section is pretty interesting, though, like instead of messing with tokenizers, changing vocab is pretty big for small models.

Check out the podcast or if it comes up in discussion, we'll discuss it. Scaling laws was another interesting one for the paper itself. Um, basically traditional like Chinchilla scaling laws used to have this whole, like they're predicting what's the optimal for your compute budget, like what's the optimal model parameters, all that stuff, how many tokens you train on.

We thought that they were just scaling and praying and trading like, you know, fixed cost training run for cheaper inference, but this stuff is actually grounded. So they developed new scaling laws where TLDR of what they did is previously we used to tr- pr- we used to do scaling laws where we're just predicting on next token prediction accuracy, right?

So we're trying to predict on like perplexity and just how good is next token prediction. Instead, they do all this fancy math, and they change the training objective to be like more representative of a reasoning benchmark. They use the ARC challenge, where basically they have a reasoning benchmark, and now instead of doing scaling laws to predict next token prediction, they've changed it so that they're doing scaling laws to predict optimal model stuff based on actual reasoning, and that's where they come up with this, like their scaling laws show that for a 402B model, you wanna train on sixteen and a half trillion tokens.

Based on that, they did a flagship 405B based on fifteen trillion tokens. And then this is where they have their like infraoptimal, where they started to do the 8B, the 70B. They just reused their fifteen trillion tokens and just over-trained, and that works.

The other really cool section, the sections that didn't make it on Twitter were like their training infrastructure. So they give out everything, right? They give out like the full pre-training stack of like they have a section in here on how they do their pre-training.

So like, um, one is like the whole hardware configuration, so sixteen thousand H100 hours, what failures they hit, why they went for simplicity. This was a pretty interesting section. Like over their fifty-four day training, they had like four hundred job inst- uh, interruptions, four hundred nineteen unexpected interruptions, and like seventy-eight percent of these were like GPU hardware issues.

And then they have a section on like if they did MOE, all this stuff compounds, so we just wanted something like simple, scalable that we could deal with well. And like this is stuff that you don't really see in papers anymore, right?

It goes further with like what is the pre-training set. So like these formulas we don't really see anymore, right? So it's like when they pre-trained it, here's their like peak learning rate, here's their warm-up, here's their decay, here's how many training steps, here's the batch size.

Little nuggets like this haven't really like come up on Twitter yet, but like, you know, at first, they have a batch size of four million tokens with a small sequence length. So like the first bit of training is a sequence length of four thousand, then they double it to like eight million sequences at eight thousand for the next two hundred fifty-two million tokens.

After they've trained on two hundred million tokens, they double it again to like larger batch size for the next three trillion tokens, and then they do most of the training at eight thousand token sequence length. So like little stuff like this, I feel like we still need to digest.

There's, there's reasons for why they did this, but basically TLDR, no other open source paper has like a formula like this. And then that's kind of what the next like hundred pages is. I feel like at that point, instead of finding what I found interesting, like- I found all this stuff really interesting.

They talked about the batching, GPU utilization, memory, like utilization, all that stuff, like CUDA optimizations, their whole training recipe, um, what they released performance stuff. Instead, I feel like that's enough of a high level overview of the paper.

The more fun stuff is like, yeah, so how does it perform? Uh, they're all better. Infra companies are pretty cheap, and this is also where, like everyone else can hop into discussion. Um, Eugene, Sean, other Eugene, hop in now.

Um, you know, Fireworks is somehow really undercutting inference price. The scale leaderboard is a held out leaderboard. It does pretty good here. Um, what else? Groq has it. So some insider info for all the infra companies, they gave access to randomized weights that were the same size about six days before launch.

So six days ago, infra companies started playing around with it. They started working out how they're going to do inference, what type of decoding they need, but they didn't have the paper, they didn't have the actual weights, and then day of, they released weights.

But like, yeah, stuff like Groq is serving at 1,000 tokens per second. Um, what other discussions have we have here? Kyle did pretty good evals on, um, performance. He started doing it on his own fine-tuning stack. So he started fine-tuning it, um, compared it to 4o Mini.

OpenAI within hours responded with like 4o Mini fine-tuning, but fine-tuning the Llama 3.1 8B is kind of on par with 4o Mini. 4o Mini fine-tuning is kind of broken and free, but it gets worse. What other fun stuff?

Um, there's a comparison of model pricing here that's being updated live. Other tweets, George Hotz, Karpathy tweeted, um, vLLM supports it. Other more independent benchmarks coming in. Basically, it's good. The other interesting part was the licensing. So they changed up their like Llama license to proper full open source everything.

Um, we have more infra providers, NVIDIA stuff, but yeah, that's kind of where I feel like we should open it up. That's the quick 15, 10 to 15 minute overview. Whatever people found interesting, like I know there was a lot of discussion about synthetic data, Gen, Sean, and uh, Eugene, you had good tweets about this.

So I think this is where we open it up to whatever people found interesting, and then we dig into those topics because we're not getting through the rest of it. Um, I'm going to open up chat and see what people are up to.

But yeah, thoughts, everyone, hop in.

Guest 215:00

Yeah, I wanted to jump-

Pipeline Parallelism15:00

Vibhu15:01

The other jump racing, by the way. Yeah.

Guest 215:05

Every, if, uh, uh, one thing to warn about pricing is that, uh, you're going to see a lot of providers jumping in, and everyone's just trying to get their piece of the pie. So, so, uh, so like with some of the previous model launches, you see some pe- people coming in at lower and lower price, and then they'll increase later.

But I wanted to jump in on the training side because I'm quite sure, uh, Vibhu, Eugene Yan will have lots to say on the data. So, uh, I think I'll start with that. Um, I can't share the screen, by the way.

Uh.

Vibhu15:34

Do you wanna sh- do you wanna take over or you want me to scroll?

Guest 215:37

Uh, I want to take over slightly because-

Vibhu15:39

I can stop here.

Guest 215:39

Yeah.

Vibhu15:40

Yeah.

Guest 215:40

Uh, because I want to jump through a few mater- uh, a few things there. So let me share my screen.

All right. So I didn't see, uh, too much, uh, too much talk about on this, but, uh, but for me, right, one of the big ones is actually pipeline parallelism. Uh, uh, not ma- not sure how many people...

Can you see my screen?

Vibhu16:05

Yes.

Guest 216:07

Yeah. So, so, uh, so if you're looking at this and it's like, what is this crazy freaking schedule that they are doing here? But, uh, TLDR, uh, pipeline parallelism is the, generally the idea of like scaling up your training across multiple GPUs, let's say.

Uh, and, and to, and to build around optimizing that. That has its own benefits. Uh, it has also its own downsides. Uh, and the, the major downside that the reason why people try to avoid pipeline parallelism at all cost, uh, and they use like DeepSpeed3, for example, where the weights are shuttered around all the other GPUs, is, is that, is that if you look at pipeline parallelism or model parallelism, there's this problem called the bubble.

The bubble is basically as, as, as your dataset goes through the, the different devices, there, the, so the forward pass and then the backwards pass, uh, you have all this GPU time here where some of the GPUs are waiting for other GPUs and are doing nothing, and basically they are wasting compute.

And because, uh, because, uh, everyone want- uh, wanted to avoid wasting compute, the, uh, it went on to, to a, to a search of like, uh, the algorithm to figure out how to do pipeline parallel. And one major one is actually, uh, uh, SaleSG, coincidentally Singapore, where they created like this crazy ass algorithm, right, to basically train, train without any wasted time.

So you see the gray, the gray spots are with the wasted time respectively. And Facebook is now, now embarked on their own journey on this. And the reason why this is exciting even for, for smaller models is that this kind of algorithmic changes on the training, right, is what's going to allow you to train bigger models easier on lower end GPU.

So this concept could apply to, let's say, training a 70B model on 24 GBG, uh, GPUs and things like that. And the reason why they probably need it for, for the 80 GB is because they're training 405B. And yeah, and, and a lot of people thought, uh, like academia thought that this was a, uh, treated it as a dead end because of the bubble problem.

And then Facebook's like, "You know what? We are going to do that." And, and that to me is one of the more exciting thing. Uh, the, the other one that I saw some people tweeted out is about batch sizing being smaller, constraint or supporting the batch size.

Guest 318:29

Well, I thought Google has pipeline parallelism in their, uh, JAX, the distributed training repositories. They don't?

Guest 218:36

Yeah, they do. They do. Uh, but the thing is, no offense to Google, no one really took, uh, everyone just interpreted it as TPU has two, 2 liter VRAM.

Kind of, kind of, kind of thing. And they had the basic pipeline parallel button, which still suffered from the bubble problem. The, this weird scheduling, which I'm quite sure they are go- people are gonna start replicating it, is to reduce the bubble, the wastage, the wastage.

Guest 319:01

So, so I also saw lots of papers on this from, um, maybe NVIDIA and Matei Zaharia from Berkeley or Stanford. Like, they had lots of interleaved pipeline parallelism updates.

Guest 219:12

Correct.

Guest 319:13

So, so you're saying no one is using it, just Facebook has used, used it more recently? I, I find that pretty, um-

Guest 219:21

Or at least no one published it within, within their, their training processors, 'cause this is the first major model that, of this cell class size, right, that's saying, "Hey, we are doing pipeline parallelism." Pipeline parallelism has-

Guest 319:34

I mean, the Goo- Google models, so they have some, um, uh, these pathways, distributed training architecture systems, and they publish in, uh, maybe OSDI, which is kind of the biggest distributed systems conference. So they publish these trainings, and they can do all sorts of parallelism within their systems, and even mixture of experts parallelism and, and stuff like that.

So they do quite, quite heavy stuff. I'll look it up and p- post some papers if I find them, um, in the, in, in the messages. But yeah, th- this,

my mental model was that people are actually doing this at scale. Thanks.

Guest 220:11

Yeah, so I'll, I'll draw the distinction between pipeline parallelism and techniques like use DeepSpeed3, which is essentially where the GPU has, uh, uh, NVLink connectivity to other, other GPUs to actually read the, the model weights. Pipeline parallelism is really more of like instead of going cross GPU to read the, the weights of the other models or the other half of the model, you actually just focus on the half of the model that you, that you're working on.

And, and the t- and this has the, the, the trade-off respectively of saving VRAM and allowing you training larger model and larger batch size, but it means you have the bubble problem. And I, I think the focus is really more about the bubble problem here rather than w- rather than, than anything else.

And yeah, like I say, I, I, I do expect more people to replicate this part. Yeah. So, uh, that's the part that I wanted to jump in on. The other major one that I wanted to jump in on, uh, is just multilingual.

Uh, I'm so happy that I'm seeing this. We try to avoid using machine translated data to fine-tune the model. Um, and, um, this is something that I, I think multilingual people know that I've been shouting on the roof about saying, "Hey, can we stop using machine translated data for other languages?"

And then assuming that's great because when you speak to the other language, uh, natives, because they've been saying that sucks. And finally someone is also, uh, at least on the bigger, bigger model side, is doing that as well.

So particularly excited about that part. But yeah, uh, I think I'll hand off to the, the whole data stream.

Vibhu21:44

The, the interesting little section there of translated data is I've still seen it used where, like, they have a Llama 3 filter that extracts out what's the highest quality data, what's the highest quality reasoning code data and whatnot.

And in other work they'll still do ver- this is, like, very traditional pre-training data set stuff, right? Where you need more data augmentation to get more high quality data and translation. So, like, one thing is you can train on multiple rounds of that, right?

So, like, more epochs on high quality data, so you can just resample it. But then there was a paper that I'm forgetting that tested this. Do they wanna only use a little bit? Do they wanna train on multiple rounds of pass-throughs of the same high quality data?

Or they, do they wanna do basic augmentation, like translate and translate back, and somehow translation through other languages work better? Like, that was the best option. Translating it to high quality in another language as opposed to translate and translate it back.

So there's still, like, some value, but interesting little piece.

Guest 222:41

Yeah. So, uh, I think I want to hand off to the people who are gonna tear all the data parts into bits, 'cause I just wanted to jump in on training side because that's what I can uniquely offer.

Guest 322:54

Awesome. Appreciate that. I think Cameron has his hands up.

Guest 423:00

Hey, um, did they make any claims around it being good for code generation?

Code & Benchmarks23:00

Guest 223:06

Yeah.

Guest 423:06

I'm interested in whether... Yes, versus, versus Claude.

Guest 323:11

Yeah. Uh, they, they out- like, uh, it, this is a big contrast to Llama 2, where they were intentionally not training for code and then they put out CodeLlama separately. Uh, now they explicitly outline code as a separate modality, like separate from text.

Uh, Vibhu, I, I, I don't know if you have a slide on this, this stuff. Uh, and then they also did synthetic data for code as well. Um, yeah, they just, uh, they, they, they spend a lot more time on code this time around.

Guest 223:38

Basically both of those-

Guest 423:38

Has, has anyone looked at it versus Claude 3.5 Sonnet yet?

Guest 323:43

Uh, yeah, that's-

Vibhu23:44

We vibe checked it.

Guest 323:45

Yeah, they, they will, they will be doing benchmarks.

Vibhu23:48

We vibe checked it and-

Guest 423:48

They did what check? Vibe?

Vibhu23:50

Yeah. So, like, you know, it's not rigorous evals, but, like, we vibe checked it, and, like, it, it does pretty good. So in the paper they did explicitly mention as well, like, yeah, they used to have previous, um, they, they used to have previous Code Llama models, right?

And part of their, like, second step of post-training was to add in this section on code, but they explicitly no longer need to do that. And I'll, I'll pull up the section of the paper basically. But they, they mentioned that this is, like, natively trained in, in pre-training as well, but, um, it's, it's a good code model.

They also have a-

Guest 524:25

So, um-

Vibhu24:27

The, the base models-

Guest 524:27

There's a Scale AI-

Vibhu24:28

... have these models. Yeah.

Guest 324:31

Jeremy, can you repeat? Uh, we can't hear you very well.

Guest 524:34

Okay. Yeah. There's a Scale AI benchmark where, um, Sonnet and, uh, 4o were compared against the new 4o5B model, and, uh, 4o5B is found to be basically on par with GPT-4o, which is worse than both Sonnet and G- GPT-4 Turbo preview.

Um, uh- There's a tweet thread and a Reddit comment that I'll just drop. Um, but it outperforms Gemini 1.5. Uh, the thing I like about the scale benchmarks is that they are holdout. That is, like none of the companies have access to them, and they're private.

So there's probably more durability to the benchmarks, and they don't have as much of a conflict of interest, though they did co-watch with Llama, um, so, um, yeah, there may be a little bit of conflict of interest.

Guest 425:22

Thank you. Thank you, Jeremy.

Vibhu25:25

Yeah. Uh, um, so o-

Guest 525:26

I think there's-

Vibhu25:27

Overview, their scale-

Guest 525:29

Go ahead

Vibhu25:30

... scale leaderboards aren't just coding. So for people that don't know, it started out with, uh, the GSM 8K, where they tried to recreate it, and they made a GSM 1K, which is meant to match the actual benchmark, and just be a heldout that they'll run models, they'll evaluate them.

And then that turned into now they have heldout benchmarks that no one can see what the actual examples are of coding, instruction following, math, Spanish. There's a bunch of these, and yeah, they're, they're kind of like pretty good in the sense of like no one can directly train on them.

There was a piece that said, like when they put out their first one, what's the delta between companies, like models, that do really well on traditional like GSM 8K, but don't do well on 1K, where it's like they haven't seen it before.

So they basically tried to test who overfit to the benchmarks, and this is trying to solve that. So if we go through it real quick, this is kind of where the 405b sits in coding. It's like a step right below, um, GPT-4s and Sonnet.

Sonnet's still slightly better. And then we can kind of go, go through it. I think they're still testing the 405b, 'cause I'm not seeing it through the rest of them. But, um, they're being, they're being updated in tweet threads and whatnot.

And then Jeremy shared a link to the Reddit that talks about this, where they're, they're basically going through them, and then there's discussion here if anyone's interested. But yeah, someone was also talking then.

Guest 426:59

Thanks very much.

Guest 327:01

Hey, yeah. Well, I have something to share. The coding evaluation, um, it seems like they... So there's... I can share my, s- my screen. Um, but-

Vibhu27:14

Right

Guest 327:14

... uh, basically HumanEval. Um, HumanEval is, uh, kind of... One second. Let me try and share it. Um, can you see? Yeah. So HumanEval, um, is one of the benchmark data sets that people use to, to benchmark coding, and it's very simple.

Like, they have 150 questions, and it's almost like auto-complete. Like, solve this simple puzzle in Python or things like that. It's very, like one, two lines. And, um, you can see that... Let's see. So the Llama 405b is not state-of-the-art, so Cloud Sonnet beats it by a few percentage points.

Uh, it's close to the GPT and the Sonnet models but, uh, slightly worse. Uh, and I think this kind of conf- is similar to the Vipe checks. My understanding was on the initial Llama stuff that, um, Meta didn't focus that much on reasoning or on code, 'cause they're a social company, so maybe reasoning is not as super important, uh, for them.

But then they hacked, um, focused coding data collection session and, uh, send, uh, shared the big code model, which kind of wasn't that great. Maybe if you don't put the data in from the beginning, um, just trying to fine-tune on code, um, by itself, um, uh, doesn't work that well.

The other thing I wanted to share, can you see this other page, uh, now? Um, basically it seems they spend quite a bit to, to make their coding much better in, in Llama 3. Um, and, uh, they actually trained the code expert and then, uh, try to use that code expert to maybe, um,

I guess collect high-quality human annotations and, and, uh, do some more, more post-training. And then they also did some synthetic data generation, um, um, to, to improve coding. So I think they spent quite a bit to, to work on reasoning and coding.

I didn't read this section carefully, but yeah, they have a full section on, on, uh, trying to get better code data to generate, to, uh, incorporate feedback, um, and do analysis. Like they, they did quite a bit on coding.

Um, yeah.

Vibhu29:36

Yeah. There's, there's two sections there. One is the synthetic data gen with coding, and the other is the pre-trained mix of their code and reasoning sample, where they also ha- they trained a second classifier to... So, like, one of the takeaways there was like when you're doing pre-processing of 15 trillion tokens, you actually can't just run inference of, like, even Llama with all, even Meta with all the GPUs they have, they couldn't afford to just throw Llama 3 inference as, like, this whole 15 trillion token set.

So they trained, like, a code and reasoning classifier on DistilRoBERTa, which is, like, a small original encoder-decoder transformer, to try to, like, annotate out their web scrape data for quality and whatnot. Um, so they have it both there in the pre-training set and in the synthetic data gen.

There's a really good, um, quote tweet that went on about all this code gen. It's by, uh, Eugene. I will share screen and throw him on the, on the stage if he wants to talk about it.

Synthetic Data30:41

Eugene Yan30:41

Yeah. Thank you.

Vibhu30:42

Um-

Eugene Yan30:42

Um, I'm currently commuting-

Vibhu30:44

So yeah, no problem

Eugene Yan30:44

... but I'll, I'm, I'm finding-

Vibhu30:46

Okay

Eugene Yan30:46

... a quiet space right now. All right, great. Thank you, Vibhu. Um, yeah, I can, I can talk through it.

Vibhu30:50

No.

Eugene Yan30:50

I think... Can you hear me fine?

Vibhu30:52

We can come to it in a few minutes. If you... Yeah, if you're, if you're commuting, we can come to it in a bit. We can talk about it-

Eugene Yan30:56

No, I'll be commuting for a while.

Vibhu30:58

Okay.

Eugene Yan30:58

I'm walking to the stadium right now.

Vibhu30:59

Okay.

Eugene Yan30:59

Team event.

Vibhu31:00

Okay.

Eugene Yan31:00

But I'm finding a good space to sit.

Vibhu31:02

Okay.

Eugene Yan31:02

Okay, so I think what really stood out for me in this paper was that how much automation and augmentation was, was there, right? Like in the first one, you can see they actually use Llama 2 to filter out bad data, right?

If you see, um, and this is in the pre-training step. So essentially what they're saying is that, hey, we trust Llama 2's judgment well enough to be able to do that. And if you, if you scroll down, next slide.

And then over here you can see that they actually trust Llama 3 to do tag intention. They actually tag, um, tag the generated data or the, the responses based on intention, and they also classify things based on difficulty, right?

And, and the thing is, the diff- the mod- uh, they actually adopt some kind of curriculum learning where they start, they start with a single shot prompt or rather a, a single res- a single turn, uh, prompt and response.

And then after they, they move on to multi-turn. Next slide.

And then after that, uh, and, and this is the, this is the, this is the code expert that everyone's been talking about, right? So what it means is that in order to get Llama good at code, as an intermediate step, they had to train a code model.

And that sounds quite crazy, right? Uh, I mean, for me, I mean, sometimes training such large models just seems so, to take so much effort to curate the data, um, to set up the info and everything, but it seems completely essential in, in this case.

They, they could not have done it without that. And Andrej Karpathy had a great tweet about this, whereby, um, every model distillation and models and synthetic da- data generation is really now a stepping stone for the next better model.

Next, please.

And then the same thing here is... Oh, okay. And of course, here, th- this is just an example of how, how much you trust the synthetic data, right? Uh, the model was prompted to generate problems, then solve each problems.

So I'm focusing only on the green highlights here. Solve each problems, and then they give the model the errors, and then they ask the model to solve the errors. And then the model also generates the unit test, which they then use to evaluate the generations on the unit test itself.

It's, it's like you see that the human is very minimally in the loop. Um, and then if you, if we move on, um, and, and you see this pattern everywhere, like multilingual, Eugene Chi already talked about it. They, they use the...

One thing that's interesting here is that they generate, they use Llama to generate data for target capabilities, and then they back translate it into docstrings and comments. So that's how they can teach the model to, uh, explain code.

And then they use those tweets and comments, uh, uh, those docstrings and comments to actually create code again. And then we're gonna, we, we're gonna go through the rest really quickly. It's like multilingual, the same pattern here. Math and reasoning, the next one.

You see it's the same pattern whereby the model actually augments the training data with the step-by-step. So one thing that's really interesting here, right, in the sense that they actually went the extra step to, no pun intended, to actually train stepwise reward models.

That's kind of crazy, no? I mean, they, they wanted each step in the chain of thought to be so good that they actually took the extra effort to change step, to train stepwise reward models, which they then compli- combined with Monte Carlo tree search to, um, to improve the reasoning traces.

And then you see long context for, uh, uh, synthetic data for long context is the same pattern. Q&A. And then you, a- as you scroll down, you see synthetic data for image captioning and, uh, synthetic data for factuality.

All, like factuality, essentially all of, all of it is just synthetic data if you look at this. Uh, I think time will tell whether this really works out well or not. I think we're still too early on the evals.

And then you see synthetic data for adversarial examples, synthetic data for the image encoder training, uh, where they use image captions. And then they augment existing datasets with new instructions and responses. And what's really interesting here, uh, on the second last tweet, is that their human annotators were actually, um, augmented with model in the loop, right?

And, and if you think about it, this slightly represents a shift in how some folks are thinking about it, right? I mean, a lot of people is like thinking human in the loop. But no. Now it's model in the loop.

Like I ver- whereby you use the model to, to create an initial generation that the human then can edit, and it just makes the edit so easy for the human, right? And then the one big takeaway from all of this, um, and, and that's all I had from this paper.

But the one big takeaway from all of this pa- from all of this is that can you imagine how much Meta had to educate and upskill their SDEs or their existing scientists to use this new technology to be trusting of it, and their annotators to, to trust the new technology and, and to just work based on that.

So that was quite eye-opening for me, and I think it sort of suggests that the long term, here's where the puck is heading. Um, and that's all I had. Thank you.

Vibhu36:10

Awesome. Um-

Eugene Yan36:11

Thank you, Vibhu.

Vibhu36:12

Eugene coming in clutch with, I threw him on the spot. He's commuting and already had slides and tweet thread. Um, but yeah, uh, what other topics are there? I've, I've got chat.

Guest 436:23

Did, did they mention-

Vibhu36:24

Eugene, you're muted.

Guest 436:25

Did they mention using chain of verification prompting,

Eugene?

Eugene Yan36:35

Um, do you mean chain of thought prompting or chain of verification where they try to verify the chain?

Guest 436:41

It- the latter.

Eugene Yan36:43

Okay. I don't think they actually did that, but they did mention they had stepwise reward models that actually chaps, checks every step in the chain of thought. So but I didn't, I don't recall seeing chain of verification. Sorry.

Guest 436:57

Okay. Thank you.

Eugene Yan36:59

Welcome.

Guest 337:00

Eu- Eugene, uh, like early last year there was tree of thought and, uh, some iterations, uh, with Monte Carlo, uh, search and this tree of thought stuff, but, uh- At that point, LLMs weren't good enough to verify, uh, or provide enough signal for this multi-step reasoning things to happen and, and things in the loop.

Do you know, do you have some idea how they solved it or why they were able to make all this progress? Basically lo- use all the tricks we were reading about maybe half a year, a year ago, but, uh, and it seems they, they actually got them to work, so the, the...

Yeah, I, I wonder what made it work.

Eugene Yan37:38

Yeah, I don't know. I'm very interesting, I'm, I'm very curious about that as well. I wish there was more papers showing how to use m- uh, Monte Carlo tree search and actually get it to work, uh, and share more details about it.

I'm afraid I haven't, I haven't seen too much of that in this current paper.

Guest 337:53

Is it mostly for coding that they, they employ this, or for other tasks as well? 'Cause for coding you could signal back some reward, but for other things, uh, like I don't know how you evaluate things and propagate, uh, information and validate the chains of thoughts-

Eugene Yan38:08

I-

Guest 338:08

... that one.

Eugene Yan38:09

Yeah, if I recall correctly, it was actually in the math and reasoning, uh, section. So whereby they actually use stepwise reward models to evaluate each step, to score each step in the chain of thoughts. Um, so that the final output gets better.

Guest 338:26

Thank you. I'll look into it more. Thank you.

Guest38:28

Yeah, it's, it's the math. Uh, Lightman et al. is the, uh, is the citation.

Eugene Yan38:38

And, and I guess I'll wrap up with one final thing. I'm sorry it's a bit noisy. I think y'all should-

Vibhu38:43

Oh, thank you.

Eugene Yan38:43

I think, um, Swix's Latent Space podcast with Thomas, he really goes really deep, and he's, he has a strong opinion on, on synthetic data, right? I think, I think listening to that podcast will give you a lot more insight into how Meta is really embracing synthetic data.

Um, so I, I found that podcast quite helpful.

Vibhu39:05

And this was the, um, this was the Karpathy tweet about synthetic data?

Also, yeah.

Eugene Yan39:15

Yes, I think-

Vibhu39:15

Great, great podcast

Eugene Yan39:16

... I, I think that's the one. Um, exactly... Wait a minute. Uh, I think that, that, here, one thing in the, in the sense that everything is a step for the next one. No, uh, uh, no. Not, not this one.

It w- he, it was actually a tweet about smaller models, about how the competition for smaller models is going backwards. But buried in there, uh, if you scroll down a little bit more.

Vibhu39:41

Yeah, this one backwards.

Eugene Yan39:42

Yeah, this one. Yeah, this one.

Vibhu39:43

Auto-

Eugene Yan39:43

Exactly this one.

Vibhu39:44

Okay.

Eugene Yan39:44

You can see the models have to first get larger before they get smaller, right? Uh, the, the three line paragraph. And then there's a staircase of improvement, where one model helping to generate training data for the next. It's almost like he had read this paper up front, uh, and he was alluding to that.

I don't know.

Vibhu40:03

Yeah, it's pretty interesting to see. Also, this, this was a tweet that came out even before, but, like, very much the small model distillation work, it's, like, pretty huge. And that's, I think, the big part of the license play in this, too, where, um, they, they did actually finally change their license to allow people to generate synthetic data, train on outputs of the 405b.

I think the 405b is a little over-hyped for just using it for inference. Like, when it comes to inference generation and, like, like, cost effectiveness, mixture of experts are pretty efficient, right? They use less, um, they use less RAM at inference time, and, like, they're, they're more cost effective for just the best quality.

But then this is, like, really valuable for synthetic data gen, for filtering stuff like that, and that- that's what I see more of it. Um, I know, Sean, Swix, you also had a pretty good writeup about this. Um-

Guest40:58

Sorry, about-

Vibhu40:59

Want to go over that or any other-

Guest41:00

About what?

Vibhu41:01

About, uh, how this is a synthetic data gen model or any other, any other topics we wanna dive into. Also open to, like, everyone else that's in the call, too. If anyone had anything interesting that they wanna dive into on the paper, pop in, share your thoughts, you know?

Guest41:15

Yeah.

Eugene Yan41:15

Uh, I think Sachin's hand has been raised up for quite some time.

Model Comparisons41:15

Vibhu41:18

Yeah, Sachin, go ahead.

Guest 341:19

Yeah.

Vibhu41:19

Oh, go ahead.

Guest 341:20

Yeah, so this was for, um, Eugene and Vibhu also. So we saw Sonnet actually take over some time back, right? The, so and there's still some, what do you call, gap to cover. So does Sonnet have something, some other tricks in their bag which is getting them that higher up?

Guest41:40

Uh-

Vibhu41:41

I know someone's working on a writeup about this.

Guest41:44

No, no, no. I, I, I, uh, I abandoned the idea. Um-

Vibhu41:47

Oh.

Guest41:48

Y- yeah, so they, they never published anything about, um, the, uh, the, what tricks they use. Uh, but the evidence strongly points to the fact that they used the steering vectors that they had from the scaling mono-semantici- scaling mono-semanticity paper.

Um, the, the main evidence is that they happened to do this mono-semanticity research on, uh-

Eugene Yan42:11

There we go. Yep

Guest42:13

... on, on Sonnet.

Eugene Yan42:13

What are you guys doing, you little freaks?

Guest42:15

Um, SJ is, uh, con- constantly screwing up his mic. Um, they, they did it on Sonnet, and obviously they only shipped 3.5 Sonnet. Like, s- that's the, that's, like, the, the smoking gun. Like, if they, if they actually had anything else, any other trick that caused 3.5 Sonnet to be so good, uh, uh, it would, they probably would have deployed it on Haiku and Opus as well.

Uh, the fact that they don't is, is proof positive that it's, it's basically the mono-semet- semanticity stuff. Uh, does that answer your question? Do I need to explain what that is?

Vibhu42:45

I have a hot take that it's the opposite.

Guest 342:48

No, it does.

Vibhu42:48

I think it's not, I think it's not control vectors.

Guest42:51

Uh.

Vibhu42:52

Um.

Guest42:52

Yeah. Why?

Vibhu42:53

They could be. Well, so, like, they, they did say it's a larger model. If you look at the training data and the training date for when Claude 3 Sonnet came out to 3.5 Sonnet, it also has a year and a half of significant data updates.

So I think that there was a lot of research that put out on, like, high quality synthetic data, the post-training mixture of it, and Sonnet probably just had a decent bit of, like- You know, there, there's a lot more that they could squeeze out of it.

Also, they did say it's bigger, so, like, lot more data, lot more, like, research and good quality synthetic data. The pre-training, like, data mixes started to come out. Um, it's bigger. I think it was just a lot more post-training as well because there, there was quite a bit of time-

Guest 343:36

In post-training there's a... Sorry, Vibhu. In post-training in Sonnet-

Vibhu43:40

Yeah

Guest 343:40

... I mean, you can see there's this pause, thinking pause tokens, where sometimes it generates a token, and it kind of does internal thinking and then generates the answer. So it seems like they use some recent tricks where, um, people say, "Hey, like, you kind of need to think step by step," but maybe not materialize directly in the answer the th- step by step thinking.

So, um, sometimes when you run Sonnet generations, you can see there's, there's some, uh, it stops in the middle and, uh, and people saw that, uh, those are actually paused for thinking and, uh-

Vibhu44:12

Yeah

Guest 344:12

... and, um, side thoughts. So it seems that really helps with, with the reasoning tasks quite a bit, so that's one additional trick that you, they use. But, uh, yeah, uh, I agree that it's been one year of work, so they probably have lots of tricks in there, not, not just one, two, three.

Like, uh, much like in the Llama paper, you'll see that it's, uh, hundreds and hundreds of people. Still less than 1,000, but, um, yeah, it's like a Manhattan Project to build one of these things.

Guest44:40

Uh, I actually counted the number of people. La- Llama 3 had 200 core contributors, um, which is good. It's pretty small. It's less than Gemini, which had 950. Uh, so um, uh, Sebastian says, "What does thinking mean?" Okay, here, here's, uh, here's where we get philosophical.

Vibhu45:01

Um, my, my quick take on that is this used to be a thing with, like, the original ChatGPT web UI stuff of, like, why is it pausing at stuff? I think some of this is also just the, the way the API works, right?

So, like, what's the inference it's running on? How's the streaming? What's the API service like? Sometimes when there's blocks, it's not that something else is going on, it's just that, like, there's a delayed stream of your API response, and sometimes people overanalyze that.

Is it that it's thinking? I- is it what's going on? Um, maybe, maybe not.

Guest45:31

Oh, no.

Guest 245:32

Um, but I think for the Sonnet's case that, uh, some, some users have already used prompt in- I guess, prompt injection to trick it into, instead of doing the XML thinking block, and to use it to output it in a different format.

It's, then you literally get to see the thinking process.

Guest45:50

Yeah. So f- for what it's worth, uh, we actually, I, w- I inter- I went to iClear and interviewed the, uh, pause token author. Um, it's, uh, it's on the iClear episode if people want to check it out.

Um, I do not think that, uh, Claude specifically implemented that version of thinking. Um, I, I think it's much simpler. It, it is just chain of thought. It is just prompted, uh, you know, XML chain of thought, uh, that, that, that is then subsequently post-processed and removed inside of Claude Artifacts.

Um, yeah, and, and, but it, it's still, it's still a form of thinking. It's, it's, it's a form of chain of thought. Um, it definitely improves the performance. Um-

Guest 246:24

Right. Sorry. Yeah, that's what I was aware of. What... I'm sorry, I missed what Eugene mentioned for, like, an alternative to what you described as like a ch- a chain of, a chain of thought that is not presented to the user.

Eugene? Yeah. So i- it's instead of, like, creating a custom token, which is the pause token concept, they literally just got the mod- through prompting, got the model to, to reply with a thinking XML block, which then you can trick it through prompt engineering to, like, substitute the tokens respectively.

Then, then suddenly this technique becomes very blatant when, when, when you get to see it, see it respectively because it's no longer hidden from the UI. Um, yeah. I think also r- uh, on a separate line with that, because since a, a lot of people are looking to the eval, and then they are like, "Hey, some of these evals are doing worse than, let's say, 40, right?"

Uh, or things like that, right? The fact that it's already closed itself, right, means that, yeah, if you are just, like, a few weeks away until, like, every single benchmark, right, you're going to see, like, a, a point jump because someone fine-tuned a code specific version of the L3 model, uh, or a code specific ver- uh, or, or, uh, a medical reasoning specific version of this model.

It's gonna take slower than normal because I spoke to s- some people in the fine-tuning community, um, the biggest hurdle has been, what do you mean you need at least three nodes of H100 to start the process? Yeah.

The, the, the amount of vRAM requirement is kind of huge. Uh, I suspect we are going to see more LoRAs first before we get full fine-tunes. Um, also the most random part, um, I know Meta did this for, for good reasons, uh, because they, they, th- uh, they basically did a lot more keyword filtering from the sources.

But a lot of, uh, people in the AI, uh, companion space, they were like, "Nooo," basically.

Yeah, makes sense. I'm gonna look into that. Thanks.

Guest48:26

Um, I, yeah. Do we have more things on the, on the paper? I mean, there's, there's more to discuss. I feel like everyone's being too polite.

Vibhu48:34

There's a lot of, uh, new scaling laws they brought up. They, they had a whole recipe for post-training, how they did it, how they did their SFT, how they did... They also released both the base and the instruct models, how much of this was done by synthetic data, um, how they trained their, like, image video adapters, all that stuff for multilingual stuff.

They, like, give out a whole recipe on, like, how to do this, and it's a, it's a long, long read. But for anyone that hasn't read a lot of papers, this is also probably a really good one that's like a very approachable, very readable, not too crazy technical one to at least understand what's going on.

They go into some of their, like, evals on their multimodality, how their adapters work, how it performs, and they're like- That's pretty good. They added speech into, um, speech understanding, how to train a speech encoder, how many hours of recording they used, how they filtered it.

Like, they go through literally all of this. And, like, this is probably where, like, the, you know, you could have an hour on this paper that's like every step of it. Um, but it's, it's an interesting one where, yeah, they do go into all that data set, how they transcribed it.

Like, just little, little stuff, too. Like, in their, in their speech understanding section, there's a section on, like, our ASR training data contains 230 hours-- 230,000 hours of manually transcribed screech recording that spans 34 languages. So, like, you know, just a little one line of, like, "Oh, we casually ma- manually transcribed 230,000 hours of, like, 34 languages of speech, and we're just training a little adapter for this that we're not releasing."

So, like, they, they put a lot of work into that. And then it goes even deeper into, like, how do you use this for pre-training? What about, like, spoken dialogue? How do we fine-tune? So, like, what's the recipe for a speech adapter in an LLM?

Like, yeah, we did a lot of pre-processing to the base data set of, like, manually transcribe a lot of speech recording, have multi-languages, train it out, h- here's, like, the speech length segments that we want. Then we fine-tune this, like, adapter for spoken dialect.

How do we do that? Well, we synthetically generate responses for prompt. We ask, like, for transcripts. We generate them. They generate, like, 25,000 hours of speech synthesis through voice box, which is, like, a whole nother series that Meta has put out around everything, like how did they, um, how they do, like, voice generation.

So, like, they have a whole really good breakdown paper of th- of that, how they use that to generate model, to fine-tune, and generate synthetic data for this. So, like, there's a lot in here if anyone's interested in.

Um, a lot of that doesn't make it to Twitter, but dig in, present it.

Um, but yeah.

Architecture stayed the same. I thought the interesting parts were also just, like, they wanna keep it very foundational and see what works and what they can just scale up. Um, the, the second aspect to that is, like, it'll be fun to see when they start cooking with, like, actual MOEs, how do we, like, you know, go outside the box.

But high level also, like, it's, it's nice to have clarity on their scaling laws. Like, I've definitely presented too many times that they just scaled up and prayed, and they took an AP to 15 trillion and, like, they were very inefficient.

And then other, other papers, like five three took a big shot at this, right? So five three's whole paper is about how we had chinchilla scaling, then we have, like, this, uh, inference optimal Llama pre-- like, scaling, and then here's how you can do what we think is right.

And then now Llama puts out a paper, and they're like, "No, this is actually all based. Here's, like, new scaling that's not just next token prediction. It's grounded on reasoning. Here's how scaling laws work. Here's how you can use it.

Here's why we did it." And it's like, well, it's gonna make sense. Um, yeah.

Guest52:32

Yeah. On the scaling part, I find, I found it interesting and funny that they, they were using ARC as a measurement for-

Vibhu52:38

Yeah

Guest52:38

... for the, the scaling training. Uh, one of the weirdest take that I had in my head, it was like, so Facebook spent over $100 million at the ARC Challenge w- to try to win the million-dollar prize.

Vibhu52:50

I think it's a different ARC, right? If I'm not mistaken. It's not the same as the million dollar-

Guest52:54

Yeah, but it, it's not the same data set.

Vibhu52:55

But still, like.

Guest52:55

Yeah, but-

Vibhu52:56

It's, uh, it's-

Guest52:56

It's still in the lines.

Vibhu52:58

Yeah. Um.

Guest53:03

Um, and-

Vibhu53:03

But yeah, lot of...

Guest53:05

Yeah. Lot, lot of, lot of good stuff. Uh, in the five to seven minutes we have left, uh, I wanted to give some time to, I guess, Hassan, if you wanna... Hassan's actually built an, an app that's kinda cool with, uh, Llama 3.1, and, uh, maybe there's something to learn about, you know, prompting it, building with it.

Anything surprising?

Building Apps53:22

Hassan53:25

Yeah. Thanks. Uh, thanks Sean. Hey, everybody. Uh, I just wanna talk about this, this app that I built real quick, more definitely a lot less technical than we're talking right now. This is-

Vibhu53:35

Whoo

Hassan53:35

... dropping all the way down to the-

Guest53:36

All about building

Hassan53:37

... the ground layer.

Guest53:38

Yeah.

Hassan53:38

Yeah. I, I guess it is all about building. Uh, but I, I just built this little app, you know, it's, um, it uses a search API to put in whatever topic you wanna learn about, like, um, quantization, and it can explain it to you at kind of any level you want.

So let's learn about quantization at an elementary school, uh, elementary level. Uh, so it'll basically use a search API, grab all these sources, and put it into context, 'cause obviously Llama 3.1, larger context, you can fit in a lot of stuff, which is great.

And, uh, I've noticed that it's, um, it's pretty good at, you know, kind of at dumbing down concepts, but also responding to the prompt a little bit better. Um, and, and almost responding, like, if I set a system prompt it, uh, you know, kind of details how it should behave over the next few messages, I found that it responds a little bit better.

Like, for example, for, for this, um, for this system prompt, I had, like, make sure you make the overview really short, 'cause, um, when I was testing on, on just Llama 3 and on other platforms, it was giving me a really, really long initial answer.

Um, and so I want it to give a really short overview, but at the same time, I want it to be detailed in the future. And I also want it to include quizzes at certain times. And so, um, I just found that it's, that's a little bit, uh, it, it was a little bit smoother at, at kind of, um, at, at responding.

So yeah, here it's, it's gonna, you know, actually try to give me a quiz, um, and try to just be interactive and kind of teach me a subject at, at any level that, that I want. So, um, it's, uh, fully open source.

Llama Tutor, uh, llamatutor.com. Um, it's definitely open source. I misspelled this. Uh, and, uh-

Guest 455:10

Did it give an example of quantization math?

Hassan55:16

I don't actually know over here. I think it's talking about, uh-

Guest 455:20

I just got into a debate on a call where someone was making an obvious error in quantization math and how much the memory before versus after

Hassan55:31

I'm not sure.

Vibhu55:31

I can send this to them.

Hassan55:32

I'm not sure if it'll... I, I know it does formulas as well. Yeah, so it'll put some... I should probably format these a little bit, a little bit nicer. Um, but-

Vibhu55:38

That's awesome, though.

Hassan55:40

Yeah. Thank you. I... It has... Yeah, I got about 4,000, uh, visitors and, and if anybody's curious about cost as well, so about 6,400 requests. Um, I had some errors 'cause I was playing around with, uh, I was hitting, like, the context limit and, and a bunch of other stuff.

Um, and also, you know, using the Together API, uh, very biased, I, I work over at Together. Um, but the cost for anybody curious, so we have about 12,000 average tokens per request. We have about 5,900 successful tokens.

Um, and so if you do the math, that's like 74 million, uh, tokens, and if I go over to our pricing, right now we're at 18 cents per million tokens for 8B 3.1. Um, and so that comes out to about $12, um, from the 72 million tokens that, that I used in the last 24 hours from, from these, like, 4,000 people and, and 6,000 requests.

Vibhu56:30

Wow.

Hassan56:31

So that's, uh, that's all I have.

Vibhu56:34

And so... Wait. C- um, wait. Can you show... Did you put a link in the chat? Or can you type the screenshot?

Hassan56:41

To that? No, I'll, I'll do that. Yes, I'll put in a link to the, uh, Tutor-

Vibhu56:47

Oh, Llama Tutor. Okay.

Guest 656:49

Hasan.

Hassan56:50

LlamaTutor.com.

Guest 656:53

Hasan.

Vibhu56:54

Beautiful.

Guest 656:55

Quick question at a high level. Um, is this using some sort of, like... Have you ever used the GPT's actions? Is it using a system similar to that?

Hassan57:08

I actually haven't used actions. Can you tell me about that?

Guest 657:11

So actions is where you input an API spec and the model itself can make the calls, um, the model itself can decide to make the calls to that API, provided that API spec.

Hassan57:23

Oh. Oh, so like function calling, basically? Um-

Guest 657:26

Yes. It, it is-

Hassan57:28

No, I'm not us- yes. Uh, I'm not using it on, on this app. This was, like, the most simple example I could do where I give it a system prompt, I do this API call for search, I parse all the sources, and I just drop everything, and I'm like, "Hey, build something interactive."

Um, so this is kind of step one. There's obviously so much more I wanna do. I wanna try to do, um, generative UI maybe, uh, with the Versailles ISK, where I show for, for the quiz example, I show an actual, like, quiz component that, that renders.

Um, I can, like, generate a little, like, report out of all the sources and have like a, like read this, you, you know, like, you, you could read this to check it out or, or flashcards or... Like, I, I feel like there's a lot of directions to go with it, but I kind of just wanted to build something really, really quickly.

I just built it over the weekend. Um, we, we got early access to, to 3.1 'cause we were a, a launch partner with, with, uh, with Meta. So kind of just playing around with it and, and try to build something really quick over a weekend.

Guest 658:22

Ooh. Okay. I missed that detail. That's really cool. All right. Thanks.

Serving & Quantization58:27

Guest 758:27

So you're doing the retrieval-

Guest 358:28

I'm curious if you could share what does it take to serve a 405B model on Together?

Hassan58:36

What does it take to serve 405B? So we're serving, um, we're serving... What is it? FP8. It takes eight A100s, for instance.

Yeah. But we're, we're looking into quantizing down to int4 and trying to serve on four H100s, but, um, in progress.

Vibhu58:58

The math is roughly one gig to one gig of VRAM, right? So at, at FPA, um, yeah, one to one, and then you could scale that out. There's, there's pretty good resources on this. I feel like we've shared them in Discord.

And then someone asked about quantization stuff. Together put out pretty good, like you're now running quantized models and a good blog post about this. It's somewhere in the Zoom chat, and we'll probably share it in Discord too, but good resource on all that.

Uh, al- also the paper has a section on inference, of course. So how do you run a 405B? What does that look like? What's efficiency in that? They have a whole section on this. Um, they have sections on, like, previous meta work of, like, how do we have more token efficiency, so, like, multi-token prediction as a pre-training task.

How do we have, like, you know, other parts of what you might see? Meta's, Meta's done a lot of work on this. But, uh, overall, that's kind of a high level overview and our thoughts on the paper. If anyone else has questions, we can discuss them.

Otherwise, next week, um, we were supposed to do WizardLM and Orca 3, so, like, heavy synthetic data gen stuff this week, today. Uh, we pushed that to next week, so sorry this was, like, not super prompted. Last, like, you know, like, basically less than 24 hours ago we decided, let's switch to this paper club.

So we didn't have crazy slides, but next week we'll have stuff for, um, Orc- uh, yeah, Orca 3 and WizardLM. Same paper club, same time. And then at some point I might do a deeper dive into this, like a proper one-hour breakdown of all this.

If anyone's interested I'll, I'll share somewhere.

Guest 71:00:36

I, I'm curious about the quantization thing and whether, like, if you had the choice of a larger, like a... Sorry, more parameters, but worse quantization, how do you, how do you figure out the trade-off between that and running, say, the, the 8 bill- like, say, the 70 billion model without quantization versus the 405 with quantization to make it down to the same size?

Guest 21:00:59

Uh, people generally-

Guest 31:01:00

I think-

Guest 21:01:00

... do benchmark it anyways. Uh-

Guest 71:01:02

Okay.

Guest 21:01:03

I, I already know th- this, this already happening right now because a lot of the enterprise communities are actually trying to figure this out. Um, and for anyone who wants to try at 4-bit quantized, right, you can actually run it on 16 4090 cla- uh, uh, uh, GPUs, uh, for if you, if you're doing it, uh, 4-bit quantized.

Uh, and I'm really seeing folks doing that. Uh, whether that's a good idea, whether the model will get dumber or not is something we'll probably find out in the next few days.

Vibhu1:01:34

At a high level though, the, the larger ones see less of a hit when quantizing than the small ones. Uh, r/localllama has a lot of comparisons and benchmarks of different quantizations. So basically, like, a lot of the prosumer rig is a dual 4090 or dual 3090 system, right?

So you've got about 48 gigs of VRAM. With 48 gigs of VRAM, you can run, like, a 70B at 4-bit quant. And then, you know, a lot of people will spend, like, $3,000 on a local rig. They can run a 4-bit 70B, or they'll look at, like, how does that compare to, you know, a 34B at higher precision or a single GPU where you're running an 8B at, like...

So there's, there's solid breakdowns of benchmarks of these. Uh, this is where Reddit people are doing good, good work, you know? Um, as opposed to, like, when you look at it from the infra provider side of, like, what's the benefits of quantizing.

There's, there's also efficiency in, like, QLoRA fine-tuning and quantizing and doing it. But benchmarks-wise, like, the LBR is, like, bigger models take less of a hit, smaller models take a bigger hit, and then, like, speed inference. Alex, you have thoughts?

Alex1:02:45

Not thoughts. I have a question, actually. I don't usually have questions, but this time I have a question about effects of quantization. W- what, uh, in your guys' experiences, uh, are the most visible effects of something like quantization?

What comes to mind when, like, you see that the model is quantized, and how quickly-

Vibhu1:03:06

It's just dumber

Alex1:03:06

... you realize, "Oh, okay, this is effects of quantization"?

Guest 21:03:08

It's just dumber. That's all. Um-

Vibhu1:03:11

Well-

Alex1:03:11

Dumber in any specific areas? Is it dumber knowledge-wise? Is it dumber logic-wise or any specific areas, or just, like, overall dumber?

Guest 21:03:20

Uh, one thing that I've seen a lot of community members do, right, is that when you over quantize, right, uh, at longer context length, right, it starts going into repetition rapidly.

So that's like we have gone too far line.

Vibhu1:03:38

Yeah. In, in a lot of the, like, 1-bit quants as people ineffectively quantize stuff, it starts to just go into pure chaos, right? So it g- it becomes stochastic random next token prediction. Uh, the, the little trade-offs that you see at, like, what type of quantization you wanna do, it starts to perform worse at, like, chain of thought, at long context, at, like, needle in the haystack.

Some of those things start to degrade earlier on in quantization as opposed to, like, pure performance. And then, yeah, sometimes it's just dumber. Like, it's just worse on, like, some of the trivia QA type benchmarks, which is interesting because it's like...

Or sorry, it's, it's, it's worse on reasoning benchmarks, but decent on trivia. So trivia is where it's, like, under internal knowledge, what does it already know. It doesn't lose facts, but it loses reasoning. It loses, like, verboseness of responses and whatnot.

So it, it, it degrades in that sense. Um, but then you, you do have benefits of, like, yeah, it's, it's significantly more efficient to run, right? 4-bit versus 8-bit, you can run in half the hardware. You get speed ups.

You can fine-tune more efficiently, so there's trade-offs. Though the one interesting thing I'll note there is if anyone saw the Apple LLMs, like, at their WWDC stuff, they put out a Apple research blog on how they have all their local LoRA swaps, and they dropped in, like, a little paragraph where they're like, "We te- we, we tested our models at, like, a net of, like, 1 or 2-bit quantization, and we saw no performance drop compared to full precision."

So they're basically like, "Yeah, we benchmarked it. We got lossless quantization." No details, no nothing. But they, they seem to have stated they figured it out, which, you know, skeptical, but they're running quantization for on-device.

Alex1:05:22

Yeah, but on specific tasks for them. On very specific-

Guest 21:05:25

It's also gonna be very... Sorry. It's also gonna be task specific. We have seen weird edge cases where s- where when you quantize a model one step, for certain tasks it actually improve, uh, degrading all other task.

Vibhu1:05:42

That's true, too.

Alex1:05:42

I'm writing, uh, I'm writing an email now between cloud providers for 3.1 70B just to see if, like, you know, Together's, uh, FP8 versus, I don't know, Fireworks or something like this has a difference, or versus Groq. So if I have any, uh-

Vibhu1:05:58

Have you seen a Artificially analysis?

Alex1:06:03

Oh, yeah. I think I saw something.

Vibhu1:06:05

The, um-

Alex1:06:06

The quantitative, the quantitative-

Vibhu1:06:07

These guys, right, they run pretty good benchmarks.

Guest 21:06:09

Yeah, but they, they don't do quality. They just do reported, uh, numbers, I think. Um, so they just-

Alex1:06:15

Yeah, they do-

Vibhu1:06:15

What they do is compare infra providers. Oh, they, they're not doing their own benchmarks? I thought they have a whole heap of benchmarks themselves.

Alex1:06:20

They're doing pricing and they're doing speed. I don't think they're doing quality. They're not-

Guest 21:06:24

They don't, yeah, no, they're not doing benchmarks. Yeah.

Alex1:06:26

Yeah.

Guest 21:06:27

I, I think-

Guest 81:06:27

Quality is much more difficult to benchmark, to be honest.

Alex1:06:30

Yeah.

Guest 21:06:32

So I'll, I'll add to Alex. So I'll add to Alex, like, what he was asking. I know credible reports of people where they're saying between Groq and some other providers, the quality is different in the sense everything else remaining the same, the inference engines are giving wrong answers.

So I'll just leave it at that though.

Alex1:06:56

I mean, one thing I can tell you right now that I noticed that Groq, uh, temperature zero returns different responses every request. So I don't know if temperature zero actually works there. Uh-

Guest 21:07:05

Yeah.

Alex1:07:05

But yeah.

Guest 21:07:06

Yeah, that, that, that has been, that has been, uh, a lot of people have said that to the CEO also, so.

Guest 41:07:16

Wouldn't, wouldn't temperature zero supposed to be the most deterministic-

Alex1:07:21

Yep

Guest 41:07:21

... instead of-

Guest 81:07:23

Depends on the framework, to be honest. For opinion, I think it's, uh, temperature is zero.

Guest 41:07:31

Well, but theoretically it's supposed to be the most determinist, what I would expect based on- What temperature is supposed to do that it would be the most deterministic?

Guest 21:07:42

Yeah, I guess-

Vibhu1:07:46

So there's a re-

Guest 21:07:46

Oh, go ahead, Vibhu.

Vibhu1:07:48

So sometimes there's little reasons, like when they're running inference, right? They might run it on different hardwares. Like, maybe there's a server, server in West Coast, East Coast that's different GPUs, different CUDA kernels, different drivers, and some of that affects how you do rounding, how you do next...

Like, what... Yeah, basically, how is inference being done, so you see slight variation. Then across inference providers, this is their, like, secret sauce, right? Are they doing self speculative decoding? Are they doing speculative decoding? How are they doing it?

So, like, all these little things have differences, but there's also the reason of temperature is like, yeah, different hardware, different, uh, drivers, different all that. So some of that answers it, but I don't know, Eugene, maybe you have more to add.

Guest 21:08:29

I was just gonna say that-

Guest 41:08:30

Mm

Guest 21:08:30

... for GPUs, floating point just on work... I mean, for GPUs with floating points, and you push it through so many calculations and so many math models, the floating points aren't just, just, just not gonna be precise. So that's why even if temperature is zero, it's not gonna be the same throughout, uh, for multiple requests.

Guest 81:08:49

Yeah, even if it's like the same-

Guest 41:08:51

Uh-

Guest 81:08:51

... GPU on the same request. Like, the, the order you do reduction on the math models, as Eugene mentioned, will affect the results because, like, floating point calculations is not deterministic. Like, uh-

Guest 41:09:03

Yeah

Guest 81:09:03

... you can try this in Python and add 0.1 plus 0.1, you might get 0.19999, not 0.2. So, like, the, the order you're doing the addition and multiplication can affect the results, and if you're doing like 1,000 addition multiplications, it's difficult to get the same order every time.

Guest 21:09:21

It's about that very small 0.0000001% noise that ends up cascading all the way. Um, if you want true temperature equals to zero, the, the known reliable way is to run it on CPU, 'cause the CPU do, do have the floating point precision, uh-

Guest 41:09:38

Mm-hmm

Guest 21:09:38

... reliability. But if you're running on CPU, you're gonna take forever. So...

Vibhu1:09:44

Yeah. And, and to add, um, to Eugene also, there are actually like five or six different class of, uh, floating point, uh, what do you call? Classes that there are in the CUDA, and I believe, like, people have to hand code certain, uh...

Have to account for these, uh, what do you call? Differences, and that's where they have to be actually shown the errors, saying that on some other, uh, platform we see this error, and over here we see this answer, and now someone has to painfully go back and figure out the bug in their code.

So most of this is basically bugs in the translation layer, they're not bugs in the hardware. That's what I was trying to say.

Guest 41:10:26

I'm definitely gonna try that CPU approach. So I'm trying to get a project approved and not, uh, with the ba- the rationale being this type of inconsistency.

So also-

Guest 81:10:46

Uh, also, like, depends on the entry engine. Like, XLama 2 is not deterministic at all compared to maybe... I think Transformers is more deterministic. You can set the torch seed and the inference seed. So also the inference engine is very important.

Guest 41:11:11

Thank you.

Vibhu1:11:14

Awesome. Well, um, thanks everyone for joining in on our last-minute drop-in. Um-

Guest 91:11:19

Can I share something really quick?

Vibhu1:11:21

Yeah, yeah, go ahead.

Guest 91:11:23

Just a quick-

Vibhu1:11:23

I've, I've got to drop, but this, this, this room will still stay open. I'm gonna make, uh, someone else host. Feel free to keep discussing. Next week we have another one at 12:00, but I'm gonna make, uh, Eugene host.

Live Demos1:11:23

Guest 91:11:35

Oh, it's okay. I can do it next time. Thanks.

Vibhu1:11:37

No, no. Go ahead. Go ahead. Go ahead. Sure.

Guest 21:11:39

No, don't make me host. I'm at a baseball game.

Guest 81:11:42

No, uh, the other Eugene.

Guest 91:11:43

So, uh, it's, uh, quick. So basically-

Vibhu1:11:46

Great

Guest 91:11:46

... um, so basically I found like a Llama 3.1, like a 405b actually might do better in some domains specific question than like a ChatGPT 4o. Do you want, guys, guys want to see a little bit, just really quick?

Guest 21:12:02

Sure. Uh-

Vibhu1:12:03

Yeah

Guest 21:12:03

... and it, it's no surprise actually. I actually think that w- more and more people will find specific domains- ... a- as time will happen.

Guest 91:12:11

That's great. So let, let me show you really quickly.

So here is a s- little summary. So basically, this is the specific question I was asking. Here, you're expert in mechanistically interpretable research. How do you use residual stream? Yeah. And so here's my summary. Overall, I think, you know, uh, 405b did a better job to explain, and also there are some important f- uh, information like is missing in ChatGPT 4o, not expressly, uh, explained.

So let me show you the, uh, answer from, um, uh, Llama 3.4 405b first. I feel, I find it's really easy to follow for someone like me, like I'm now doing research in this field, and it also give enough technical detail.

So, uh, if anyone want to read a, a little bit.

Yeah. So, and now I can, uh, move to the Cha- what are the answer from ChatGPT-4 and the 4o. Do you guys want to see it?

Guest 41:13:16

Yeah.

Guest 21:13:17

Yeah.

Vibhu1:13:18

Yeah.

Guest 81:13:18

Yeah, I'm done.

Guest 91:13:19

You're done? Okay, cool. So this one is from ChatGPT-4. I feel like, uh, ChatGPT still kind of like, uh, answer things in a sort of generalized, summarized way.

Let me know if you want me move, scroll down,

or if I'm scroll down too fast.

Guest 21:13:49

I think, uh, if-- I, I think Eugene, the other Eugene would, would, uh, would, would chime in that actually a lot of these things are very dependent on the prompt. So-

Guest 91:14:00

Mm-hmm

Guest 21:14:00

... so it won't surprise me for this question-

Guest 91:14:02

Yeah, so I use the same prompt. That's a good point. Mm.

Guest 21:14:03

No, but it won't surprise me if you-

Guest 91:14:04

So I use the exact same kind. Okay. Sorry. Mm.

Guest 21:14:08

Yeah. Uh, what I meant is, uh, it won't surprise me if you tweak the prompt right. The winner ends up flipping around. That's, uh, that's what I meant.

Guest 91:14:16

Yes. Just for the argument, I didn't tweak the prompt just for, uh, Llama 45B. So I just, "Hey, have this one." I just throw it out. But, uh, definitely, you know, more detailed, uh, study need to be done.

And let me show you, uh, uh, uh, ChatGPT 4. Oh. It's better than 4, but I don't think, uh, I still think, uh, uh, Llama 3, 4, 5B did a better job. Just a second. Let me see.

So I'm going do the same thing, scrolling down. Let me know if I'm moving too fast.

So basically, I think, I kind of feel, of course, need to do more detailed study. I feel like, uh, um, uh, Llama 45B may have been doing a much better job of organize its, uh, knowledge. You know, how, the way to handle, how to answer question.

You know, how to organize the knowledge together. That's-- I think that's pretty important.

Guest 101:15:32

So I'll give my personal example that, um, when I look at some healthcare related data, and if I'm looking it through, um, basically ChatGPT 4 and then even when I use the similar prompt that I give through Perplexity, I get different answers.

But since I know the domain, I know that sometimes like ChatGPT is just bullshitting. And Perplexity-

Guest 91:15:56

Yeah

Guest 101:15:56

... actually gets references and it gives you a slightly better or more answer which I can take and go back to the references and dig deeper.

Guest 91:16:06

Mm-hmm.

Guest 101:16:06

So you, you will def- once you know the domain better, you should not fall in for the English is my take. The English may look all correct, but it might be nonsense at the end of the day.

Guest 91:16:15

Oh, I ha- I, uh, looking into a little bit about the restaurant stream. I can, I can, uh, have pretty good confidence it's, it's not bullshitting from 45B. Yeah. So the, the things that, uh, really amazed me is like how it explain.

And, uh, as me like I say, "Hey, I'm just starting getting into like a lot of technical thing I haven't done before," it seems like it just really, uh, explained really well. I can just go down to, you know, uh, do look into specific in- information, stuff like that.

But definitely I see your point, like, uh, it's just a one shot. But still seems pretty amazing because I didn't trying to tweak the prompt just for, uh, 45B. Yeah. Thank you. Yeah. So that's all that I want to share.

Guest 21:16:58

Yeah, definitely. Uh, we are going to see more and more of these. Like, like, I think what's interesting, es- especially when you're having different foundation models, uh-

Guest 91:17:06

Mm-hmm

Guest 21:17:06

... you're going to have very different default outputs because, um, yeah, we can probably steer all models, especially the bigger ones with prompt engineering to, to, to, to a certain desired effects or even like, like fine-tuning. But the default out of the box would be, uh, what's interesting because that's what, what a lot of day-to-day users-

Guest 91:17:25

Yes

Guest 21:17:26

... is what they'll experience. So like for me, for the longest period of time, I didn't care that GPT-4 had a better eval score across the board than the Claude model. The Claude model just seems friendlier and nicer to me, and I like it better.

And, and, and, and maybe that's what just matters sometime when, when they are all good enough to get the job done.

Guest 91:17:49

Yes. So I wish you, we have a better model where we don't need to tweak the, uh, the prompt, you know. In a way it can reflect the how well the model is organized its knowledge, you know. And so if it's easy, just talk to it without tweaking the prompt, actually it's, you know, in a much better, uh, advanced stage, I think.

Guest 21:18:11

But I think at the long run, that would be hard to achieve because if you just view each model a bit-

Guest 91:18:17

Mm-hmm

Guest 21:18:17

... like a individual, so you-

Guest 91:18:19

Mm-hmm

Guest 21:18:20

... you can call it Llama-kun, you can call it, uh, uh, GPT-4-kun, we- we- et cetera. It's just really, uh, uh, a preference thing when it boils down to it because like, like, like, like as I said, well, I prefer Claude because of the way he did output.

I know someone who prefer the other way. And-

Guest 91:18:38

Mm-hmm

Guest 21:18:38

... that, that's a challenge here, right? You're not, you're not-- You can't possibly just train the model to, to satisfy everyone in the world. There'll be that back and forth.

Guest 91:18:47

Yes. Mm-hmm.

Guest 21:18:47

And that's where all those like more, uh, few short prompting will come in or fine-tunes will come in. And I think that pro- is probably the s- more exciting thing, thing about Llama because we are gonna see lots of fine-tuned.

Guest 91:18:59

Mm-hmm. Yes.

Guest 21:19:02

Yeah. So-

Guest 91:19:02

Yeah. Sorry, go ahead.

Guest 21:19:04

I think, I think because since we are, we are well past the usual, uh, I-- and there's a lot of new faces here, uh, I'm just going to like re-recap and, and wind back up. Uh, so the latent space Discord, um, we host this weekly paper club, and this is why, uh, this session happened.

Uh, this week is just a spec- very special one because of the whole Llama 3 model came out. Uh, prior to that, uh, prior to that, we were, we were planning to do other, uh, other papers, and we'll probably go back to our usual, usual schedule.

So if any of you, uh, the, especially the, the new folks that joined in because this, this got broadcast on a much larger audience. Um, yeah, feel free to, uh, feel free to join us next week, and we will be talking other papers.

Guest 91:19:50

Great. Thank you. Yes. Bye.

Guest 81:19:54

Thank you, Vibhu

Guest 41:19:57

Thanks a lot.

Guest 101:19:59

All right. Have a great day, everyone.

Okay, I'll stay. Uh...

It's got organized.

See you later, Eugene. Peace.