# How to train a Million Context LLM — with Mark Huang of Gradient.ai

Latent Space · 2024-05-31

<https://addtry.com/9adf13dc-0536-4319-a0fa-ee61b066d8e3>

Mark Huang of Gradient.ai explains how his team extended Llama 3 to a 1 million token context window using curriculum learning, RingAttention, and EasyContext, achieving near-perfect GPU utilization. They employed theta scaling from the RoPE paper to interpolate positional encodings, trained on carefully curated datasets including SlimPajamas and synthetic data from GPT-4, and validated with benchmarks like Ruler and needle-in-a-haystack. Huang details the trade-offs between full fine-tuning and LoRA adapters, the challenges of pushing to 4M tokens (degradation from floating-point precision limits), and why long context matters for state management across sessions and grounding multimodal inputs. He calls for community collaboration on long-context evaluations and pairwise multimodal datasets.

## Questions this episode answers

### What role did theta scaling play in Gradient's extension of Llama 3 to a million-token context window?

Mark Huang explains that Gradient increased Llama 3’s RoPE theta value using an empirical formula from the RoPE scaling paper. This enabled positional interpolation rather than extrapolation, allowing the model to handle longer sequences as if they were seen during pre-training. Combined with curriculum learning and Ring Attention, this let them scale from 8K to 1M tokens. At 4M, degradation appeared due to floating-point precision limits.

[18:53](https://addtry.com/9adf13dc-0536-4319-a0fa-ee61b066d8e3?t=1133000)

### What datasets and synthetic data techniques did Gradient use for Llama 3’s long-context fine-tuning?

Mark describes a two‑stage approach: continuing pre‑training with filtered and concatenated Slim Pyjamas data, then supervised fine‑tuning on a curated version of UltraChat. They used GPT‑4 to synthetically rephrase chat data and generate new tokens, and ensured the datasets forced the model to attend to information across the entire context. Diversity and avoiding over‑reliance on recent tokens were key priorities.

[28:01](https://addtry.com/9adf13dc-0536-4319-a0fa-ee61b066d8e3?t=1681000)

### How does Gradient evaluate long-context models beyond the needle-in-a-haystack test?

Mark says they rely on the Ruler benchmark suite, which includes multi‑needle retrieval, variable tracking, and tasks like counting common words across the full context. These harder evaluations test holistic understanding and instruction following, pushing models beyond simple look‑up. He notes that even multi‑needle tests are insufficient, and the community needs benchmarks that measure state management over long, evolving sessions.

[42:25](https://addtry.com/9adf13dc-0536-4319-a0fa-ee61b066d8e3?t=2545000)

## Key moments

- **[0:00] Intro & Gradient**
  - [1:40] Mark Huang's founding story: from quant finance to Box and Splunk, then launching Gradient to deliver AI's full enterprise value.
  - [4:17] Gradient is a full-stack AI platform enabling agentic workflows, says Mark Huang.
  - [4:55] Q: What is the minimum viable agent? Mark Huang defines it by improvement in probability of success.
  - [8:18] Gradient pivoted to handling out-of-domain data, building systems that learn with the user.
  - [9:14] Swyx: 'In AI we definitely care a lot more about out-of-domain generalization,' defining the difference between ML and AI.
- **[9:53] The 1M Bet**
  - [10:14] Gradient extended Llama3 to 1M context length, inspired by Gemini and compute from Crusoe.
  - [13:05] Q: What is Crusoe? Mark Huang explains it as an alternative GPU cloud provider.
  - [14:17] Mark Huang explains why models lack long context: quadratic scaling of self-attention and need for curriculum learning.
- **[14:32] Scaling Obstacles**
  - [17:45] Q: Is there a minimum context size for extension? Mark Huang says no minimum but model needs good perplexity.
  - [18:58] Llama3's theta parameter governs RoPE scaling; increasing theta enables longer context via positional interpolation.
  - [19:30] Q: What is the theta value? Mark Huang explains it controls rotational embeddings, with interpolation over extrapolation.
  - [22:23] Mark Huang: 'It's not a mathematical tautology or proof; it's an empirical formula that actually worked really well' for theta scaling.
- **[22:42] Attention Flavors**
  - [22:43] RoPE scaling was chosen over ALiBi and YaARN for its 'empirical elegance,' while Ring Attention was used for training.
  - [25:25] Zhang Peiyuan's EasyContext was the only PyTorch Ring Attention implementation that worked for Gradient.
- **[28:01] Data Engineering**
  - [28:01] Gradient trained the extended Llama3 on SlimPajamas and filtered UltraChat, emphasizing data diversity to avoid shortcutting.
  - [31:07] Q: Can you put out-of-distribution data in context extension? Mark Huang says it's a drop in the bucket vs trillions of tokens.
  - [33:29] CodeLlama lost language capabilities after overfitting to code, warns Mark Huang, illustrating the risk of full fine-tuning on narrow data.
  - [34:14] Swyx vs Mark Huang: Can smarter loss functions solve overfitting to recent data? Mark Huang says it's 'provably solvable' but hard.
- **[34:27] Synthetic Data**
  - [36:18] Gradient used GPT-4 to rephrase chat data and generate synthetic tokens, with data pipeline being a key moat.
- **[38:22] LoRa & Merging**
  - [41:23] Mark Huang: model merging with LoRAs is 'like a singular value decomposition on top of weights' to get the most important ones.
- **[42:25] Benchmarking**
  - [42:53] Mark Huang calls needle-in-a-haystack a 'primitive you have to pass' but says the community must move to harder benchmarks like Ruler.
  - [44:00] Ruler suite tests multi-needle retrieval, variable tracking, and common word counting, pushing models to understand full context.
  - [47:20] Gradient's 4M context extension saw quality degradation from floating-point precision and overly large theta values.
  - [48:54] Q: Do people care about context beyond 1M tokens? Mark Huang says use cases are still being explored, maybe 10M or 100M needed.
  - [49:40] Long context enables shoving entire code repositories into models and managing evolving state in sessions, says Mark Huang.
- **[50:08] Use Cases**
  - [51:54] Q: How does long context impact chat vs document summarization? Mark Huang says documents can use RAG, but session state is harder.
  - [54:33] Mark Huang highlights a CMU paper showing many-shot in-context learning reduces sensitivity to example selection.
  - [55:58] Mark Huang sees healthcare and finance as key long-context domains, especially for grounding on medical images and filings.
- **[56:25] Multimodality**
  - [57:45] Mark Huang predicts: 'Multimodality is going to be pivotal for long context,' with early fusion like Chameleon leading.
- **[59:38] Staying Current**
  - [1:00:59] Mark Huang's routine for staying updated: AI News newsletter, Twitter, Discord, and trying products.
  - [1:05:11] Mark Huang: 'If you do not try out the tooling… you are missing out on someone's compression algorithm.'
  - [1:05:56] Q: What is a good perplexity score for context extension? Mark Huang says around 4, with quick drop indicating correct theta.
  - [1:06:20] Mark Huang's tip: if perplexity oscillates early in training after theta scaling, cut training and retry a new theta.
  - [1:10:29] Mark Huang calls for collaboration on long context evaluations and multimodal dataset construction for grounding.
- **[1:10:37] Call to Action**

## Speakers

- **Alessio** (host)
- **Swyx** (host)
- **Mark Huang** (guest)

## Topics

Language Models, Compute, Benchmarks

## Mentioned

Crusoe (company), Google (company), Gradient (company), Meta (company), Mistral (company), Mosaic (company), OpenAI (company), AlphaCode (product), Chameleon (product), CodeLlama (product), EasyContext (product), GPT (product), Gemini (product), Llama (product), MPT (product), Ruler (product), Slim Pyjamas (product), UltraChat (product), Yi (product), Zero Scrolls (product)

## Transcript

### Intro & Gradient

**Alessio** [0:01]
Hey, everyone. Welcome to the Leading Space Podcast. This is Alessio, partner and CTO in residence at Decibel Partners, and I'm joined by my co-host Swyx, founder of Small AI.

**Swyx** [0:11]
Hey, and today we're in the remote studio with Mark Huang from Gradient. Welcome, Mark.

**Mark Huang** [0:16]
Hey, glad to be here. It's really, uh, you know, a great experience to be able to talk with you all. I know, I know your podcast is really, really interesting, and I always am listening to it every time you guys have a release.

**Alessio** [0:29]
He's not a paid actor. He, he said that out of his own will.

**Swyx** [0:34]
We, we will give you the check later. Um, so-

**Mark Huang** [0:36]
Yeah

**Swyx** [0:37]
... uh, Mark, you're unusual in the sense that you and I go back to college. Uh, I don't exactly remember where we overlapped, um, but, uh, you know, we both went to, uh, to Wharton, um, and went into the sort of quantitative developer realm.

**Mark Huang** [0:50]
Yeah, exactly. Um, kind of crazy, right? So all goes full circle. Uh, I was a quant for quite a few years and then made it out into, uh, Silicon Valley. And, um, now we, we intersect again when it kind of feels like more or less the same, right?

Like the AI wars, the, the trading wars back in the day too, to a certain extent, and the grab for talent.

**Swyx** [1:14]
Yeah. There's, there-- I think there's definitely a few of us, uh, ex finance people moving into tech and then finding ourselves, uh, gravitating towards data and AI. Uh, it seems like you did that. You were, uh, uh, you, you were at a bunch of sort of quant trading shops, but then as you moved to tech, you were lead data scientist at Box and staff ML scientist at Splunk.

Uh, and then before working on, uh, the startup that eventually became Gradient. Uh, you wanna tell that story?

**Mark Huang** [1:40]
Yeah. I think, um, part of the reason why I, I, I came over from the quant finance world is to get more collaboration, learn about, um, what big data and, and scaling, uh, machine learning really looks like, uh, when you're not in this, like, this bubble, right?

Uh, and, um, working at Box, I, I worked mostly in a cross-functional role, um, helping, you know, product analytics and go to market. And then at Splunk, it was a lot more of, um, specific, uh, role where I was helping with streaming analytics and in, uh, search and in deep learning.

And, uh, for Gradient, like, really why we started it was whether it was in finance or whether it was in tech, I always noticed that there was a little bit more to give in terms of what AI or ML could, um, contribute to the business.

And, uh, we, we came at a really good time with respect to wanting to, um, bring, like, the full value of what that could be into the, into the enterprise. And then obviously, OpenAI created this huge vacuum into the industry to allow for that, right?

So, um, I myself felt like really, really empowered to, to actually ship a product and ship stuff that I could think could really help people.

**Alessio** [3:04]
Maybe just to touch a little bit on Gradient. I, I know we have a lot of things to go through, Gradient, Llama3, uh, context extension. There, there's a lot, but what exactly is Gradient? And you have an awesome, uh, design on your website.

It's like really retro and, uh, I think people that are watching Fall Out on Amazon Prime right now, uh, can maybe feel nostalgia just, just looking at it. Um-

**Mark Huang** [3:25]
Yeah

**Alessio** [3:26]
... w-what, what exactly is it? Because I know you have the foundry, you have the agent SDK. There's like a lot of, a lot of pieces into it.

**Mark Huang** [3:32]
Yeah, for sure, and, and appreciate the, the call out for the, for the design. I know my co-founder, Chris, uh, spent a lot of thought in terms of how he wanted the aesthetic to look like, and, um, it reminds me a lot about Mad Men.

So, um, that's how I, I-- that was the in-initial emotional shape that I felt when I saw it. Uh, well, quite simply, like Gradient, we're a full stack, uh, AI platform, and what we really wanna do is we wanna enable, um, all of the, you know, RPA workloads or the, uh, codified, uh, automation workloads that existed in enterprise before.

Um, we really wanna enable f-people to transition into more autonomous agentic workflows that are less brittle, um, feel more seamless as an interface too, so and, uh, able to empower, like, what we really think l- uh, the new AI, um, uh, workforce should look like.

And, um, you know, that, that kind of required us to build a fairly horizontal platform for those purposes.

**Alessio** [4:35]
We had this discussion in our AI In Action club on Discord, like the minimum viable agent or like kinda how you define an agent. Uh, w- yeah, what's the-- in your mind, what is like the minimum thing that you can call actually an agent and not just like- ...

a for loop, you know? And, uh, how, how do you see the evolution over time, especially as people adopt it more and more?

**Mark Huang** [4:57]
Yeah. I-- so I kind of, uh, stage it where, um, everybody, first of all, at the, at the lowest level thinks about like nondeterminism with respect to how the pipeline looks like when it's executed. But even beyond that, um, this goes back into effectively evaluations.

It's like on each stage of the node, you're gonna have to see a marginal improvement in the probability of success for that particular workload because of, uh, nondeterminism. So, um, yeah, I think it is an overloaded term to a certain extent 'cause like everything is an agent if it calls a, a language model, uh, or any sort of multimodal model these days.

But, um, for us, it's like, you know, my background is statistics, so I wanna see like improvements in the probability of the, the success event or outcome happening because of more nodes.

**Swyx** [5:49]
Yeah. I think, you know, the one thing that, uh, makes this sort of generative AI era very different from the sort of data science-y type era is that it is very nondeterministic and it's, it's hard to control. Um, yeah, I mean, so like, you know, I, I think what's the- Uh, founding story of Gradient, like, uh, how, you know, of, uh, of, of all the problems that you chose, like, uh, why choose this one?

Um, you know, uh, how did you get together your co-founders? Anything like that, that, uh, that, like, bring us up to the present day.

**Mark Huang** [6:21]
One of my co-founders is Chris, and he's, he's a really good friend of mine as well. I don't know if you intersected with him at Penn as well, but-

**Swyx** [6:26]
Chris Cheng

**Mark Huang** [6:27]
... um, yeah, yeah. Ch-Chris Cheng, who-- he wa- he was at Penn as well, di-did banking for maybe one or two years and then, um, you know, was a, a software engineer at Meta, um, also was at Google and then most recently he was, like, a director at Netflix in product.

And, uh, we always wanted to do something together, but we felt the, you know, what really came to fruition was wanting to develop something that is enterprise-facing for once, um, mostly because of our experience with internal tooling and i-in-in-in-inability for something to, like, uh, basically, um, exist through, like, a migration, right?

Like, all, all the time with every ML platform that I've ever had to experience or he had to experience, it's like a rebuild and you rip it out and you have a new workflow or automation come in, and it's this huge multi-quarter, maybe even multi-year, uh, project to do that.

And, uh, we also teamed up with, um, a former coworker of Chris's from Opendoor, Forrest, uh, who's also, um, on Google, uh, Cloud Platform and, um, you know, him seeing the scale, uh, and actually the, uh, the state-of-the-art in terms of Google was using AI for systems before everybody else too, right?

They invented the transformer and their internal set of tooling was just so far superior to everything else. Like, it's really hard to-- for people to go back after seeing that. So what we really wanted was to, um, reduce that, that friction for, like, actually shipping, um, uh, workloads in, in, in, in product value when you have all these, like, uh, types of operational frictions that happen inside of, uh, these large enterprises.

And then, um, really, like, the main pivot point for all of it was, like you said, things that can ha- uh, handle out-of-domain problems. So, like, out-of-domain data that comes in, having the flexibility to, uh, not fall over and, um, having something that, uh, you build over time that continues to improve.

Like, machine learning is about learning, and I feel like a lot of systems back into place, they weren't-- they were learning a very specific objective function, but they weren't really natively learning with the user. So, like, that's the whole, you know, we use the term assistant all the time, but, you know, my, my vision for the assistant was always for the system to, to grow alongside me, right?

Like, almost like an embodied, you know, second, uh, uh, a limb or something that will be able to get better as you, you also learn yourself.

**Swyx** [9:20]
Yeah. I, I might maybe call it, um, you know, people are always trying to define a different-- difference between ML and AI and, uh, I think in AI we definitely care a lot more about out-of-domain generalization. Um, and that's all under the umbrella of learning, but, uh, it is a very specific kind of learning.

Uh, I'm gonna try to make a seg-segue into, you know, today's, like, main topic of con-conversation that's something that you've been blowing up on, which is the long-context learning, right? Which is also f- some form of out-of-topic-- out-of-distribution generalization.

And, and in, and in this context, uh, you're extending the context window of an existing open source model. Um, maybe if you want to, like, just bring us all the way back to, uh, towards, like, what got you interested in long context, um, why did you, you, you find it, like, an interesting investment to work on and then, you know, the, the story of how you did your first, uh, extensions.

### The 1M Bet

**Mark Huang** [10:14]
Yeah. I think, um, it came, um, th- for Llama3 it's specifically, uh, we chose that model because of the main, um, criticisms about it before when, when it first got released. Uh, eight thousand context lengths just seemed like it was too short 'cause it seemed like Mistral and even Yi came out with, like, a two thousand token, uh, uh, model-- context length, uh, model.

Um, but the, the really-- the inception of all of it was, um, us, like, s- fine-tuning so many models and working on RAG so much and having this, and it still exists today, this basically pedi-pedagogical debate with everybody who's like, "Hey, is it fine-tuning versus RAG?

Is it this versus that?" And, like, at the end of the day, it's, it's just all meta-learning, right? Like, all we want is, like, the best meta-learning workflow or meta-learning setup possible to be able to, um, adapt a model to do anything.

So wh- naturally, long context had a place in that, but nobody had really, um, pushed the limits of it, right? Like, you would see, like, ten-shot, maybe one hundred-shot prompting, uh, uh, for, uh, improving the model's capabilities. But it wasn't until Google comes out with Gemini with the w- first one million context length model that a lot of people-- a lot of people's jaws dropped and, um, that hunger for, uh, understanding what that could really facilitate in the new workflows, uh, came about.

So, um, we were staged to actually train, um, uh, other open source models to do that. But the moment Llama3 came out, we just went ham against that specific model 'cause, uh, the two things that, um, were particularly appealing for that was the fact that, like, I see a lot of these language models as compression algorithms to a certain extent, like the way we have, like, fifteen trillion tokens into a specific model.

That definitely made me feel like it would have, um, a lot of capabilities and- Be more, uh, adaptable towards extending that context length. So we went in there and, um, you know, the, the one million number was always, uh, that was more of just, like, put the North Star up there and see if we can get there.

Um, and then see what was happening along the way as we did that. So, um, yeah, uh, also shout out to Crusoe who facilitated all that compute, because I would be lying if I was to say, like, anyone could just go out and do it.

It, it does require quite a bit of compute. It requires, like, a lot of preparation, but it just, like, all the, all the, uh, stars kinda aligned for that moment for us to, to go after that problem.

**Alessio** [13:05]
I'll take a side, uh, note on Crusoe since you just brought it up. Um, yeah, like, can you explain what Crusoe is? Uh, you know, I, I have this mental image of putting GPUs on top of oil rigs.

Um- So what, what i- what is it? What do they do? Uh, how do you work with them? You know, just any- anything, anything nice. I'm sure they appreciate nice things that you say about them too.

**Mark Huang** [13:26]
Oh, for sure. For sure. Um, so, uh, they, they came to us in, in, through a collaborative effort where we basically were in search of, um, a cloud, you know, a, a GPU provider. I don't wanna call cloud service provider quite yet, because then, you know, you think about the hyperscalers.

But for them, uh, you know, they're one of the biggest alternative GPU, uh, cloud providers. And, um, they were offering up, like, we wanna do a collaboration to showcase, uh, their technology, and, uh, it r- it just made it really easy for us to, like, scale up with their L40Ss, and those were the specific, uh, GPU instances we used.

And, um, uh, coordinating that effort with them to get, you know, that dedicated cluster, uh, first to do the project, um, it, it became a really good relationship and we still work with them today, 'cause, like, we're trying to evaluate more of these models and possibly train more of them.

And anyone could go up to them and, and basically, um, get your compute from them, and they, they have a lot of, uh, uh, GPUs available for the- for those type of projects.

### Scaling Obstacles

**Alessio** [14:35]
I would love to maybe have you run people through why the models don't come with longer context sequences out of the box. Like, obviously, um, you know, the TLDR is, like, self-attention is, like, quadratic scaling of memory, so the longer the, the context size, the more compute you have to spend, the training time, uh, and that's why you have to get Crusoe to, to help you, uh, expand it.

How do you actually train a large language model that is, like, a very long context, and then how does that differ from just tacking it on, on top later? Uh, and then maybe we'll dive into performance and some of those things.

But, uh, I, I think for a lot of folks in our audience that are more AI engineers, they use models, but don't necessarily build the models themselves, a lot of time it's hard to understand what goes into actually making a, a long context model.

**Mark Huang** [15:20]
Yeah. In, in terms of, you know, all the literature out there, I would say, um, honestly, it's, it's probably still TBD as to, like, the trade-offs between the approach we did, which is more of a curriculum learning approach, uh, after the fact, versus inherently training a model with a long context throughout, because I just don't think people have looked at the scaling properties of it in de- deep detail.

But as stylistic facts exist out there with res- research papers, from Meta themselves actually, um, they've al- already shown in a paper that, uh, if you train a model on a shorter context and you progressively increase that context to, like, you know, the final limit that you have, like 32K is usually the limit of Llama 2 was, was that long, um, it actually performs better than if you try to train 32K the whole time.

And, uh, I like to think about it intuitively as if you're trying to learn probability theory, you're not gonna go and read the book cover to cover and then do all the exercises afterwards. What you're gonna do is you're gonna do each chapter, do an exercise, read the chapter, do an exercise, and then finish, right, with the final set of, like, holistic exercises or examination.

So attention is exactly what it sounds like to a certain extent. Like, you have a bunch of indices, and you are making the model attend to localized context and concepts, uh, across the entirety of its, um, uh, uh, encoding, right?

Like, whatever the text that-- the sequence that you're giving it. So, um, when you're doing the curriculum learning aspect of things, you are kind of trying to give it the opportunity to also attend to all the concepts. So data actually, and the curation and the creation of that context, plays a huge role 'cause, um, a lot of times people make the mistake of trying to extend the context length by giving, just giving it raw text that, uh, doesn't have, uh, the necessity for the model to go all the way in the beginning of the sequence and then connect an idea to the end of the sequence.

**Alessio** [17:45]
So data, data quality is one thing, uh, but it sounds like as long as the base model is at least... I- I-- Would it still work, like, the one million context, if Llama 3 was 2K context size? Like, is there, like, a minimum context size that you need to then be able to generalize or does it not, not really matter and fine-tuning kind of takes care of it?

**Mark Huang** [18:06]
There's no minimum, I would say, uh, or at least I, I can't make such a strong statement as to say that that does not exist. But if you have a 4K, any, any regular model out there, like, if you, you can, you can progressively increase the, uh, context length of it, so long as it has shown The really good perplexity scores, uh, prior to your context length extension.

So if it hasn't shown, uh, good perplexity, you-- it basically can't even predict the next token. You're kind of, like, out of luck, right? Um, but then from there, the other component that we actually just released a blog on maybe last Friday, is, like, you gotta pay attention to that, um, to the theta value that the model starts off with.

What was fairly unique about the Llama 3 model was the sh- their choice of the theta parameter, which gave some suspicion as to how long the context could be extended, uh, for the model. So, um, that aspect of, like, we can go into, you know, a huge, uh, uh, uh, lesson in terms of positional encodings and, and, uh, RoPE scaling and stuff, but, um, the, that-- those concepts and, uh, the-- that aspect of things, uh, enables you to, to scale out the, the length much more easily.

**Swyx** [19:30]
What's the, um, tl;dr of what the theta is for a model? If I, if I haven't built a model before, what-

**Mark Huang** [19:37]
Yeah, yeah.

**Swyx** [19:38]
Not, not me, obviously, I know what it is, but for, for people that don't know, right?

**Mark Huang** [19:43]
Right.

**Swyx** [19:43]
I- I'm totally an expert.

**Mark Huang** [19:45]
Yeah. Um, well, so not all models have it, but, you know, so some models will employ, uh, uh, RoPE, um, scaling and, uh, specifically Llama 3 does that. But there's also other positional encoding and embedding mechanisms that other models employ.

But tl;dr is, if you think about, um, most architectures, they employ, uh, basically like a-- it's kind of like a sine or cosine curve. And you're thinking about, like, the different, uh, you know, you've the amplitudes that occur there to allow for, uh, the model to, like, see different types of, um, distributions of data.

Um, really what the, the theta value does, it's-- it, it governs, like, how often, like, a, a pattern's s- uh, going to appear in the embedding space. So, like, you, uh, basically are able to make-- uh, sh-shift that, um, rotational, uh, curve, uh, by increasing the theta value and allow for, uh, different types of, um, distributions to be seen as if they had actually occurred, uh, in the training data before.

So there's, uh... It's super confusing, but it's like there's extrapola- uh, positional extrapolation, and then there's interpolation. You want interpolation. It's been shown that just pure extrapolation makes the model a lot worse, and it's harder to attend to stuff.

Whereas the interpolation is like you're squeezing everything back in to what the original context length was to a certain extent, and then allowing for it to, um, overlap different sequences that it's already seen, as if it actually occurred when you see a billion contexts of, of, uh, of sequence tokens.

So, um, yeah, I, I think, uh, that aspect, like, we didn't know how well it would scale. I think that's one thing. So, like, um, I'm not gonna lie and tell you, like, right off the bat, we're like, "We're definitely gonna hit a billion."

It was more like we're getting to two fifty-six, and it looked good. We did our evals, we scaled it more, and then, uh, what was really good was that we established the formula, uh, at the start. So, like, it's actually a formula that we actually took, um, from, uh, the paper.

Uh, I think it's, uh, I think it's the r- the RoPE scaling paper. And we looked at that particular formula, and then we backed out the values. And they're-- it's all empirical, so, like, there's no... It's not like a mathematical tautology or proof.

It's just like, it's an empirical formula that actually worked really well, and then we just kept scaling it up and it, and it held. It's kind of like the scaling laws, you know. You-- nobody knows, like, the scaling laws exist, but you don't know if they're gonna continue.

So...

### Attention Flavors

**Swyx** [22:43]
Yeah, like, um, are you able to compare it with, like, other forms of scaling that, uh, people have been talking about? Um, ALiBi comes to mind. YaARN, uh, is being talked about a lot by News Research. Um, and then there's, there's other forms which are, like, not exactly directly related, but, like, Ring Attention comes up a lot.

Uh, we had a, we had a really good session with Strong Compute in our-- in the Latent Space Discord talking about s- uh, all these approaches. Um, I was just wondering if you want to compare and contrast, like, RoPE versus the other stuff.

**Mark Huang** [23:11]
Yeah, I think, um, AL- uh, I can never pronounce it right, but A-ALiBi, ALiBi.

**Swyx** [23:17]
ALiBi? Yeah.

**Mark Huang** [23:18]
Uh, yeah, ALiBi. Um, uh, we, we haven't compared with that one specifically, mostly 'cause the... I've, I've noticed the, um, some of the newer architectures don't actually em-employ it a lot. I think the last architecture that actually really employed it was the Mosaic MPT model class, and then almost all the models these days are all, uh, RoPE scaling, a-and then effectively, you can use YaARN with that as well.

Um, we just did the theta scaling specifically because of its, like, empirical elegance. Uh, it was really easy and, like, it, it was well understood by us. Um, the other one that I know that in the open source that people are applying, which uses more of a LoRa-based approach, um, which is really interesting too, is the one that Wing has been, uh, employing, which is Pose.

Um, you know, we, we've, we sort of helped them evaluate some of the models. Uh, with respect to, like, the performance of it, it, it does start to break down a little bit more on the longer, longer context, so, like, five hundred thousand to a million.

Um, it appeared that, uh, it, it, it doesn't hold as well specifically for, like, needle in the haystack. But, um, you know, it's still TBD as, like- You know, evaluations, I call it just like a high dim-- it's a sparse, high-dimensional space where you're just, like, evaluating, uh, performance across so many different things and then trying to map it back to, like, "Hey, here's the thing that I actually cared about from the start, and I have, like, a thousand different evaluations, and they tell me something, but not the entire picture," right?

Um, and as for, like, Ring Attention specifically, like, we, we employed Ring Attention in order to do the training. So we combined Flash Attention and Ring Attention together, uh, with, like, a really specific, uh, network topology on our GPUs to be able to maximize the, the memory bandwidth.

**Swyx** [25:07]
Yeah. Uh, yeah, as far as I understand, like, Ring Attention, uh, you know, a lot of people credit it for Gemini's, um, uh, million token context, but actually, it's just a better utilization of GPUs, right? Like-

**Mark Huang** [25:17]
Yeah

**Swyx** [25:17]
... that's, that's really, that's really what it is. Um, you mentioned in our show notes, um, Zhang Peiyuan's EasyContext repo. Uh, I have seen that come up quite a bit. Um, what does that do as a-- you know, like, how important is it as, as a Ring Inten-Attention implementation?

I know there's, like, a, maybe an-an-an-another one that was done by Lucid Raines or one of the other-

**Mark Huang** [25:36]
Yeah

**Swyx** [25:37]
... sort of open source people. Um, but, like, what, what is EasyContext? Like, is, is that, is that the, the place to go? Like, did you evaluate a bunch of things to implement Ring Attention?

**Mark Huang** [25:46]
Yeah. I-- w-we evaluated all of them. Like, we-- it was, uh, um, uh, I-I would say the original authors, you know, Matei and all the folks at, uh, uh, Berkeley-

**Swyx** [26:00]
The other group. Yeah

**Mark Huang** [26:00]
... they created the JAX implementation for it. And, um, unfortunately, uh, not to discredit, like, you know, TPUs or whatever, like, the JAX implementation just does not work on, on GPUs very well. Like, uh, any naive setup that you do, like, it just won't run out of the box very easily.

And then, uh, unfortunately, that was probably the most mature repo with, uh, a lot more configurations to set up interesting network topologies for your cluster. Um, and then the other PyTorch implementations outside of EasyContext, they, um, they just didn't really work.

Um, maybe we weren't implementing one small aspect incorrectly, but, like, there was an active development on it at a certain point. Like, even Lucid Raines, I think, uh, he's interesting 'cause for once he was actually, like, he was, like, taking a job somewhere and then just stopped, you know, commit, doing commits, and his-- as we were working to try to find it, like, we never really wanna jump in on a repo where someone's, like, kind of actively committing breaking changes to it.

Otherwise, we have to, like, eat that repo ourselves. And, um, yeah, EasyContext was the first PyTorch implementation that applied it with native libraries that, uh, worked, um, pretty well, and then we adapted it ourselves, uh, in order to, um, configure it for our, our, our cluster network topology.

So, um, you know, shout out, uh, to Zhang Peiyuan for, um, his open source contributions. I think that, um, we look forward to possibly collaborating him and push that further in the future because I think more people, if they do wanna get started on it, I would recommend that to be the easiest way.

Like, I don't know how many people know JAX. Me personally, I don't really know it that well. So I'm more of a PyTorch guy. So, um, yeah, I, I think that it, he, he provides a really good introduction to be able to, um, try it out.

**Alessio** [28:01]
And so on one side you had the technical discovery. What about the actual customer interest, customers that you work with? I, I feel like sometimes the context size can be a bit of a marketing ploy. You know, people are like, "Oh yeah, well, no, one million, two million, three million, four million."

### Data Engineering

**Alessio** [28:17]
So that, that, that's kind of the algorithms side of it. How do you actually, you know, how do you power the, the training? But the other side is obviously the, the data that goes into it. There's both quantity and quality.

Uh, I think I saw one of your tweets, uh, you trained on about two hundred million tokens for the AP model to the context extension. But what's ac- uh, what are the tokens? You know, how do you build them?

Um, what are like maybe some of the differences between pre-training datasets and context extension datasets? Um, yeah, any other color you could give there would be, would be great.

**Mark Huang** [28:49]
So specifically for us, um, we actually staged two different updates to the model. Um, so our, uh, initial layer that we trained was, um, just a basically like a pre-training layer, so continual pre-training where we took the Slim Pyjamas data, and then we filtered it, uh, and, um, concatenated it so that it would reach the context lengths that we were trying to extend out to.

And then we took the UltraChat dataset, filtered it down, or maybe some other, um, you know, second order, uh, derivative of the UltraChat dataset that was curated and, and then filtered it down, and then reformatted it for our, uh, chat use case.

Um, for those two datasets, I think, uh, you always have to really keep in mind, uh, for the pre-training data, um, whether or not, um, you may be, like, cutting off tokens in weird ways, um, whether or not, um, you know, the content is actually, uh, diverse enough to retain the ability of the model.

So Slim Pyjamas tends to be one of the best ones, mostly because it's a diverse dataset, and, um, you can use, uh, embeddings too as a, as a pre-filtering step as well, right? Like, how diverse are your embedding space ver-- uh, to the original corpus of the model?

And then train on top of that to retain its abilities. And then finally, for the, the chat dataset, making sure that it's, um, attending to all the information That would be expected to really stretch its capabilities, 'cause you could, you could create like a long context, uh, dataset where, like, every single time, the last 200 tokens could answer the entire question, and that's never gonna make the model attend to anything.

So it's even something that we're doing right now is trying to think about, like, how do we actually improve these models, and how do you uplate the datasets such that it can expose, like, even more nuanced capabilities, um, that aren't easily measurable quite yet.

**Swyx** [31:07]
Is there a ratio between diversity of the dataset versus, um, diversity compared to what the model already knows? Like, does the model already need to understand a good part of the new, uh, like the context exten-extension data to, to function?

Like, can you put a context extension dataset that is, like, very far from, like, what was in the pre-training? I'm just thinking as, as the model get older, you know, s-some of the, the datasets that we have might not be in the knowledge of the existing model that you're trying to extend.

**Mark Huang** [31:38]
I think that's always a consideration. I think specifically, um, you really gotta know how much-- how many tokens were expended into that particular model from the start. Um, and all models these days are s- now double-digit trillions, right?

So it's kind of a drop in the bucket if you really think, "I can just put, you know, a billion tokens in there, and I actually think that the model's gonna truly learn new information." Um, there is a lot of research out there between the differences with respect to, like, full fine-tuning, which we applied full fine-tuning versus LoRa-based fine-tuning.

Um, it's a trade-off, and my opinion of it is actually that, uh, you can test certain capabilities, and you can, um, kind of inject, uh, new knowledge into the model. But to this day, I've not seen any research that does, like, a strong, well-scaled-out empirical study on how do you increase the model's ability to understand, like, these decision boundaries, uh, with, with a new novel data.

Most of it is taking, like, holding out a portion of the data, um, as, like, novel and then needing to recycle some of the old knowledge, so it just doesn't, it just doesn't forget and get worse at everything else, right?

Um, which, which was seen, like, we do have historical precedent where CodeLlama was, you know, trained further from the, the, the original CodeLlama was trained further from Llama 2, and it just lost all its language capabilities basi-basically, right?

So it's not, you know, I, I, I don't wanna call that project, like deem it as like, uh, a failure, but it wasn't like a really successful generalization exercise because, you know, these models are about, like, flexibility and being, like, generic to a certain extent.

**Swyx** [33:35]
So one thing I see in the recent papers that have been coming out, um, uh, is this, so this concept of multi-stage training of data. Um, and if you're doing full fine-tuning, maybe the, the, the move or the answer is don't train five hundred billion tokens on just code because then, yeah, it's gonna massively overfit to just code.

Uh, instead, like, maybe the move is to slowly change the mix over, uh, over the different phases, right? So in other words, you, like, you, you still need to mix in some of your original source dataset to make sure it doesn't deviate too much.

Um, I feel like that is a very crude solution. Like- ... maybe there's some smarter way to adjust, like, the loss function so that it doesn't, like, uh, deviate, uh, or overfit too much to more recent data. Um, it's-- it seems like it's a solvable thing, is what I'm saying.

Like, the, this, this overfitting to more recent data issue.

### Synthetic Data

**Mark Huang** [34:28]
Well, solvable is hard. I think, uh, provably solvable is always something that I know is extremely, um, difficult. Uh, but, uh, from a heuristical standpoint as well as, like, having, like, some sort of statistical, um, uh, efficiency on, like, how you can converge to, uh, the downstream tasks and improve the performance that way, um, in a targeted manner, I do think there are papers that, um, try to do that.

Like the DoReMi paper, um, I think it was released last year. It was really good about doing an empirical study on that. Um, I think the th- one thing that I-- uh, people struggle with though is the fact that they always try to do it on pretty naive tasks.

Like you, you target, like, a naive task, and then you, uh, create your data mixture, and you try to, um, show some sort of, uh, algorithm that can, that can, um, retain the performance, uh, for those, uh, downstream tasks.

But then what do we all care about are actually, like, really, really interesting, complex tasks, right? And we barely have good evaluations for those. Like if you, if you, uh, do a deep dive at the Gemini, uh, one five, uh, technical paper, which they just updated with, like...

It was a fantastic paper with new updates. Uh, if you look at all of their long context evaluations there, like, a lot of them are just not something that the open community can even do because they just hired, like, teachers to evaluate whether or not this model generated a huge lesson plan that is really coherent.

Or, like, you hire a bunch of subject matter experts or, you know, they taught, uh, the model how to do language translation for, uh, an extinct language where only two hundred people in the world know. It's like it's kinda hard for us to do that same study, right, as a, as an early-stage startup.

**Swyx** [36:28]
I mean, technically now you can use Gemini as a judge. Um, Gemini is touting a lot of their capabilities in low-resource languages. Uh, o-one more thing before on, on the sorta data topic. Um, did you have any Um, exploration of synthetic data at all.

Um, you know, use, use Mistral to rephrase some existing part of your dataset to generate more tokens, anything like that, or, or any other form of synthetic data that you, that you choose to mention. Um, I think you, you also mentioned the large world model paper, right?

So, um-

**Mark Huang** [36:56]
Yeah

**Swyx** [36:57]
... yeah, anything like that.

**Mark Huang** [36:58]
Yeah, yeah. So, um, yeah, we, we used, like, GPT-4 to, uh, rephrase, uh, certain aspects of, um, the chat data, um, reformatting it or, uh, kind of generating new types of, uh, tokens and language in, in, you know, types of data that the model could see.

Um, and, uh, also, like, trying to take the lower, uh, correlated instances of out-of-domain data in, in-- that we wanted to inject into the model too as well. So, um, I actually think a lot of the moat is in the, the data pipeline.

Um, you'd, you'd notice, like, most papers just don't really go into deep detail, uh, about the dataset creation because they probably know... I mean, there's, there's some aspects-

**Swyx** [37:51]
Yeah

**Mark Huang** [37:51]
... that are, like, uninteresting, right? Which is like, "We paid a bunch of people and, like, generated a lot of good data." But then the synthetic data-generating pipeline itself, that, that, you know, sometimes that could be, like, twenty-five percent or, or fifty percent of the entire dataset that you've been used to do pre-training.

**Swyx** [38:07]
Yeah, I think it's just for legal deniability rather than- than, uh, "No, it's just too boring," you know. "I'm not gonna say anything 'cause it's too boring." No, it's-

**Mark Huang** [38:15]
Yeah

**Swyx** [38:15]
... actually really interesting, but, uh, it, it, and, and in fact, it might be too interesting, so, uh, we're not gonna say anything about it.

**Mark Huang** [38:22]
Yeah.

### LoRa & Merging

**Alessio** [38:23]
One more question that I had was on LoRa and taking some of these capabilities out and bringing them to other model. Uh, you mentioned Wing's, uh, work. Um, he tweeted about, "We're gonna take this LoRa adapter for the Gradient One Million Context ext-extension, and you're gonna be able to apply that to, to other model."

Can you just generally e-explain to people how, uh, these things work with language models? I think people understand that with stable diffusion, you have these, like, LoRa patches for, like, different types of styles. Um, does that work similarly with LLMs, and is it about functionality?

Can you do LoRa patches with specific knowledge? Like, what's the state-of-the-art there?

**Mark Huang** [39:03]
Yeah, I think there's a huge, uh, kind of resurgence in what I would, uh, call, like, model alchemy to a certain extent 'cause you're, like, taking all of these, uh, LoRas and you're mixing them together and then, uh, like that's a lot of the, the model-merging stuff that, um, I think Charles Goddard does, um, uh, in, in a lot of others in the open community, right?

'Cause it, it, it's a really easy way, like you don't need training, and you can test and evaluate models and take the best skills and mix and match. I don't think there has been as much empirical study, like you're saying, for how it shows the same type of like-- it, it's not as interpretable as like stable diffusion to a certain extent 'cause, um, you know, even we have experimented with, with taking like deltas in the same methodology as Wing, where we'll take a delta of like an already trained model, try to see how that has, uh, created like, in a sense, an RoHF layer, right?

Taking the Llama Instruct layer, subtracting the base model from that, and then trying to apply that LoRa adapter to like another model and seeing what it does to it. It does seem to have an effect though. Like, I will not lie to say I'm really surprised how effective it is sometimes.

But I do notice that for more complex abilities, um, other than like more stylistic stuff, it do- it kind of falls through 'cause, uh, maybe it's requires a much deeper path in the neural network, right? Like all these things, these weights are just like huge trees of paths that, um, the interesting stuff is like the road, uh, you know, the road less traveled to a certain extent.

And when you're just like merging things brute force, uh, together that way, um, you don't quite know what you'll get out all the time. Like, there's a lot of other research that, you know, you have merge ties, and you have all these different types of techniques to effectively just apply like a singular value decomposition on top of weights and just get like the most important ones and prevent interference across, you know, all the other layers.

Um, but, uh, yeah, I, I, I think that that is, uh, extremely interesting from the, uh, developer community, and I, I wanna see more of it, uh, except it is to a certain extent, kind of polluting the leaderboards these days 'cause it's so targeted and like now you can, you can kind of game the, the metric by just finding all the best models and then just merging them together to do that.

Um, and I, I'll just add one last bit is basically, um, the most interesting part about all that actually to me is when people are trying to take the LoRas as a way of like, uh, short-circuiting the training process.

So they take the LoRas, they merge it in, and then they'll fine-tune afterwards. So like the fine-tuning and the reinitialization of a little bit of noise into, uh, uh, all of the, the new merged models provides like a kind of, uh, kind of a learning tactic for you to get to that, um, capability a little bit faster.

### Benchmarking

**Swyx** [42:25]
There's a lot there. I really like the comparison of, uh, ties merging to singular value d-decomposition. Um, that's, uh, that's something that I get... I, I looked at the paper and I, I, I don't really think I understood it on, on that high level until, until you just said it.

Very cool. Um, we have to move on to, to benchmarking. Um, uh, this is a very fun topic. Uh, needle in a haystack. What are your thoughts and feelings? And then we can discuss the other benchmarks first.

**Mark Huang** [42:51]
Ah.

**Swyx** [42:52]
Needle in a haystack-

**Mark Huang** [42:53]
You want to-

**Swyx** [42:53]
It's very, very popular

**Mark Huang** [42:54]
... put me on the spot with that one. Um, yeah, I think, I think needle in the haystack is definitely, like, the standard for presenting the work in a way that people can understand and also proving out. I, I would say, like, I view it as, like, a primitive that you have to pass in order to give the model any shot of doing something that combines both, like, a more ho-holistic language understanding and, like, instruction following, right?

Like, honestly, like, it's mostly about, um, if you, if you think about the practical applications of, like, long context and what people complain most about models when you stuff a lot of context into it, is either the language model just doesn't care about what you asked it to do, or it cannot differentiate, like, you know, context that you want it to use as a source to prevent hallucination versus, like, instructions.

Um, I think that, you know, when we were doing it, it was to make sure that we were on the right track. Uh, I think Greg did a really great job of creating a metric and a benchmark that everybody, uh, could understood.

It was intuitive. Even he says himself, we have to move past it. But, um, to that regard, um, it's a big reason why we, we, we did the evaluation on the Ruler suite of benchmarks, which were, uh, way harder.

Um, they actually include needle in the haystack within those benchmarks too. Um, and, uh, I would even argue is more comprehensive than the, uh, benchmark that, that Gemini, uh, released for their, like, multi-needle in the haystack.

**Swyx** [44:30]
Yeah. You mentioned, uh, quite a few. You mentioned Ruler, Lugo, Infinite Bench, Bamboo, Zero Scrolls. Um, like, y-y- do you wanna, do you wanna give us, like, uh, maybe two or three of, of those that, that you thought were particularly interesting or challenging and, uh, you know, what made them stand out for you?

**Mark Huang** [44:45]
There's just so many, and, uh, they're so nuanced- ... that I would say like, yeah, Zero Scrolls was the first one I'd ever heard of, uh, coming out last year, and it was just, like, the extent-- like, making...

It's, it's got-- It was more of, like, tracking, um, uh, variable over long context. Um, I'll, I'll go into Ruler 'cause that's the freshest in my mind, and, like, we're just scrutinizing it so much and running the evaluation in the previous two weeks.

But, like, Ruler has, um, basic-- uh, has four different types of evaluations. So the first one is exactly needle in the haystack, it's that you throw multiple needles. So you gotta retrieve, you know, multiple key value pairs. There's another one where that basically you need to differentiate-

**Swyx** [45:28]
Multi-key, multi-value, multi-query.

**Mark Huang** [45:31]
Yeah, yeah. Multi-value, multi-query. Um, those are-- that's the ablation.

**Swyx** [45:34]
Yeah.

**Mark Huang** [45:35]
Um, uh, there's also a, uh, a variable tracking one where you go, "Hey, if X equals this, Y equals this, you know, Y, uh, Y equals Z, like, what was-- what is this variable?" And you have to track it through all of that context.

And then finally, um, there's one that is, uh, more of like creating a summary statistic, so like the common words one, where you choose a word, uh, that goes across the entire context, and then you have to, like, count it.

So it-it's a lot more holistic and a little bit more difficult that way. Um, and then there's, you know, the, uh, a few other ones, uh, that escape me at this moment. But, um, yeah, it-- Ruler really pushes you...

If I think about the progression of the evaluations, it pushes it to, uh, start to force the model to actually understand, like, the totality of the context rather than, um... Right? Like, everybody argues to say like, "I, I, I'll-- Couldn't I just use, like, a retrieval to, like, just grab that variable rather than, like, pay ten dollars for one shot or something?"

Um-

**Swyx** [46:43]
Yeah

**Mark Huang** [46:43]
... although it's not as expensive-

**Swyx** [46:44]
You still have to pay what you need.

**Mark Huang** [46:45]
Yeah. Exactly, exactly. So being able to actually, like... I, I think the, the main thing that, like, I struggled with, with even some of our use cases, um, were, like, when the context is scattered across multiple documents, and you have, like, really delicate plumbing for the retrieval stuff.

But, um, it only works for that one, that really specific instance, right? And then you, you throw in other documents and you're like, "Oh, great, like, my retrieval doesn't grab the, the relevant context anymore." So, like, that's the dream, right, of getting one model, a, a model that can generalize really well that way.

**Swyx** [47:20]
Yeah, totally. Um, th-that, that-- And I think that probably is what Greg mentioned when, uh, saying that he has to move beyond the needle in the haystack. Um, you also mentioned, uh, so you, you extended from one million to four million token context recently, um, and you saw some degradation in, in the benchmarks too.

Like, uh, do you wanna discuss that?

**Mark Huang** [47:39]
So if you look at our theta value at that point, it's getting really big. So think about floating point precision and thinking about propagating, like, uh, a va- basically, now you're starting to run into problems where in a, a deep enough network and having to pr- um, you know, to do joint probabilities across, like, so many, uh, uh, tokens, um, you're hitting the kind of the, the upper bound on, um, accuracy there.

And, uh, there's probably some aspect of, um, kind of, uh, clamping down certain activations that we need to do within training. Uh, maybe it happens at inference time, uh, as well, uh, with respect to, like, the theta value that we use and, and how do we, uh, ensure that it doesn't just explode.

Like, um, if you've ever had to come across, like, the exploding gradients or the vanishing gradient problem, you will know what I'm talking about. Like, a lot of the empirical aspect of that and scaling up these, these things is, uh, is, is- You know, experimentation and figuring out, like, how do you que-- how do you kind of marshal these really complicated, uh, composite functions such that they don't just, like, do a divide over zero problem at one point.

**Alessio** [49:06]
Awesome. Um, the-- just to wrap on the, uh, on-- there's the evals and then there's what people care about, you know? But there's two things. Do you see people care about above one million? Because Gemini had the two million announcement, and I think people were like, "Okay, one million, two million, it's whatever."

**Mark Huang** [49:24]
Mm-hmm.

**Alessio** [49:24]
Like, do you think we need to get to ten million- -to get people to care about again?

**Mark Huang** [49:28]
Yeah.

**Alessio** [49:28]
Like, do we need to get to a hundred million? Um, yeah.

**Mark Huang** [49:32]
Um, I mean, that's-- it's, uh, that's an open question. I would certainly say a million seemed like the number that got people really excited for us, and then, you know, the four million is kind of like, okay, it's like that's seen as more...

But rather than like a breakthrough milestone, it's like just the next incremental, uh, uh, checkpoint. Um, I, I, I do think, like even Google themselves, they're evaluating and trying to figure out specifically how do you, how do you measure the, the quality of these models and how do you measure and, and, and map those to capabilities, uh, that you care f- that, that you care about going down the line, right?

### Use Cases

**Mark Huang** [50:19]
And, um, I, I, I think I'm still-- like us as a company, like we're figuring out how to saturate, um, the context window in a way that's like actually, uh, adding incremental value. So the obvious one is code, 'cause code repositories are huge.

So like, can you stuff the entire context of a repo into the co-- uh, into a model and then make it produce like some module that is useful or some suggestion that is useful. However, I would say like there are other techniques like, you know, AlphaCode and flow engineering that, um, if you do iterative things in a more agentic manner, uh, it may actually produce better quality.

I would preface and I would actually counter that maybe start off with, um, the use case that is a little bit more-- that people are more familiar with right now, which is constantly evolving context in like a session.

So like, whereas you're coding, right? If you can figure out evals that actually, um, work where you're constantly providing it multiple turns, and each incremental turn has a nuanced aspect, and you have a targeted, uh, uh, generation that you know of, making the model, um, track state and have state va- management over time is really, really hard.

And it's an incredibly hard evaluation that will probably only really work when you have a huge context. So, um, that's sort of what we're working on, trying to figure out those types of aspects. You can also map that, like it's not just code.

State management exists in like, you know, we work in the finance sector a lot, like investment management, like having management, uh, uh, a state management of like a concept and, and stuff that evolves over like a long session.

So, um, yeah, I, I, I'm super excited to hear like what other people think about the longer context. Um, I don't think Google's probably in-investing to try to get a billion quite s- quite yet.

**Alessio** [52:34]
Mm-hmm.

**Mark Huang** [52:34]
I think they're trying to figure out, um, how to fully leverage, uh, w-what they've done already.

**Alessio** [52:41]
Yeah. And does this change in your mind for very long chats versus a lot of documents? The, the chat is kinda interactive, you know, and the information changes. The documents, you're just trying to synthesize more and more things.

Um, yeah. Any thoughts on how those two workloads differ?

**Mark Huang** [52:58]
Yeah, I, I mean, I would say, like with the document aspect of things, um, you probably have like a, a little bit more ability to tweak, um, other methodologies. Like you can get around the long context, uh, sometimes where you can do retrieval augmented generation, or you do like, um, uh, hierarchical, like recursive summarization.

Um, whereas like evolution in like a session, because that state variable could undergo like pretty rapid changes, um, it's a little bit harder to, um, imagine like you getting around that without codifying like a really specific workflow or like some sort of, um, uh, you know, state clause that is going back to like determinism, right?

Um, and, and then finally, like what I really think, uh, people are trying to do is like figure out, um, how do all these sh- like shots, uh, progress over time. So like, how do you get away from the brittleness of like the retrieval step to like shoving-- if you shove in a thousand shots or two thousand shots, will it just make the retrieval aspect of good examples irrelevant?

And like it's sort of, kind of like a randomly sampling is fine at that point. There's a, there's a-- there's actually a paper on that, that, that came out, uh, from CMU-

**Alessio** [54:24]
Hmm

**Mark Huang** [54:24]
...that they, they, they showed, um, with respect to a few, uh, extraction or classification, high cardinality benchmarks. Um, they tracked like fine-tuning versus int- in-context learning, uh, versus like many, many shot in-context learning. And they basically showed that like many, many shot in-context learning, uh, helps to prevent as much sensitivity around the examples themselves.

Right? Like the distraction, the distraction error that a lot of LLMs get, where you give it irrelevant context and it literally can't-

**Swyx** [55:00]
Yeah

**Mark Huang** [55:00]
... do the task because it just is-- Like, it's sorta like a person too, right? Like, you gotta be very specific about, "I don't wanna distract this person because then, you know, they're, they're gonna go down a, a rabbit hole and not be able to complete the task."

**Swyx** [55:14]
Yeah. Well, that's kinda the flip side of the needle in a haystack thing too, in a bit. It's like n-now the models pay attention to, like, everything so well-

**Mark Huang** [55:22]
Mm-hmm

**Swyx** [55:22]
... that, like, sometimes, yeah, it's hard to get them to like... "I just said that once. Please do not bring that up again." You know?

**Mark Huang** [55:28]
Guilt-ridden.

**Swyx** [55:28]
Like, it happens to me with code. Yeah, yeah, it happens to me with, like, a, a CSS style sometimes and, like, things like that. If I have a long conversation, it's like it tries to always reapply certain styles, even though I, I told it maybe that's not the right, the right way to do it.

Um, but yeah. There, there's a lot, a-again, of, uh, empirical work that people will do. Um, and just-- I know we, we kinda went through a lot of the, the technical side, uh, but maybe the flip side is w-why is it worth doing, you know?

Like, what, what are, like, the use cases that, that people have, um, that make long context really useful? I know you had-- I, I think you have a lot of, uh, healthcare use cases. I saw on your Twitter.

You just mentioned the, the finance use case, obviously. Uh, some of the filings and documents that people, that companies publish can be quite worthy. Um, any other things that you wanna bring up, uh, maybe how people are using Gradient, anything like that, I think that will help, um, cl- uh, have a clearer picture for, for people.

### Multimodality

**Mark Huang** [56:25]
Yeah. Um, so beyond, like, just using the context for, uh, you know, sessions and evolving state management, it really comes down to something that's fairly obvious, which everybody's trying to do and work on, is how do you ground the, the language model better?

So I think when you think pure text, that's one thing. But then multimodality is, in my opinion, going to be-- it's going to be pivotal for long context. Just because, like, videos, uh, when you're getting into the frames per second, um, uh, and you're getting into lots of images and, like, things that are a lot more, like, embodied, you're-- you, you need to utilize and leverage way more, way more tokens.

And that is probably where, you know, us as a company, like, we're exploring more and trying to open up the doors for a lot more use cases, 'cause, um, I think in financial services, uh, as well as h-uh, healthcare, um, we've done a good job on the tech side, but we still need to push a little bit further when we combined, like, you know, a picture with words, like a chart with words or, um, somebody's, uh, medical image with, with words, stuff like that.

Like, you definitely can do a better job, um, and, you know, it's timely too, 'cause Meta just released their Chameleon paper, uh, the new Chameleon paper that does multimodal training, and it shows that early fusion helps you to-- It's, like, more sample efficient, right?

So having that kind of view towards the future is something that, um, you know, we wa- we, we wanna be primed to do because, you know, it's similar to what Sam Altman says himself too, right? Like, you need to just assume that these models are gonna be 10X better, um, in the next few years.

And if you are primed for that, like, that's where, uh, you have kind of a business that, you know, you're not just, uh, pivoting after every release or every, uh, uh, event, uh, you know, that drops.

**Swyx** [58:32]
Uh, I think the, the thing about this 10X issue is that, um, the, the 10X direction moves all the time, you know. Um, some people were complaining about GPT-4.0 that, um, yeah, look, like, the ELO scores for GPT-4.0 actually in reality weren't that much higher than GPT-4 Turbo, and really the-- You know, so it's not 10X better in reasoning, it's just 10X better in the integration of moda- uh, multi-modalities and...

Oh, by the way, look over here. There's a really sexy voice chat app that they accidentally made, that they had, they had to deprecate today. Um, i-it's like the 10X direction keeps moving. Now, now it's like, you know, fully in, like, sort of multimodality land, right?

And, like, um, the question is, like, what next, right? Like, so you, you can 10X in, in, in various ways, but, like, you, you guys have 10X context length. Um, but, like, you know, it-- Are we, are we chasing the last war?

'Cause, like, now, now, now, like, nobody cares about context length, and now it's, now it's, now it's, like, multimodality time, you know. I, I'm, I'm joking, obviously. People do care about it. I, I just, uh, I wonder about this, uh, how, uh, this, this, this comment about this 10X-ing every single time.

### Staying Current

**Mark Huang** [59:39]
You know, that's honestly why we kind of have our eye on the community- ... as well as you, right? Like, you, you, you, you, you know, with your community and, and the things that you hear, um, you know, you wanna build.

We're, you know, we're a product company. We're trying to build for users and, uh, trying to listen to, uh, understand, like, what they, what they actually need. Like, obviously, you know, you don't, you don't build everything that people ask you to build but have a-- know, know what's useful, right?

'Cause I think that, uh, you're, you're totally right there. Where, like, if we, if we want to make something, um, 10X better in a certain direction, but nobody cares, and it's not useful for somebody, then, um, it, it wasn't really worth the, worth the, the while.

And if anything, maybe that's, like, bitter, the bitter lesson 2.0 for so many tech startups, is, like, build technology that people care about and will actually 10X their value, rather than, like, build technology that's just, that's just 10X harder.

**Swyx** [1:00:38]
I mean, no, that's not, that's not bitter lesson. That's just Paul Graham. That's, that's-

**Mark Huang** [1:00:42]
Yeah

**Swyx** [1:00:43]
... what he want. That's... One more thing on the Chameleon paper. I was actually just about to bring that up, you know. So on AI News, like, my sort of daily newsletter, it was literally my most, my most recent featured, uh, paper.

And, uh, I always wonder if the-- you can actually sort of And, like, train images onto the same latent space as words. That was kind of done with, like, you know, what we now call late fusion models with, like, Llava and, um, Flamingo and, uh, you know, all the others.

Uh, but now the, the early fusion models like Chameleon seem to be the way forward. Um, like, obviously, it's more native. Um, I wonder if you guys can figure out some kind of weird technique where you can take an existing, like, Llama 3 model and, like, you know, early fuse the, the, the images into the, the, the text encoder, um, so that we, we just retroactively have the early fusion models.

**Mark Huang** [1:01:32]
Yeah. Um, even before the early, you know, that-- the Chameleon paper came out, I think that was on our big board of next to-dos to, to possibly explore, um, or our backlog of, of ideas, right? Uh, because, uh, as you said, early fusion, like, even before this paper, I, I can't remember, uh, I think Meta even had, like, a scaling laws for multimodality, uh, paper that does, uh, explore more early fusion.

Like, the moment we saw that, it was just kind of obvious to us that, um, eventually it'll, it'll get to the point that that becomes a little bit more mainstream. And, um, yeah, like, that's a, that's a cool twist that we're, we're, we've been thinking about too as well.

Uh, as well as, like, other things that are kind of in the works that are a little bit more agentic. But, um, yeah, if open collaboration interests you, we can always work on that together with the community.

**Swyx** [1:02:24]
Ooh, okay. Shout out there. Um, uh, cool. Uh, we-- you can leave that, uh, in the call to action at the end. Um, I just wanna, you know, we have a couple more questions to, to round this out.

Um, you mentioned a lot of papers in your work. Uh, you're also building a company. You're also looking at open source projects and community. Um, what is your daily or weekly routine to keep on top of AI?

**Mark Huang** [1:02:45]
Uh, um, so one, subscribe to AI News. He didn't have to pay me to say that. I actually really think, like, it's a good aggregator. I, I think it's a good aggregator. I'll tell you why. Most of the fastest-moving, like, um, research that's being done out there is, like, it's showing up-- it's mostly on Twitter.

Like, my Twitter's like... I, I wasn't a power Twitter user at all before three years ago, but I had to use it, and I had to always check it in order to keep on top of, like, early work, right, that people wanted to talk about or present.

'Cause nothing against, uh, submitting research papers to, like, ICLR or ICML. Like, know-knowing the state of the art, like, those are, like, six months, uh, uh, late, right? Like, people have already have it, dropped it on archive, or they're just openly talking about it.

**Swyx** [1:03:39]
The submission process. Yeah.

**Mark Huang** [1:03:41]
Yeah. And then being on Discord to see, uh, when the rubber hits the road, right? Like, the im-implementations and the practices, um, that are being done or, like, the, the datasets, like you said. Like, a lot of, um, conversations about really good datasets and how do you construct them are done, um, in the open and figuring that out.

For people that don't have, like, budgets of, like, ten million dollars to just pay a bunch of annotators. Um, so I, you know, my, my routine daily is, like, the second thing I do when I wake up is to look on Twitter, uh, to, to, to, uh, uh, see what the latest updates are from, uh, you know, specific people that do really, really great work.

Um, Armin, uh, at Meta, who did the Chameleon paper, is like, everything he writes on Twitter is like gold. So, like, anytime he writes something there, like, I really try to figure out what he's a-he's actually saying there and then tie it to, to techniques and research papers out there.

And then, um, you know, sometimes I, I try to use, uh, certain tools. Like, I myself use AI itself, uh, to, to search for the, the latest, uh, papers on a specific topic, if that's the thing on the top of my mind.

And, um, at the end of the day, uh, trying out the products too. I think if you do not try out the tooling and some of the, the products out there, like, you are missing out on, like, someone's compression algorithm.

Like, they compressed all the research out there and all the thought and all the state of the art into a product that they're trying to create for you. And then, like, really backing out and reverse engineering, like, what it took to build something like that, like, that's, you know, that's huge, right?

Like, if you can actually understand, like, perplexity, for instance, like that's... You'll, you'll-

**Swyx** [1:05:32]
Yeah

**Mark Huang** [1:05:32]
... already be well ahead on the research.

**Swyx** [1:05:34]
Oh, by the way, you mentioned, uh, what is a good perplexity score? I-I-if there's, like, just a number, right? Like, it's, like, five to eight or something. Like, what, what, what-- Do you have a, do you have a number in mind when you, when you said that?

**Mark Huang** [1:05:46]
Yeah. I mean, um, what was the one that we had? Like, flipping between train loss and perplexity is actually not native to me, uh, quite yet. But, like, yeah, between, like, if you can get, like, a four, uh, using the context length extension on, on, on, on Llama, like, you're in the right direction.

And then obviously you'll see spikes, um, and specifically when the one trick you, you should pay attention to is, um, you know that your, uh, context, uh, length and theta scaling is working right if the early steps in the perplexity go straight down.

So, like, when it wasn't correct, it would oscillate a lot in the beginning, and we just knew that we cut the training short and then retry a new theta scale.

**Swyx** [1:06:29]
W- 'Cause you're, you're-- In fact, you're properly continuing the, the, the fine-tuning or the, the full pre-training.

**Mark Huang** [1:06:34]
Yeah, yeah.

**Swyx** [1:06:34]
Um-

**Mark Huang** [1:06:35]
The model just, like, im-

**Swyx** [1:06:36]
Yeah

**Mark Huang** [1:06:36]
... it saw something out of domain immediately and was like, "I have no idea what to do."

**Swyx** [1:06:40]
Yeah.

**Mark Huang** [1:06:40]
And you needed to- ... to be able to overlap that, that, uh, positional-

**Swyx** [1:06:46]
Yeah

**Mark Huang** [1:06:46]
... uh, embedding on top of each other.

**Swyx** [1:06:47]
And one follow-up, right, uh, before, before we sort of close out. Um, like, y- I think being on Twitter and, like, looking at all these new, uh, h- new headlines is, is really helpful. But then, um, it only gets you, like, a very surface level understanding, and then you still need a process to decide which one to invest in.

Um, so I'm, I'm trying to dig for, like, what is your formula for, like, deciding, you know, what to go deep on and what to kinda skip?

**Mark Huang** [1:07:13]
From a practical standpoint, as a company Like I already know the-- there are like three to five things that will be valuable and useful to us, and then there's other stuff that's like out of scope m- from, from-- for different reasons.

Some stuff is like out of scope from, um, "Hey, this is not going to impact or help us," and then other things are out of scope because we can't do it. You know, like the, the sta-- Like different tech-- So a g- really good instance for that is, um, specific algorithms for, um, you know, improving extremely large-scale distributed training.

Like that's the-- We're not gonna have the opportunity to get two thousand H100s. If we do, it'd be really cool. But like I'm just saying like as for now, like you gotta, you gotta reach for the things that would be useful.

Things that would be useful for us, for instance, um, are-- for everybody actually, to be honest, is like, um, evaluations, uh, different post-training techniques, and then synthetic data, uh, uh, construction. Like we're always on the, uh-- I'm always on the look for that.

And then how do I figure out whether these things, um, you know, which new piece of news is actually novel? Um, well, that's sort of my like mental cache to a certain extent. Like I've built up like this state of like I already know like all the things that have already been written, uh, for the state-of-the-art, uh, for, for certain topic areas, and then I know what's being kind of recycled as like an empirical study versus like something that actually is very insightful.

Underrated, uh, uh, uh, specific instance would be like the DeepSeek paper. I'd never seen it before, but like the, uh, the multi-head latent attention, like that was really unexpected to me because like I thought I'd seen every ty-- not every type obviously, but like every way that people wanted to cut like mixture of experts into interesting ways, and I never thought something would like catch my eye to be like, "Oh, this is totally, um, new, and, and, and it really does have a lot of value."

Um, yeah, so like that-- I, I think that's, that's mainly, uh, uh, how I try to do it and like, um, you talk to your network too. Like I just, you know, talk to the people in the know and, and, and make sure like I have certain subject matter experts, uh, uh, on, on speed dial that, uh, I also like to share information with and, and, and, and understand like, "Hey, is this, um...

Does this catch your eye too? Do you think this is, uh, valuable or real?" 'Cause yeah, right, Sean, we-- it's a noisy space we're in right now, um, which is cool 'cause it's, it's really interesting and, uh, people are excited about it.

But at the same time, there, there is actually a 10X or more explosion of information coming in that all sounds really, really unique and new, and you could spend like hours, you know, down a rabbit hole that, that isn't as useful.

**Alessio** [1:10:24]
Awesome, Mark. I know we, we kept you in the studio for a long time. Uh, any final call to actions for folks? That could be roles you're hiring for, uh, requests for startups, a-anything that comes to mind that you wanna share with the, with the audience.

### Call to Action

**Mark Huang** [1:10:37]
Yeah. I think on the line of, um, we definitely have a call to action to get more people to work together with us for long context evaluations. Um, that is the sort of the it topic throughout like every-- like even Meta or Google or any of the other folk, uh, are, are focusing on, 'cause I think, um, we lack an understanding of that within the community.

And then, um, can we, uh, as a community also help to construct like other modalities of data sets that would be interesting? Um, like pairwise data sets, right? Like you could get just straight video and then straight text, but like getting them together that have like, uh, for, for grounding purposes will be really useful for training the nes- next set of models that I know are, are coming out and, um, uh, the more people we have contributing to that would, would be really useful.

**Alessio** [1:11:36]
Awesome. Thank you so much for coming on, Mark. Uh, this was a lot of fun.

**Mark Huang** [1:11:39]
Yeah. Thanks a lot.

**Swyx** [1:11:40]
Yeah. This was great.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
