LALatent SpaceAug 2, 2024· 1:23:36

The Winds of AI Winter (Q2 Four Wars of the AI Stack Recap)

Swyx and Alessio recap Q2 2024 through their 'Four Wars' framework, arguing that the AI landscape is shifting from frontier model dominance to commoditization and vertical applications. They highlight Claude 3.5 Sonnet overtaking OpenAI on coding benchmarks, Llama 3.1's synthetic data approach enabling 7B models to rival GPT-4, Mistral Large 2's non-commercial license and lost open-source crown, and on-device models like Gemini Nano and Apple Intelligence. The Quality Data Wars see NYT suing OpenAI, Reddit licensing data for $200M+, and synthetic data proving real for math (AlphaProof near IMO gold) and code. The Multimodality War includes ChatGPT Voice Mode delayed, Meta's Chameleon for native fusion, and Google's PaliGemma for PDF extraction. The renamed LLM OS War covers agent protocols, memory databases, and the collapse of model cost by an order of magnitude every four months, pushing startups toward vertical services like Brightwave and Dropzone that sell labor, not tools. The episode ends with a CrowdStrike joke about agent safety.

  1. 0:00Intro
  2. 5:05Claude 3.5
  3. 8:42Llama 3.1
  4. 15:22Small Models
  5. 20:06On-Device LLMs
  6. 33:48Quality Data Wars
  7. 44:10Multimodality War
  8. 51:36RAG Ops War
  9. 1:06:01Agent Ecosystem
  10. 1:09:34Commoditization
  11. 1:15:41Service as Software
  12. 1:19:54New Benchmarks

Powered by PodHood

Transcript

Intro0:00

Alessio0:04

Hey everyone, welcome to the Late in Space podcast. This is Alessio, partner and CTO in Residence at Decibel Partners, and today we're in the Singapore studio with Swyx.

Swyx0:14

Hey, uh, this is our long-awaited one-on-one episode. Uh, I don't know how long ago the previous one was. Do you remember? Three, four months?

Alessio0:23

No. Yeah, it's been a, it's been a while.

Swyx0:25

It's been a minute. And people really enjoyed it. It's, it's just that really I think our travel schedules have been really difficult to get this stuff together. And then we also had like a decent backlog of guests for a while.

I think we've kind of depleted that backlog now, and we need to build it up again. But it's, it's been busy and there's been a lot of news, so we actually get to do this like sort of rapid fire thing.

I think some people, you know, the podcast has grown a lot in the last, uh, six months. Maybe just in reintroducing like what you're up to, what I'm up to, and, um, why we're here in Singapore and stuff like that.

Alessio0:55

Yeah. My first time here in Singapore- ... which has been really nice. This country is really amazing, I would say. First of all, everything feels like the busiest part of the city. Everything is skyscrapers. There's like plants on all the buildings, or at least in the areas that I've been in, which has been awesome.

And I was at, uh, one of the offices kind of on the south side, and from the 30th floor you can see Indonesia on one side and you can see Malaysia on the other side.

Swyx1:19

Mm. Yeah.

Alessio1:20

Um, so it's, uh, quite, quite small. One of the people there said their kid goes to school at the border with Malaysia basically, so they could drive to Malaysia every day. Just to go pick her up from school.

Yeah, and we came here, we hosted, um, with you the Sovereign AI Summit Wednesday night. We had, uh, a lot of, a lot of folks-

Swyx1:36

Nvidia, Goldman, Temasek, uh-

Alessio1:39

GIC-

Swyx1:39

Singtel

Alessio1:39

... Singtel. And we're gonna talk about this trend of sovereign AI which maybe we might cover on another episode. But basically how do you drive, if you're a country, how do you drive productivity growth in a time where populations are shrinking, the workforce is shrinking, and AI can kind of supplement a lot of this?

And then the question is, okay, should I put all this money in foundation models? Should I put it in data centers and infrastructure? Should I put it in GPUs? Should I put it in agents and whatnot? So we'll touch on some of these trends in, in the episode, but it was a fun event, and I did not expect some of the most senior people at the largest financial institution in Singapore ask about state space models and some of the-

Swyx2:17

Yeah

Alessio2:17

... alternatives. So it's great to see, um, how advanced the conversation is sometimes.

Swyx2:22

Yeah. I think that that is mostly people trying to listen to jargon that is being floated around as like, oh, what could kill transformers? And then they jump straight there without actually exploring the fundamentals, the basics of what they'll actually put to work.

That's fine. It's a, it's a forum to ask questions, so it's-- you wanna f- ask about the future, but I feel like it's not very practical to spend so much time on, on those things. You know, part of the things that I do at Late in Space, especially, um, when I travel, is to try to ask questions about what countries that are not the US and not San Francisco can do because everyone feels a bit left out.

Y- y- you feel it here as well, and, uh, I'm trying to promote alternatives. I think AI engineering is one way that countries can, can capitalize on the industry without building a $100 billion cluster, which is one-fifth the GDP of Singapore.

Um, and, and, uh, a- and, and so, you know, I'm-- what my pitch at the, uh, summit was that, uh, we would, uh, Singapore would be the AI general nation. We're also working on bringing, um, the AI General Conference to Singapore next year together with Iclair.

So yeah, we're, we're, we're just trying my best and, you know, I'm, I'm being looped into various government meetings to, to try to make that happen.

Alessio3:33

We'll, we'll definitely be here next year. We'll be-- I, I'll be back here very often. It's, uh, it's really nice. So.

Swyx3:38

Yeah. Awesome. Okay. Well, well, we have a, you know, a lot of news. Uh, how do you think we should cover?

Alessio3:44

Maybe just recap since the framework of the Four Wars of AI is something that came up-

Swyx3:50

People keep coming back to it

Alessio3:50

... end of last year.

Swyx3:51

Yeah.

Alessio3:51

Um, so basically we'll link in the show notes, but the end of year recap for 2023 was basically the Four Wars of AI, uh, which we picked GPU Rich versus GPU Poor, the Data Quality Wars, the Multimodality Wars, and the RAG/Ops Wars.

So usually everything falls back under those four categories. So, uh, I'm pretty happy that seven months later, uh, it's something that still matters given that-

Swyx4:16

It still kind of holds up.

Alessio4:17

Yeah. Most, most AI stuff from eight months ago, it's really not that relevant anymore. Um, and today we'll, we'll try and bucket some of the, the recent news on it. We haven't done a monthly thing in like three months, so three months-

Swyx4:31

Yeah

Alessio4:31

... it's a lot of stuff.

Swyx4:32

That's mostly because I got busy with the conference.

Alessio4:35

Yeah.

Swyx4:37

Um, but I, I do want to-- I actually, I, I do want to get, get, get back on that horse or, or maybe just do it weekly so that I don't have such a big lift that I don't do it.

I think the activation energy is, is the problem really. Um, so yeah, uh, I think frontier model-wise, it seems like Claude, it has really carved out a posis- persistence space for itself. You know, for a long time, Anthropic was kind of like a clear number two to OpenAI, and with 3.5 Sonnet, i- at least in like the some of the hard benchmarks on LMSYS or coding benchmarks on LMSYS, it is the undisputed number one model in the world, even with 4o Mini.

Claude 3.55:05

Swyx5:10

Uh, and we can talk about 4o Mini and, and benchmarking later on. But for Claude to be there and hold that position for, um, what is almost-- what is more than a month now in, in AI time is, is a big deal.

There's not that much that people know publicly about what they-- what Anthropic did for Claude Sonnet, but I think it's, it's still a, a huge achievement. It, it marks the beginning of a non-OpenAI-centric world to the point where people on Twitter have canceled ChatGPT.

That's been a trend that's been going on for a while. We talked about the unbundling of ChatGPT. But now, like new open source projects and tooling, they're just built for Claude. Like, they don't even use OpenAI. That's a s- a strategic threat to OpenAI, I think, a little bit.

Uh, obviously, OpenAI is so big that it d- doesn't really care about that. But, uh, for Anthropic is a big win. I, I think like to, to see that going and to see Anthropic differentiating itself and actually implementing research.

Uh, so the rumor is that the scaling monosemanticity paper that they put out, um, two months ago Was a, a big part of Claude 3.5 Sonnet. I've had off the record chats with people about that idea, and they, they, they don't agree that it is the only cause.

Um, so I, I was, I was thinking like this is the only thing that they did. But, um, you know, people, people say that, uh, there's about four or five other tricks that they haven't disclosed yet that went into 3.5 Sonnet.

But the scaling on the semanticity paper is a very, very good read. It's a very long read, but it basically says that you can find control vectors of control features now that you can turn on to make it better at code without really retraining it.

You just train a whole bunch of sparse autoencoders, find a bunch of features, and just say, like, "Let's up those features," and suddenly you're better at code, or suddenly you care a lot about the Golden Gate Bridge. These are the same things to the model.

That is a huge, huge win for interpretability because up till now, we were only doing interpretability on toy models, like a few million parameters, a model of Go or chess or whatever. Claude 3 Sonnet was interpreted and usefully improved using this technique.

Wow.

Alessio7:06

Yeah. I, I think it would be amazing if we could replicate the same on the open models to then-- Because now we can use Llama 3.1 to generate synthetic data for training and fine-tuning. I think obviously Anthropic has a lot of compute and a lot of money, so it wants to figure out, "Okay, this is what we should make the model better at."

They can kind of like put a lot of resources. I think in open source, it's probably gonna be a more distributed effort, you know? Like, I feel like Nous has held the crown of like the best fine-tuning data set owners for a while, but at some point that should change, hopefully.

You know? Like other, other groups should, should step up, and I think if we can apply the same principles to like a model as big as 4 or 5B and bring them into like maybe the 7B form factor, that would be great.

But yeah, Claude is great. I canceled ChatGPT a while ago. Every small podcast I run for latent space, it runs both on Claude and on OpenAI, and Claude is definitely better most of the time. It's not a benchmark, it's just vibes.

But when the vibes are good, uh, the vibes are good, you know?

Swyx8:02

We run, uh, most of AI News summaries on Claude as well. Um, and but and I always run it against OpenAI. Sometimes OpenAI wins. Um, I, I do a daily comparison, but yeah, Claude, Claude's very strong at summarization and instruction following, which is something I care a lot about.

So when you talk about frontier models, MMLU no longer cut it, right? Like, uh, we have reached, reached like 92 on MMLU. It's going to like 95, 97. It just means you're memorizing MMLU. Like the it there, there's some fundamental re- irreducible level of mistakes because of M- uh, MMLU's quality.

Uh, we talked about this with Clementine on the Hugging Face episode. And so we, we need to, to see what, uh, what else. What is the next frontier? I think there are 10 directions that I outlined, uh, below, but we'll talk about that later.

Llama 3.18:42

Swyx8:42

Yeah. Should we move on to Llama 3?

Alessio8:44

Yeah. 3.1. I guess that too. Make sure to good differentiate between the, the models.

Swyx8:49

Yeah.

Alessio8:50

But yeah, we have a whole episode with, uh, Thomas Shalom from, um, the, the Meta team, which was really, really good. And I'm glad we got the podcast to come out at the same time as the model. Um-

Swyx8:59

Yeah, I think we're the only ones to coordinate for the paper release for the, for the big launch, the 4.5 launch. Zuck did a few interviews, but we're the only ones that did the technical team interview.

Alessio9:08

Yeah. Yeah, yeah. Zu- I mean, they were like surfing or something with the Bloomberg person. We should get invited to surf with Zuck, but I think the-

Swyx9:16

I would, I, yeah, I would be down to-

Alessio9:17

To the, to the audience the, the technical breakdown was-

Swyx9:20

So, so behind the scenes, you know, uh, one for listeners, one thing that we have, we have a tension about is who do we invite? Because obviously, if we get Mark Zuckerberg, it'll be a big name, then it will cause people to download us more, but it will be a less technical interview 'cause he's not on the research team.

He's, he's CEO of Meta. And so like, yeah, I think it, it's this constant back and forth. Like, we want to grow as a podcast, but we want to serve a technical audience, and we're trying to do that and tr- thread that line because our currency as podcasters is the people that listen to it, and we need big names, but we also need to serve our audience well.

And I think, um, if we, if we don't do it well, this actually goes all the way back to George Hotz when, when after he re- finished recording with us, he s- he said, "You have two paths in the podcast world.

Either you go be Lex Fridman or you stay, you stay small and niche." Uh, and we, we definitely like we like our niche. We, we think it's a, it's a good niche. It's gonna grow. But at the same time, we s- I, I still want us to grow.

I, I want us to grow on YouTube, right? And, and so, uh, that's, that's always like a, a Meta thing. Like not to get too Meta.

Alessio10:16

No, not that Meta. The, the other Meta. Um-

Swyx10:19

Uh, yeah. So Llama 3, yeah.

Alessio10:20

I think to me, the biggest thing is the training on outputs. Like every company is just hiding the fact that they've been fine-tuning and training on GPT-4 outputs, and you cannot technically do it, but obviously OpenAI is not enforcing it.

I think now for the first time, there's like a clear path to how do we make a 7B model good without having to go through GPT-4 or going to Claude 3. And we'll kind of talk about this later, but I think we're seeing maybe the, you know, not the death, but like selling the picks and shovels is kind of going away and like building the vertical things is like where most of the value is actually getting captured, at least at the early stages.

So being able to make small models better at specific things through a large model is more important than yet another 7B model that I can try and use, but at the end of the day, I still need to go through the large labs to fine-tune.

So that to me is the most interesting thing. You know, it's such a large model that like it's obviously amazing, but I don't know if a lot of people are switching from GPT-4 or Claude 3.5 to run 4 or 5B.

I also don't know what the hosting op- hosting options are as far as like scaling. You know? I don't know if, if the Fireworks and Togethers of the world, how much capacity they actually have to serve this model because at the end of the day, it's, uh, it's a lot of compute if some of the big products would switch to it and you cannot easily run it yourself.

So, um, I don't know. But to me, the synthetic data piece is definitely the most, the most interesting.

Swyx11:46

Yeah. Uh, I would say that it is not enough now to say that synthetic data is real. Uh, I sh- I actually shipped that in the original email, and then I changed that, uh, in the, in the so that what you see now in the, in the podcast description.

But because there it is so established now that synthetic data is real, therefore you need to go to the next level, which is, okay, what do you use it for and how do you use it? And I think that is what It was interesting for Llama 3 for me, if you, if you read the paper, 90 pages of, uh, all filler, no killer is, is, uh, something like that is, is what the people were saying.

Very, very, like for once, a frontier model with a proper paper instead of a marketing blog post. And, uh, you know, they, they actually spelled out how they'd use syn-synthetic data for a few different domains. So they have, uh, synthetic data for code, for math, for multilinguality, for m-long context, for tool use, and then also for ASR and voice generation.

And I think that, yeah, y- okay, now you have the license to go distill Llama 3,4 or 5B, but w- how do you do that? That is the sort of the next frontier. Now you have the permission to do it.

How do you do it? And, uh, I think people are, you know, gonna reference Llama 3 a lot, but then they can use those techniques, uh, for, for everything else. W- you know, in our episode with Thomas, he, he talked about syn- uh, like I was very focused on synthetic data for pre-training 'cause that's, that's my context, that's my conversations with Technium from News and all, and, um, all the other people doing synthetic data for pre-training and fine-tuning.

But he was talking about post-training as well and for, uh, everything here was post-training. Uh, in, in fact, I wish we had spent more time with Thomas on, on this stuff. Uh, we just didn't have the paper beforehand.

Uh, but I think like when I call Llama 3 the synthetic synthetic data model is you have the license for it, but then you also have the roadmap, the recipe because it's in the paper and now, like now everybody knows how to do this.

Uh, and, and probably, you know, o- obviously, like OpenAI is probably laughing at us 'cause they w- they did this like a year ago, but now it's in the open.

Alessio13:40

I mean they can laugh all they want, but they're coming for them. Um, I, I think, I mean, that's definitely the biggest vibe shift, right? It's like obviously Llama 3.1 is good. Obviously Claude is good. Maybe a year and a half ago you didn't get the benefit of the doubt as like an OpenAI competitor to be state-of-the-art.

You know, it was kinda like, oh, Anthropic, yeah, those guys are cute over there. They're trying to do their thing, but it's not OpenAI and like Llama 2 is great but like it's really not a serious model, you know?

It's like just good enough. I think now it's like every time Anthropic releases something, people are like, okay, this is like a serious thing. Whenever like Meta releases something, it's like, okay, they're at the same level and-

Swyx14:18

Mm.

Alessio14:18

I don't know if OpenAI is kinda like sandbagging the GPT next-

Swyx14:22

They're releasing wait, wait lists.

Alessio14:22

You know? Yeah. I don't... And then they kinda, you know, yesterday or today they announced the search GPT thing-

Swyx14:30

Yeah

Alessio14:30

... behind the wait list, you know.

Swyx14:31

Th- this is the Singapore confusion. Wh- when was it that something happened?

Alessio14:33

Yeah. When was it? Yeah.

Swyx14:34

'Cause it happened yesterday US time, but today Singapore time.

Alessio14:37

Yeah. So Thursday.

Swyx14:39

Um-

Alessio14:39

It's been really confusing. But yeah, and people are kinda like, "Uh-oh, okay, OpenAI, I don't know if we can take you seriously now." So-

Swyx14:47

Well, no, uh, then one of the, um, uh, AI Grants, um, employees, I think Hirsh tweeted that, you know, you can skip the wait list, just go to perplexity.com.

And it was, it was a really, really sick burn for the, uh, OpenAI search GPT wait list. But y- their implementation will have something different. They probably like train a sp- dedicated model for that, you know? Like it-- they'll have some innovation that we haven't seen.

Alessio15:09

Yeah. Data licensing obviously. Um-

Swyx15:11

Uh, data licensing. Yes. We're optimistic, uh, uh, you know, uh, but the, the vibe shift is real, and I think that's something that is just worth commenting on and watching and, um, yeah, how the other labs catch up.

Small Models15:22

Swyx15:22

I think what you said there is actually very interesting. The, the trend of successive releases is very important to watch. If things get less and less exciting, then it's a red flag for that company, and if it's-- things get more and more exciting, it means that they-- these guys have, have a good team, they have a good plan, good ideas.

Um, so yeah, like, uh, I will call out, you know, uh, the Microsoft Phi team as well. Um, Phi-1 was kind of widely regarded to be over-trained on benchmarks, and Phi-2 and Phi-3 subsequently improved a lot as well.

I would say also similar for Gemma, Gemma 1 and 2. Uh, Gemma 2 is currently leading in terms of the, uh, local Llama sort of vibe check eval informal straw poll and, and th- and that's only like a month after release.

They released at the, um, AI Engineer World's Fair. And, um, you know, like I didn't know what to think about it 'cause Gemma 1 wasn't like super well received. It was just kind of like here's, here's like free tier Gemini, you know?

But, but now Gemma, Gemma 2 is actually like a, a very legitimately widely used, uh, model by, by the open source and, uh, local Llama community. So that's great until Llama 3 7B came along.

Alessio16:28

Yeah.

Swyx16:28

And so like the, uh... and we'll talk about this also, like just the, the winds of AI winter is als- also like what is the depreciation schedule on this, on this model inference and training class? Like it's, it's very high.

Alessio16:39

Yeah. I'm curious to get your thought on Mistral, uh, everybody's favorite sparkling weights-

Swyx16:45

Yeah

Alessio16:45

... um, company.

Swyx16:46

Yeah.

Alessio16:47

Um, they, they just released the, you know, Mistral Large Enough.

Swyx16:50

Large- Mistral Large 2.

Alessio16:51

Yeah. Large 2.

Swyx16:52

Uh, so this was one day after Llama 3, uh, presumably because they were speaking at ICML, which is going on right now. Uh, by the way, uh, Britney is doing a guest host, uh, thing for us. She's, she's running around the poster sessions doing what I do, which is very great 'cause I couldn't go 'cause of my visa issue.

I have to be careful what I say here, but I think because we still want to respect their work, uh, but, uh, Mistral Large I would say is like not as exciting as Llama 3. That-- I think that is very, very fair to say.

It is, yes, another GPT-4 class model released as open weights with a research license, not a commercial license, but still open weights, and that's good for the community. But it is a step down in terms of, uh, the general excitement around Mistral compared to Llama.

I think that would be fair to say, and I would say that to Mistral themselves. So the, the general hope is, and, and I've I cannot say too much just 'cause I have, I've had offline conversations with the people close to this.

The general hope is that th- they need something more. You know, of the 10 elements of like what is next in terms of their frontier model boundaries, Mistral needs to make progress there. They made progress here with like, um, instruction following and structured output and multilinguality and all those things.

But I think to stand out, you need to basically pull a stunt. You need to be a superlatively good company in one dimension, and now unfortunately, Mistral does not have that crown as open source kings. Uh, you know, like, like a year ago I was saying, Mistral are the kings of open source AI.

Now Meta is They've lost their ground. Uh, by the way, they've also deprecated, uh, Mistral 7B, eight by 7B, and eight by 22B, right? So n-now there's only, like, the closed source models that are API platform. So has Mistral basically started becoming more of a closed model op-- uh, proprietary platform.

I don't believe that's, that's true. I believe that they, they, they're still very committed to open source, uh, but they need to come up with something more that people can use. A-and that's a, that's a grind. I mean, they have, what?

Six hundred million dollars to do it.

Alessio18:44

Mm-hmm. Right.

Swyx18:44

So that's, that's still good. Uh, but, you know, people are waiting for, like, what's next from, from them.

Alessio18:49

Yeah. To me, the perception was interesting. In the comments of the release, everybody was like: Why do you have a non-commercial license? You're not making any money anyway from the inference. So I feel like the AI engineering tier list, you know, is kinda shifting in real time, and maybe Mistral, like you said before, was like, "Hey, thank God for these guys.

They're saving us in open source. They're kinda like speed running GPD 1, GPD 2, GPD 3 in open source." But now it's like they're kinda moving away from that. I haven't really heard of that many people using them at scale commercially just from, you know, discussions.

Um, so I'm curious to see what the next step is.

Swyx19:26

Yeah. But also you're sort of US-based, and maybe they're not focused there, right? So-

Alessio19:30

Yeah. Exactly. That's the-

Swyx19:31

The, the... It's a very big elephant, and we're only touching pieces of it. It's, it's blind, you know, blind leading the blind. Uh, I, I will call out, um, you know... They, they have some interesting experimentations with Mamba, and Mistral Nemo is actually on the efficiency frontier chart that I drew that is still relevant.

So don't discount Mistral Nemo. But Mistral Large, uh, otherwise, like it's an, it's an update, it's a necessary update for Mistral Large V1. But other than that, they're just kind of holding the, the line-

Alessio19:58

Mm-hmm

Swyx19:58

... not really advancing the field yet. That, that'll be my statement there. So those are the frontier big labs.

Alessio20:05

Yes.

Swyx20:06

And then now we're t- we're gonna shift a little bit towards the, the smaller deployable on-device solutions.

On-Device LLMs20:06

Alessio20:11

Yeah. First of all, shout out to our friend Tri Dao, uh, who released FlashAttention 3. FlashAttention 2, we kinda did a deep dive on the podcast. He came on, uh, in the studio back then. It's just great to see how small groups can make a big impact on a whole industry, just like by making math better.

So it's just great to see. I just wanted to give Tri a shout-out.

Swyx20:34

Something I mentioned there and something with-- that always comes up, even in the Sovereignty AI Summit that we did, was what-- does Nvidia's competitors have a, have any threat to Nvidia? You know, AMD, like Madex, like, um, Etched, you know, which, which caused a lot of noise with their Soho chip as well.

And just the simple fact is that Nvidia has won the hardware lottery and people are customizing for Nvidia. Like FlashAttention 3 only works for Nvidia, only works for H100s. And, like, this much work, this much scaling, this much validation going into this stuff is very difficult to replicate or very expensive to replicate for the other hardware ecosystems.

So not impossible. I actually heard a really good argument from one of, um... I think it is Martin Casado from a16z who was saying basically like, yeah, like absolutely Nvidia's hardware and ecosystem makes sense and obviously that's gon- that's contributed to its like, I don't know, like, it's like the most valuable company in the world right now.

But current training runs are like one hundred million to two hundred million in cost. But when they go to five hundred million, when they go to a billion, when they go to one trillion, then you can ac-actually start justifying making custom ASICs for your run.

Alessio21:44

Mm-hmm.

Swyx21:45

And if they, if they cut your costs by like half, then you make your money back in one run.

Alessio21:50

Yeah, yeah, yeah. Martin has always been a, a fan of, uh, custom ASIC. I think he... They wrote a really good post maybe a couple of years ago about, um, cloud repatriation-

Swyx21:58

Oh, yeah

Alessio21:59

... and, uh, custom ASIC, so-

Swyx22:00

I think he got a lot of shit for that, but, uh-

Alessio22:01

Uh-

Swyx22:01

... it's becoming more consensus now, I think. So Noam Shazeer blogging again, fantastic gift to the world. This guy, nonstop bangers. Uh And so he's at Character.AI, and, uh, they, they-- he put out a post talking about five tricks that they use to serve twenty percent of Google search traffic as LLM inference.

A lot of people were very shocked by that number, but I think you just have to remember that most conversations are multi-turn, right? Like in the span of one Google search, I will send like ten text messages, right?

So obviously there's like a ratio here that, uh, that matters. It's obviously a flex of Character.AI's traction among the kids, because I have tried to use Character.AI since then, and I still cannot for the life of me get it.

I don't-- Have you tried, uh-

Alessio22:46

I tried it, but yeah, it's definitely not, not for me.

Swyx22:48

Yeah, they launched like voice. I tried to talk to it. It was just so stupid. I just I didn't like it myself. But, uh, this is what it means-

Alessio22:56

But please still come on the podcast, Noam Shazeer.

Swyx22:57

Yeah, yeah.

Alessio22:58

Sorry. We didn't mean

Swyx22:59

No, no, no. It's, it... Because like, uh, like I, I don't really understand like what the use case is for apart from like the, the therapy role play homework assistant-

Alessio23:06

Mm-hmm

Swyx23:06

... uh, type of, type of stuff that is the norm. But anyway, one of the thing-- most interesting things, the detailed five tricks, um, one thing that people talk a lot about is native int8 training. I, I got it wrong in our Thomas podcast.

I said FP8. It's in-int8. And I think like that is something that is a easy win. Like, uh, we should, we should basically when we're getting to the point where we're over-training models one hundred times past Chinchilla ratio to optimize for inference, the next thing is actually like, hey, let's stop using so much memory, uh, when, when training because we're not-- we're gonna quantize it anyway for inference.

So like just let's pre- let's pre-quantize it in, in training. So that makes a lot of sense. The other thing as well is this concept of global-local hybrid architecture, which I think is basically going to be the norm, right?

Uh, so he has this formula of one to five ratio of global attention to local attention, and he says that that is, that is, uh, that works for the long-form conversations that Character has. Okay, that's great. And like simultaneously, we have independent research from other companies about similar hybrid ratios being the best for their research.

So Nvidia came out with a Mamba transformer hybrid research thing, and their, in their estimation, you only need seven percent transformers. Everything else can be state space models. Jamba also had something like between like six to like thirty to one.

And, uh, it-- basically, every form of hybrid architecture seems to be working at the research stage. So I think like if we scale this, uh, it makes f- complete sense that you, you just need a, a, a mix of architectures and, um, it could well be that the transformer block instead of transformers being all you need, transformers are the global attention thing, and then the, the local attention thing can be the state space models, can be the RWKVs, can be another transformer, but just, uh, limited by a sliding window.

And I, I think like we're slowly discovering like the, the fundamental building blocks of AI. One is transformers, one is something that's local, whatever that is. And then, you know, who knows what else is next. Uh, maybe-- The, the other, the other stuff is adapters, but we can talk about that.

But yeah, headline is The Gnome, maybe he's too confident, but I, I mean, I, I believe him . Gnome thinks that he can do inference at 13X cheaper than the Fireworks together, right? So, like, there is a lot of room left-

Alessio25:18

Mm-hmm

Swyx25:19

... uh, to, to improve inference.

Alessio25:20

I mean, it does make sense, right? Because, like, otherwise, I don't know-

Swyx25:23

Otherwise Character would be bankrupt by now .

Alessio25:24

Yeah, exactly. I was like, they would be- ... they would be losing a ton of money, so, um-

Swyx25:28

They are rumored to be exploring a sale. Um-

Alessio25:31

Yeah

Swyx25:31

... so I'm sure money is still an issue for them, but I'm s- also sure they're making a lot of money, so I... it's very hard to tell because it's not a very public company.

Alessio25:39

Yeah. Well, I, I think that's one of the, mm, things in the market right now too, is like, "Hey, do you just wanna keep building? Do you wanna, like, just not worry about the money and go build somewhere else?"

Kinda like a maybe inflection and adapt in some of these other non-equi-hires licensing deals and whatnot. So I'm curious to see what companies decide to stick with it.

Swyx26:00

It's strong, like y- uh, I think Google or Meta should pay $1 billion for Gnome alone.

Alessio26:06

Mm-hmm. Right.

Swyx26:06

The, the purchase price for Character is a o- $1 billion.

Alessio26:09

Mm-hmm.

Swyx26:10

Which is super underpriced.

Alessio26:11

Which is nothing at their market caps, right?

Swyx26:12

It's nothing. It's nothing.

Alessio26:13

Like Meta's market cap right now is-

Swyx26:15

Is at $2 trillion

Alessio26:16

... $1.15 trillion-

Swyx26:17

Yeah

Alessio26:17

... because they're down 5%, 11% in the past month.

Swyx26:21

What?

Alessio26:22

Um, yeah. So if you pay $1 billion, you know, that's like 0.01%-

Swyx26:27

Yeah, yeah, yeah

Alessio26:28

... uh, of your market cap. And they paid, they paid $1 billion for WhatsApp, and they buy 1% of their market cap-

Swyx26:34

Yeah

Alessio26:34

... on that at the time, so.

Swyx26:35

Yeah. That is beyond our pay grade. But the, the, the last piece of the GPU Rich Poor Wars, so, so we're going from the super GPU rich down to like the medium GPU rich, and now, now down to the GPU poors, is on-device models, right?

Uh, which is something that people are very, very excited about. So at my conference, uh, Mozilla AI, I think was kind of like the talk of the town there on Llamafile. Uh, we had Justine Tanie come in and explain like some of the optimizations that they did, and their, their just general vision for on-device AI.

I think that, like, it's basically the second act of Mozilla. Like a lot of good with the open source browser and, uh, obviously then they have since declined because, uh, it's, it's very hard to keep up in, in that field, and Mozilla has had some management issues as well.

But now, now that the operating system is moving to the AI layer-

Alessio27:22

Mm

Swyx27:22

... now they're also like, you know, promoting open source AI there and also like private AI, right? Like o- open source is synonymous with local, private, and all the good things that people want. And I think their vision of like even running this stuff on CPU is at, at a very, very fast speed by, by just like being extremely cracked .

Alessio27:38

Oh, yeah.

Swyx27:39

I think is, uh, very understated and, uh, we should probably try to support it, uh, more. Um, and it's just amazing to host these, these people and, uh, see the progress.

Alessio27:49

Mm-hmm. Yeah. I think to me the biggest question about on-device, uh, obviously there's, uh, Gemini Nano, which is getting sh- uh, shipped with Chrome.

Swyx27:56

Yeah. So, so let's survey, right? So Llamafile is one executable that runs on every architecture.

Alessio28:00

Yep.

Swyx28:01

Similar for, by the way, Mojo from, from Modular, which also spoke at the conference. And then what else? Llama.cpp, MLX, those, those kinds are, are also that, that layer. Then the next layer up would be the built-in into, uh, their products by the, by the vendors.

So Google Chrome is building Gemini Nano into the browser. The next version of Google Chrome will have Nano inside that you can use like window.ai.something, and it would just call Nano. There, there'll be no download, no latency what-whatsoever 'cause it runs on your device.

And there's Apple Intelligence as well, which is Apple's version, which is f- in, in the OS accessible by apps. And then there's, there's a bu- long tail of others. But like, yeah, your, your comments on those things.

Alessio28:43

My, my biggest question is how much can you differentiate at that model size? You know, like how big it's gonna be, the performance gap between all these models and like, are people gonna be aware of what model is running?

You know, right now for the large models, we're still pretty aware of like, oh, is this Sonnet 3.5? Is this GPT-4? Is this, you know, 3.1, 4 or 5B? I think the smaller you get, the more it's just gonna become like a neutrality, you know?

So like you're not gonna need a model router for like small models. You're not gonna need any of that. Like they're all gonna converge to like-

Swyx29:17

Well-

Alessio29:17

... the best possible performance at this point

Swyx29:19

... actually, Apple Intelligence is the model router, I think. Uh, they have something like 14... I, I did a count in, in my newsletter, like 14 to 20 adapters.

Alessio29:27

Mm-hmm.

Swyx29:27

And so based on your use case, they, uh, they'll, they'll route and, and load the adapter, or they'll route to OpenAI. So there is some routing layer. Um-

Alessio29:34

Yeah, yeah

Swyx29:35

... to me, I think a lot of people were trying to puzzle out the strategic moves between OpenAI and, and Apple here because Apple is in a very good position to commoditize OpenAI. There were some rumors that Google was working with Apple to launch it.

They did not make it for the launch. But presumably, Apple wants to commoditize OpenAI, right? So, th- y- you know, when you, when you launch, you can choose your preferred-

Alessio29:55

Mm-hmm

Swyx29:56

... external AI provider, and it's either OpenAI or Google or someone else. I mean, that puts Apple, Apple at the center of the world, uh, the, as, as with the ability to make routing decisions and, um, I think that's probably good for privacy, probably good for the planet 'cause you're -

Alessio30:12

Yeah

Swyx30:12

... you're not running like oversized models on like your, your, you know, your, your spellcheck, uh, pass. And, uh, I, I'm generally pretty positive on it. Like, yeah, I, I'm not concerned about the capabilities issue. It meets their benchmarks.

Apple put out a whole bunch of proprietary benchmarks 'cause they don't like to do anything in the way that everyone else does it . So like, you know, in the Apple Intelligence blog post, they like... I think like all of them were just like their internal human evaluations, and only one of them was an industry standard benchmark, which is, which is IF eval.

Which is good, but like, you know-

Alessio30:41

Yeah, yeah

Swyx30:41

... why didn't, why didn't you also release your MMLU? Oh, 'cause you'd suck on it. All right .

Alessio30:46

Well, I, I actually think all these models will be good. And on the Apple side, I'm curious to see what the price tag will be to be the default. Right now, Google pays them $20 billion to be the default search.

Swyx30:58

I see.

Alessio30:58

Uh, I, I wonder-

Swyx30:59

The rumor is it's zero.

Alessio31:01

Yeah, yeah. I mean, today, even if it was $20 billion-

Swyx31:03

I see

Alessio31:03

... that's like nothing compared to like, you know, NVIDIA's worth $3 trillion. So like paying 20-- even paying $20 billion to be the default AI provider like would be- Cheap compared to search, given that AI is actually being such a core part of the experience.

Like Google being the default-

Swyx31:18

Um

Alessio31:18

... for like Apple's phone experience-

Swyx31:20

Yeah

Alessio31:20

... really doesn't change anything.

Swyx31:21

Yeah, yeah.

Alessio31:22

Becoming the default AI provider for like the Apple experience-

Swyx31:25

Yeah

Alessio31:25

... would be worth a lot more than this.

Swyx31:27

Yeah. I mean, so I can justify it being zero instead of 20 billion is because OpenAI has to foot the inference cost, right? So th- that's a lot.

Alessio31:33

Well, yeah, Microsoft really is footing it. But again, Microsoft is worth 2 trillion.

Swyx31:38

Yeah.

Alessio31:38

You know?

Swyx31:39

So as someone who this is the web developer coming out, uh, as someone who is a champion of the open web, Apple has been, let's just say, a roadblock in that, in that direction. I think Gemini Nano being good is more important than Apple Intelligence being generally capable.

Apple Intelligence being like-

Alessio31:55

Yeah

Swyx31:55

... a on-device router for Apple apps is good, but like if you care about the open web, you really need Gemini Nano to work and, uh, we're not sure. Like right now we have some, some demos showing that it's fast enough, but we haven't had systematic tests on it.

Along the lines of that research, I, I will highlight that Apple has also put out DataComp LM. I actually interviewed DataComp at NeurIPS last year, and they've branched out from just vision and images to language models, and Apple has put out a reference implementation of the 7B language model that's built on top of DataComp, and it is better than FineWeb, which is huge because FineWeb was the state-

Alessio32:31

Right

Swyx32:31

... of the art last month.

And that's fantastic. So, so basically like DataComp is an open data, open weights, open model, like super everything open. So there will be a lot of people optimizing this, this kind of model. They'll be building on architectures like MobileLM and SmallLM, which basically, uh, innovate in terms of like shared weights and shared matrices for, uh, for smaller models so that you, you just optimize the, the amount of file size and, and, you know, memory that you take up.

And, um, I think just general trend of on-device models, like the only way that intelligence to cheap to meter happens is everything happens on device. So, uh, unfortunately, that means that OpenAI is not involved in this. Like OpenAI's mission is intelligence to cheap to meter, and they're not doing the one thing that needs to happen for that because there's no business plan in monetizing an API for that.

Like by definition, none of this is APIs.

Alessio33:23

I don't know. I guess Johnny AI and Sam Altman need to figure it out so they can do, uh-

Swyx33:27

Yeah

Alessio33:27

... their own AI device.

Swyx33:28

I'm excited for OpenAI phone.

Alessio33:29

Yeah.

Swyx33:29

I don't know if you would, you would buy an OpenAI phone. I mean, I'm very locked into the iOS ecosystem, but I mean-

Alessio33:33

I will not be the first person to buy it- ... because I don't wanna be stuck with like the Rabbit equivalent- ... of a AI phone, but I think it makes a, a lot of sense. I want their-

Swyx33:41

And, you know, they're building a search engine now. The next thing is the phone.

Alessio33:45

Exactly. So we'll see. I don't know.

Swyx33:48

We'll see when it comes on the wait list. We'll, we'll see.

Quality Data Wars33:48

Alessio33:50

Yeah, yeah. We'll, we'll review it. All right. So that was GPU Rich, GPU Poor. Maybe we just wanna run quickly through the Quality Data Wars. Um, there's maybe-- there's mostly drama in this section. There's not as much as much research.

Swyx34:04

Yeah. I think, I think there's a lot of news going in the background. So like-

Alessio34:06

Yeah

Swyx34:07

... the New York Times lawsuit is still ongoing.

Alessio34:09

Yeah.

Swyx34:09

You know, it's just like we won't have specific things to update-

Alessio34:13

Mm-hmm

Swyx34:14

... uh, people on. Uh, there, there are specific deals that are happening all the time with Stack Overflow making deals with everybody, with like Shutterstock making deals with everybody. It's just-- it's hard to make a single news item out of something that is just slowly cooking in the background.

Alessio34:27

Mm-hmm. Yeah. The, um-- on the New York Times thing, OpenAI's strategy has been to make The New York Times prove that their content is actually any original or like actually interesting.

Swyx34:39

Really?

Alessio34:39

Yeah. So it's kinda like, you know, the iRobot meme. It's like, uh, can a robot create a beautiful new symphony? And the robot is like, "Can you?" Um- ... I think that's the, that's what OpenAI's, um-

Swyx34:51

Yes

Alessio34:51

... strategy is.

Swyx34:52

So, so yeah. I think the, the danger with the lawsuit, um, because th- this lawsuit is very public 'cause OpenAI responded, including with Ilya, showing their emails with New York Times saying that, "Hey, we, we were doing a deal.

We, we were like very close to a deal, and then suddenly on the, the eve of the deal, you called it off." Uh, I don't think The New York Times has responded to that one, but, uh, it's very, very strange because The New York Times's brand is like trying to be like, you know-- They're, they're supposed to be the top newspaper in the country.

Alessio35:18

Mm-hmm.

Swyx35:19

If OpenAI like just-- And this was my criticism of it at this, at the point in time, like, okay, we'll just go to the next best paper.

Alessio35:25

Yeah.

Swyx35:26

The Washington Post, the Financial Times, they're all happy to work with us. And then what does New York Times have?

Alessio35:31

Yeah, yeah.

Swyx35:32

So you just lost out on like $100 million, $200 million a year of, of, uh, licensing deals just because you wanted to pick that war, which ideologically, I think they're absolutely right to do that. But, um, you know, the other people-- Uh, The Verge did a very good interview with, I think, The Washington Post.

I'm gonna get the, the, the, the outlet wrong. The, The Verge did a very good interview with a newspaper owner-editor on why they did the deal with OpenAI. And I think that listening to them on like their thinking through like the, the reasoning of like the pros and cons of picking a fight versus partnering, I think is very interesting.

Alessio36:09

Yeah. The, uh, I guess the winner in all of this is Reddit, which is making over 200 million just in data licensing to OpenAI and some of the other AI providers. Um, I mean, 200 million is like more than most AI startups are making.

Swyx36:23

Yeah.

Alessio36:23

Than to just, um-

Swyx36:23

So I think that was an IPO play because Reddit-

Alessio36:25

Yeah, yeah, yeah

Swyx36:25

... conveniently did this deal before IPO, right?

Alessio36:27

Totally.

Swyx36:28

Is it like a one-time deal and then, you know, the, the stock languishes from there? I don't, I don't know.

Alessio36:32

Yeah. No, well, their IPO has done-- Well, I guess it's not gone down, so in this market they're up 25%, I think, since IPO. But I saw the FTC had opened a inquiry into it just to like, uh, investigate.

So I'm curious what the antitrust regulations are gonna be like when it comes to data. Obviously, acquisitions are blocked to prevent kinda like stifling competition. I wonder if for data it will be similar, where, hey, you cannot actually gate all of your data only behind $100 million-plus contracts because otherwise you're Stopping any new company, uh, from building a competing product, so, um-

Swyx37:09

Yeah, I-- that's a serious overreach of the state there-

Alessio37:12

Yeah, yeah

Swyx37:12

... I would say. As a, as a, as a free market person, I want to defend... It is weird. Like I'm a free market person, and I'm a content creator, right? So I want to be paid for my content.

Alessio37:20

Right.

Swyx37:20

At the same time, I c- I believe that, you know, people should be able to make their own decisions about, uh, uh, all these deals. But UGC is a weird thing-

Alessio37:27

Mm-hmm

Swyx37:28

... 'cause UGC, UGC is contributed by volunteers.

Alessio37:31

Yeah.

Swyx37:31

And the, the other big news of Reddit is that apparently they have added to their robots.txt, like only Google should index us, right? 'Cause we did the deal with Google, and that's obviously blocking OpenAI from crawling them, Anthropic from crawling them, you know, Perplexity from crawling them.

Perplexity maybe ignores all robot.txt- ... but that's a whole different, uh, other issue. And then the other thing is I, I think this is big in the, the sort of normie worlds, um, the actors, you know, Scarlett Johansson had a very, very public Apple Notes takedown of OpenAI.

Only Scarlett Johansson can do that to, to Sam Altman and then, you know. I was very proud of my newsletter for, for that day. I called it Skyfall because the voice of the, that voice was Sky, so I called it Skyfall.

And but it's true, like, uh, you, uh, there's-- that, that one she can win. Um, and there's a very well-established case law there. And the YouTubers are in the music industry, the RIAA, like the most litigious, uh, section of the creator economy, um, has gone after Udio and Suno, you know, Mikey from our, uh, our podcast with him.

And, uh, it's un-unclear what will happen there, but, um, it's, it's gonna be a very costly legal battle for sure.

Alessio38:34

Yeah. Um, yeah, I mean, music industry and lawsuits name a more iconic duo, you know? So-

Swyx38:40

Yeah

Alessio38:40

... I think that's, uh, to be expected.

Swyx38:42

So I was-- Yeah. I think the last time we talked about this, I, I was, uh, pretty optimistic that this, something like this would reach the Supreme Court and with the way that the Su- the Supreme Court is, is making rulings, like we just need a judgment on whether or not training on, on data is transformative use.

'Cause I think it is. Literally, we're using transformers to, to do transformative use. So then it's open season for, for AI to do it, and comparatively, the content creators and owners will lose out. They just will.

Alessio39:09

Yeah.

Swyx39:09

'Cause right now we're paying them money out of fear of lawsuits. If the Supreme Court rules that there are no lawsuits to be had, then all the money disappears.

Alessio39:17

I think people are probably scraping latent space, and we're not getting a dime, so- ... it is what it is.

Swyx39:22

Yeah. Uh, no, you can, you can support with like an $8 a month subscription-

Alessio39:26

Yeah.

Swyx39:26

... and, and that pays for our microphones and travel-

Alessio39:28

Yeah

Swyx39:28

... and stuff like that. Yeah, it's, you know, it's not, it's definitely not worth the amount of time we're putting into it, but it-

Alessio39:33

Well-

Swyx39:33

... it's a labor of love.

Alessio39:34

Yeah, exactly. I-

Swyx39:36

Synthetic data.

Alessio39:37

Yeah. Yeah, I me- I guess we talked about it a little bit before with Llama, um, but there was also the AlphaProof thing.

Swyx39:43

Yes.

Alessio39:43

Uh-

Swyx39:44

Just before I came here, I was working on that newsletter.

Alessio39:47

Yeah. Yeah, Google trained. They almost got a gold medal? I forget what the-

Swyx39:50

Yes, they're one point short of the gold medal.

Alessio39:51

Yeah, one point short of the gold medal.

Swyx39:53

It's a, it's a remarkably-- I wish they had more questions. So the, the, the... So the International Math Olympiad has six questions, and each, each question is seven points. Every single question that the AlphaProof model tried, it got full marks on.

It just failed on two, and then the cutoff was, was like sadly one, one point, uh, higher than that. But still, like it was a, it was a very big... Like a lot of people have been looking at IMO as like the next gold prize, grand prize in terms of what AI can achieve.

And betting markets and Evleiza Yukovski has, has, has updated in saying like, "Yep, like we're, we're pretty close." Like we, we basically have reached it near gold medal status.

Alessio40:32

Mm-hmm.

Swyx40:32

We definitely reached, uh, silver and bronze status, and, uh, we'll probably reach gold medal next year, right? Uh, which is good. There's also related work from Hugging Face on the Numena Math Competition. So this is on the AI Mathematical Olympiad, which is an easier version of the, the human, uh, Math Olympiad.

This is all like related research work on search and verifier model-assisted exploration of, of, of mathematical problems. So yeah, that's super positive. I don't really know much else beyond that. Like it's, it's always hard to cover this kind of news 'cause it's not super practical.

Alessio41:03

Mm-hmm. Yeah.

Swyx41:03

And it also doesn't generalize. So one thing that people are talking about is this concept of jagged intelligence. 'Cause at the same time we're having this as a discussion about being superhuman. You know, one of the, um, IMO questions was solved in 19 seconds after we gave the question to, uh, AlphaProof.

At the same time, language models cannot determine if 9.9 is smaller than or bigger than 9.11. And part of that is 9.11 is an inside job. But it's, it's a funny... And that is someone else's joke. I don't know.

I, I really like that joke. Uh, but it's, it's jagged intelligence. It's, it's just a failure-

Alessio41:38

Yeah

Swyx41:38

... to generalize because of tokenization and because of whatever. And what we need is general intelligence. We've always been able to train dedicated special models to, to win prizes and do stunts. But the grand prize is general intelligence, that same model does everything.

Alessio41:52

Is it gonna work that way? I don't know. I think like if you look back a year and a half ago, and you would say, "Can one model get to general intelligence?" Most people would be like, "Yeah, we can keep scaling."

I think now it's like, is it gonna be more of a mix of models, you know? Like, can you actually do one model that does it all?

Swyx42:10

Yeah, absolutely. I think GPT-5 or Gemini 3 or whatever, uh, would be much more capable at this kind of stuff while it also serves our needs, uh, with like, with everyday things. It might be a... It might be completely uneconomical.

Alessio42:26

Mm-hmm.

Swyx42:26

Like, why would you use-

Alessio42:27

Right, yeah

Swyx42:27

... a giant-ass model to, to do normal stuff? But it is just a demonstration of proof that we can build super intelligence for sure. And, and then, you know, everything else follows from there. But, uh, right now we're just pursuing super intelligence.

I always think about this, uh, you know, just reflecting on the GPU Rich Poor stuff and then now this Alpha Geometry stuff. I used to say you pursue capability first, then you efficient- make, make it more efficient. You make frontier model, then you distill it down to the AP-7B, 70B, which is what Llama 3 did.

And by the way, also OpenAI did it with GPT-4o and then distilled it down to 4o Mini, and then Claude also did it with Opus and then with 3.5 Sonnet, right? Uh, that suitable recipe, I in-in fact, I call it part of the deployment strategy of models.

You train a base layer, you train a large one, and then you distill it down. You add structured output generation, tool calling and all that. You add the long context. You add like this, this standard stack of stuff in post-training that is growing and growing to the point where now OpenAI has opened a team for mid-training-

Alessio43:26

Mm-hmm

Swyx43:26

... that happens before post-training. I think like- One thing that I've realized from this Alpha Geometry thing is before you have capability and you have efficiency, there's an in-between layer of generalization that you need to accomplish. You need to do capability in one domain, you need to generalize it, then you need to efficient-size it.

Alessio43:46

Mm-hmm.

Swyx43:47

Then you have good models.

Alessio43:50

That makes sense. I, I think like maybe the question is how many things can you make it better for before generalizing it, you know? Yeah, I don't have a good intuition for that but-

Swyx44:00

Uh, we'll talk about that in the next, uh-

Alessio44:02

Yeah

Swyx44:02

... next thing. Yeah, so we can skip Nemo- Nemotron's worth looking at if you're interested in synthetic data. Uh, multimodal labeling I, I think has happened a lot. We'll, we'll maybe we'll just jump to multi-modal now.

Multimodality War44:10

Alessio44:11

Mm-hmm. Yeah, we got a bunch of news. Well, the first news is that, uh, 4o voice is still not out even though the, the demo was great. I think it-- they're starting to roll out the beta in the next week.

Swyx44:22

Yeah. So, so, uh, I am subscribing-- I subscribed back to ChatGPT Plus.

Alessio44:25

You gave in?

Swyx44:26

I gave in because they're rolling it out next week. So you better be on the, the cutoff or-

Alessio44:30

Nice bait

Swyx44:31

... you're not gonna get it.

Alessio44:32

Nice bait, Samo, man.

Swyx44:33

I, no, I said this. I said when I talk about unbundling on ChatGPT, it's, it's basically 'cause they had nothing to offer people. That's why people are unsubscribing 'cause like why keep paying $20 a month for this, right?

But now they have proprietary models, oh yeah, I'm back in, right? Like I want-

Alessio44:45

We're so back.

Swyx44:46

We're so back.

Alessio44:46

We're so back.

Swyx44:47

We're so back.

Alessio44:47

Yeah.

Swyx44:47

I, I would pay 200 for the Scarlett Johansson voice, but, you know- ... they'll probably get sued for that.

Alessio44:51

Mm-hmm.

Swyx44:52

Um, but, but yeah, the, the v-voice, voice is coming. We had a demo at the World's Fair that was, uh, that was I think the, the second public demo. Uh, Ro- Roman, I, I have to really give him a shout-out for that.

We had a few people drop out last minute and he was-- he say he rescued the conference and, and, and worked really hard. Like, you know, I think off the scenes, I think something that people don't understand is OpenAI puts a lot of effort into their presentations and if it's not ready they won't launch it.

Alessio45:18

Mm-hmm.

Swyx45:18

Like he, he was ready to call it off if, if we didn't make the AV work for him. And I think, yeah, they care about their presentation and how they launch things to people. Those minor polish details really matter.

Um, just for the record for, for people who don't understand what happened was, uh, you can-- first of all, you can go see, just look for like the GPT-4o talk at the AI Engineer World's Fair. But second of all, because it was presented live at a conference with large speakers blaring next to you and it is a real-time voice thing, so it's, it's listening to its own voice and it needs to distinguish between its own voice-

Alessio45:47

Right

Swyx45:47

... and between the, the human voice and it needs to ignore its own voice. So we had opening engi-engineers tune that for our stage to make this thing happen which is absurd.

Alessio45:56

Yeah.

Swyx45:57

It was so funny but also like, you know, shout out to them for, for doing that for us and for the community, right? Uh, because I, I think people wanted an update on voice.

Alessio46:05

Mm-hmm. Yeah. They definitely do care about demos. Uh, not much to add there.

Swyx46:09

Yeah.

Alessio46:10

Llama 3 voice.

Swyx46:11

Something that maybe is buried among all the Llama 3 news is that Llama 3 is supposed to be a multimodal model. It was delayed thanks to the European Union apparently. I'm not sure what the, the whole story there is.

I didn't really read that much about it. It is coming. You know, Llama 3 will be multimodal. It uses adapters rather than being natively multimodal. But I think that it's interesting to see the state of La-- of Meta AI research come together because there was this independent threads of Voice Box and Seamless Communication.

These are all projects that Meta AI has launched that basically didn't really go anywhere because they were all one-offs. But now all that research is being pulled in into Llama. Like Llama is just subsuming all of FAIR, all of Meta AI into this thing.

And, and yeah, you can see, uh, a Voice Box mentioned in Llama 3 voice adapter. I was kind of bearish on conformers because I looked at the state of existing conformer research in ICM, Clear and, and NeurIPS and they were far, far, far behind Whisper, uh, mostly because of scale, like the, the sheer amount of resources that they dedicated it.

But Meta is, is approaching there. It's-- I think it's, um, they had 230 hours, 230,000 hours of speech recordings. I think Whisper is something like 600,000. So Meta just needs to 3X the budget on this thing and they'll do it and, uh, we'll have open source voice.

Alessio47:31

Yeah. And then we can hopefully fine-tune on our voice and then we just need to write this episode instead of actually recording it.

Swyx47:38

I should also shout out the other thing from Meta which is a very, very big deal which is Chameleon which is a natively early fusion vision and language model. So most things are late fusion basically like you freeze an existing language model, you freeze an existing vision transformer and then you kind of fuse them with this, that thin adapter layer.

That, that is what, uh, Llama 3 is also doing. But Chameleon is slightly different. Chameleon is interleaving in the same way that Ida Fix, uh, the sort of, um, dataset was, is doing. Interleaving natively for image generation and, uh, vision and, and, and text understanding.

And I think like once that is better understood that is going to be better. That, that is the more deep learning pilled version of this, uh, the more GPU rich version of, of doing all this. I asked Yi Tay this question about Chameleon in his, in his episode.

He did not confirm or deny but I think he, uh, he would agree that that is the right way to do multimodality. And now that we have-- we're proving out that multimodality is valuable to people, basically all these half-ass measures around adapters-

Alessio48:41

Mm-hmm

Swyx48:41

... is going to flip to natively multimodal. To me that's what GPT-4o represents. It is the train from scratch-

Alessio48:48

Yep

Swyx48:49

... fully omni modal model whi- which is early fusion. So if you want to read that you should read the Chameleon paper basically. That's what is my, is my whole point.

Alessio48:56

And there was some of the Chameleon drama because the open model doesn't have image generation.

Swyx49:02

Yeah.

Alessio49:02

And then there were, uh, fine-tuning recipes.

Swyx49:05

It's so funny.

Alessio49:05

And then the leads were like, "No, do not follow these instructions to fine-tune image generation."

Swyx49:10

Yeah, yeah. Don't do that at all. Yeah, yeah. That's just really funny. I don't know what the, the-- I, I-- okay, so yeah, whenever image generation is concerned obviously because of the Gemini issue, you know, it's, it's very tricky for, for large companies to, to release that but they can remove it, say that they re- point out exactly where they remove it and let the open source community put it back in.

The last piece I had, which I kind of deleted, was, uh, ju-just a special mention, honorable mention of Gemma again with PaliGemma, which is one of the smaller releases from Google I/O. I think you went, right? So PaliGemma was, was mentioned in there?

Alessio49:43

Uh, I-

Swyx49:44

I don't know. It was, it was one of the-

Alessio49:45

Not, not on the... Yeah, yeah, one of the workshops, I think.

Swyx49:47

Very, very small release.

Alessio49:48

Yeah.

Swyx49:48

Um, but, uh, CoPaliGemma now is being talked a lot about as a, as a late fusion model for extracting structured text out of PDFs. Very, very important for business work-

Alessio49:57

Yeah, I know

Swyx49:58

... workhorses.

Alessio49:59

Yes.

Swyx49:59

Uh, so apparently it is doing better than Amazon Textract and all the others, uh, state-of-the-art, and it's just, it's a tiny, tiny model that does this, and it's really interesting. It's a combination of the Omar Khatab's, like, ColBERT sort of retrieval approach on top of a vision model, which I was severely underestimating PaliGemma when it came out, but, like, it's continuous to come up.

Like, there's a lot of trends, and it... Again, this is making a lot of progress here, just in terms of their, their applications in, in real world use cases. Like, these are small models, but they're very, very capable, and they're a very good basis to build things like CoPaliGemma.

Alessio50:31

Yeah, no, Google has been, it's been doing great. I think maybe a lot of people initially wrote them off, but between, um, you know, some of the Gemini Nano stuff, like, uh, Gemma 2, PaliGemma, we'll talk about some of the KV cache and context caching, um-

Swyx50:46

Uh, yeah. Yeah, that's, that's a RAG wars, yeah

Alessio50:48

... for it to be building. So there's a, a lot to like, and our friend Logan is over there now, so he's excited about everything they got going on, so yeah.

Swyx50:55

I think there's a little bit of a fight between AI Studio and Vertex, and what Logan represents is, is... So he's moved from DevRel to PM, and he was PM for the Gemma 2 launch. Vertex has this reputation of being extremely hard to use.

It's one, one reason why GCP has-

Alessio51:10

Yeah

Swyx51:10

... kind of fallen behind a little bit. And, and so AI Studio represents, like, the developer-friendly version of this, like the, the, the Netlify or Vercel to, to the AWS, right? And I think it's Google's chance to reinvent itself for this audience, for the AI engineer audience that doesn't want like five levels of auth IDs and org IDs and-

Alessio51:29

Right. Yeah

Swyx51:30

... policy permissions just to get something going.

Alessio51:32

True, true. Um, yeah, we wanna jump into RAG Ops Wars.

RAG Ops War51:36

Swyx51:37

What to say here? I, I think that what RAG Ops Wars are to me, like the, the tooling around the ecosystem, and I might need to actually rename this war, I think.

Alessio51:47

War renaming alert. What are we calling it?

Swyx51:49

Uh, LLMOs.

Alessio51:51

LLMOs.

Swyx51:52

Because it, it used to be when you're-- when, when the only job for AIs to do was chatbots.

Alessio51:58

Mm-hmm.

Swyx51:59

Then RAG matters-

Alessio52:00

Yeah

Swyx52:00

... then Ops matters. But now we need AIs to also write code. We also need AIs to work with other agents, right? That's not reflected in any, any of the other wars. So I think that just the, the whole point is what does an LLM plug into with the, with the broader ecosystem to be more capable than an LLM can be on its own?

Alessio52:19

Yep.

Swyx52:20

I just announced it, but yeah.

Alessio52:23

Yeah, yeah.

Swyx52:23

This is something I've been thinking about a lot. I-a-it's a blog post I've been working on. Uh, basically my tip to other people is if you wanna see where things are going, you go open up the ChatGPT GPT Creator.

Every single button on the GPT Creator is a potential startup.

Alessio52:35

Mm-hmm.

Swyx52:36

Um, EXA, uh, is, is for search, the knowledge RAG thing is for RAG.

Alessio52:40

Yeah, yeah.

Swyx52:40

The code interpreter-

Alessio52:41

We invested in E2B.

Swyx52:42

Yeah. Congrats. Is that announced? I don't know if

Alessio52:44

Oh, it's announced now.

Swyx52:45

It's announced now.

Alessio52:45

By the time this goes out, it'll be.

Swyx52:47

What, uh, briefly, what is E2B?

Alessio52:48

Uh, so E2B is basically a code interpreter SDK as a service, so you get a code interpreter to any model. Uh, they partner with Mistral to add that in. They have this open source, uh, Claude Artifacts clone using E2B.

It's a-- I mean, the amount of, like, traction that they've been getting in open source has been amazing. I think they went in, like, four months from, like, 10K to a million containers spun up, um, on the cloud.

So I mean, you told me this maybe, like, nine months ago, 12 months ago, something like that.

Swyx53:16

Yeah.

Alessio53:16

Uh, you were like, uh, what you literally just said, every ChatGPT plugin can be-

Swyx53:21

A business, a startup

Alessio53:22

... or like, uh, can be a business startup.

Swyx53:23

Yeah.

Alessio53:24

Um, and I think now it's more clear than ever than the chatbots are just kind of like the, uh, Band-Aid solution-

Swyx53:31

Yeah

Alessio53:32

... you know, before we build more, more comprehensive systems and, um, yeah. EXA just raised a Series A, uh, from Lightspeed, so, uh-

Swyx53:39

I tried to get you in on that one as well.

Alessio53:40

Yeah, yeah. No, I read that.

Swyx53:42

I'm trying to be a scout, man. I don't know.

Alessio53:46

Um, so yeah. This is a giving as a VC, early-stage VC, like giving capabilities to the models is like way more important than the actual LLM Ops, you know, the observability and like all these things. Like, those are nice, but like the way you build real value for a lot of the customers is like how can this model do more than just chat with me?

So running code, doing analysis, doing web search.

Swyx54:11

Hmm. I might disagree with you. I, I, I think it's com-- They're all valuable. They're all valuable.

Alessio54:15

Yeah, well-

Swyx54:16

They're all valuable. So I will disagree with you just on like, it, it, uh, I find Ops my number one problem right now building Smalltalk and building AI news, building anything I do, and I, I don't think I'm happy with all the Ops solutions I've explored.

Alessio54:29

Mm-hmm.

Swyx54:29

There are some 80-something Ops startups.

Alessio54:31

Right.

Swyx54:32

I nearly, you know, started one of them, but we'll briefly talk about this Ops thing-

Alessio54:36

Yeah, yeah, yeah

Swyx54:37

... then we'll get back to, go back to RAG. The central way I, I explain this thing to people is that all the model labs view their job as stopping by serving you their model over an API, right?

That is unfortunately not everything that you need-

Alessio54:48

Mm-hmm

Swyx54:49

... in order to productionize this API. So obviously there's all these startups. They're like, "Yeah, we are Ops guys. We've, we've done this for 30 years. Uh, we will now do this for AI." And the, the 80 of them show up, and they all raise money.

And the, the question is like, what do you actually need as, as like sort of a AI native Ops layer versus what did you just plug into Datadog?

Alessio55:09

Mm-hmm.

Swyx55:09

Right? Uh, I don't know if you ever, you have dealt with that because I'm not like a super Ops person, but I appreciate the importance of this thing. I think there's three broad categories, which is frameworks, gateways, and monitoring or tracing.

We've talked to like, uh, I interviewed Human- Humanloop in London and after. You've talked to a fair share of them. I've talked to a fair share of them.

Alessio55:29

Yep.

Swyx55:30

So the frameworks would be, uh, honestly, I won't name the startup But basically what this frame company was doing was charging me forty-nine dollars a month to store my prompt template, and every time I make an inference, it would F string-

Alessio55:42

Mm-hmm

Swyx55:43

... call the prompt template on some variables that I supply, and it's charging forty-nine dollars a month for unlimited storage of that.

Alessio55:49

Yeah.

Swyx55:50

It's absurd, but, like, people want prompt management tools. They want to interoperate between PM and developer. There's some value there. I don't know what the right price is.

Alessio56:00

Yeah.

Swyx56:00

There's some price.

Alessio56:01

Well, I was at... I'm sure I can share this. I was at the Grab office-

Swyx56:04

Mm-hmm

Alessio56:05

... um, and they also treat prompts as code, but they build their own thing to then import prompts and-

Swyx56:09

Yeah, I wanna check prompts into my code base as a developer, right? But maybe do, do you want it outside of the code base?

Alessio56:13

Well, I think it's like how do you... Well, you can have it in the code base, but, like, what's like the prompt file? What's like, you know, it's not just a string. How do you better-

Swyx56:21

There's string and model-

Alessio56:22

How do... Yeah

Swyx56:23

... and config.

Alessio56:24

Exactly. How do you pass these things? Uh, but I think, like, the problem with building frameworks is, like, frameworks generalize things that we know work, and, like, right now we don't really know what works. So-

Swyx56:36

Yeah, but some people have to try. You know, and the whole point of early stage is you try it before you know it works.

Alessio56:40

Yeah, but I think like the, the past, if you see the most successful open source frameworks that became successful businesses are frameworks that were built inside companies and then were kind of spun out-

Swyx56:51

Yeah, okay

Alessio56:51

... uh, as projects. So, uh, I think it's more of ordering.

Swyx56:53

So we're, we're more vertical, vertical-pilled instead of horizontal-pilled.

Alessio56:57

I mean, we try to be horizontal-pilled, right? And it's like, where are all the horizontal startups? So that's it.

Swyx57:02

There are a lot of them, they're just not that... They're not gonna win-

Alessio57:06

Well, my-

Swyx57:07

... uh, by, by, by themselves. By themselves.

Alessio57:08

Yes.

Swyx57:09

I think some of them will win by sheer excellent execution and, and then, but, like, the market won't pull them. They will have to pull the market.

Alessio57:16

Well, but that's the thing is like, uh, you know, take like Julius, right?

Swyx57:20

Yeah.

Alessio57:20

It's like, "Hey, why are you guys doing Julius like the same as Code Interpreter?" And yet they're pretty successful. A lot of people use it because they're like solving a problem, and then, uh-

Swyx57:30

They're more dedicated to it than Code Interpreter.

Alessio57:32

Exactly.

Swyx57:32

And so this-

Alessio57:33

So it's like I, I think-

Swyx57:34

Just take it more seriously than ChatGPT, you, you will.

Alessio57:36

I think people underestimate how important it is to be very good at doing something versus trying to serve everybody, uh, with some of these things. So yeah, I think that's a learning that a lot of founders are having.

Swyx57:48

Yes. Okay. So to draw out the ops worlds, uh, so it's a, it's a three circle Venn diagram, right?

Alessio57:52

Mm-hmm.

Swyx57:52

It's frameworks, it's gateways. So the only job of the gateway is to just be one endpoint that, uh, proxies all the other endpoints.

Alessio58:00

Mm-hmm.

Swyx58:00

Right? And, and it, it normalizes the APIs mostly to OpenAI's, uh, API just because most people started OpenAI. And then lastly, it's monitoring and tracing, right?

Alessio58:10

Mm-hmm.

Swyx58:10

So logging those things, understanding the latency, like P99 or whatever, and like the number of steps that you take. So LangSmith is obviously very-

Alessio58:17

Yeah

Swyx58:17

... very early onto this stuff. Uh, but so is Langfuse. So is... Ah, my God, like there's so many.

Alessio58:24

There's so-

Swyx58:24

Like, I'm sure like Datadog has some, like, uh-

Alessio58:26

Yeah, yeah, they have, yeah

Swyx58:26

... Weights & Biases has some.

Alessio58:28

Yeah, yeah.

Swyx58:28

Um, you know, there's, it's, it's very hard to, for me to, to choose between all those things. So I, as a, as a small team developer, want one tool that does all these things, and my discovery has been that there's so much specialization here.

They're like, everyone is like, "Oh yeah, we, we do this, but we don't do that. For the other stuff, we recommend these, these two other friends of ours." And I'm like, why am I integrating four tools when I just need to have one?

They, they all, they all the same thing.

Alessio58:51

Mm-hmm.

Swyx58:52

Um, that, that is my current frustration. The obvious frustration solution is I build my own, right? Which is-

Alessio58:59

Yeah, yeah.

Swyx58:59

You know, we have 14 standards, now we have 15.

Alessio59:00

Yeah.

Swyx59:01

So it's just a very messy place to be in. I, I wish there was a better solution to recommend to people, because right now I cannot clearly recommend things.

Alessio59:09

Yeah. I think the biggest change in this market is like latency is actually not that important anymore. Like we lived in the past 10 years in a world where like 10, 15, 20 milliseconds made a big difference. I think today people will be happy to trade 50 milliseconds to get higher quality output-

Swyx59:26

Mm

Alessio59:26

... from a model.

Swyx59:27

Mm.

Alessio59:27

So, but still all the tracing is all like, how long did it take? Like, what's the thing? Instead of saying, is this quality good for this output? Like, should you use another model? Like we're just kind of taking what we did with cloud and putting it in LLMs instead of saying what actually matters when it comes to LLMs, what you should actually monitor.

Like I don't really care what my P99 is if the model is crap, right? It's like also like I don't own most of the models, so it's like this is the GPT-4 API performance. It's like, okay, am I going into a moment?

It's like I can't do anything about it, you know? So I think that's maybe why the value is not there. Like, you know, am I supposed to pay 100K a year like I pay to Datadog or whatever to tell me, for have you tell me that GPT-4 is slow?

It's like- ... you know, and just not, uh, I don't know.

Swyx1:00:14

I agree. It's, it's challenging there. Um, okay. So the, the last piece I'll, I'll mention is briefly MLOps is still real, and I think LLMOps or whatever you call this AI engineer ops, the ops layer on top of the LM layer, uh, might follow the same evolution path as the MLOps layer.

And so the most impressive thing I've seen from the MLOps layer is from Apple. When they announced Apple Intelligence, they also announced Telaria, which is their internal MLOps tool, which way you can profile the performance of each layer of a transformer.

Alessio1:00:43

Mm-hmm.

Swyx1:00:43

And you can AB test like 100 different variations of different quantizations and stuff and pick the best performance. And I could see a straight line from there to like, okay, I want this, but for my, my AI engineering ops.

Like I, I want this level of clarity on like what I do and, uh, there's a lot of internal engineering within, within these big companies who take their ML training very seriously, and I see that also happening for AI engineering as well.

And let's briefly talk about RAG and context caching maybe, unless you have o- other like LLM OS stuff that you're excited about.

Alessio1:01:14

Um, LLM OS stuff I'm excited about.

Swyx1:01:17

Any cool LLM OS stuff.

Alessio1:01:17

No, I think, I think that's really a lot of it is like move beyond being, uh, observability or like help for like making the prompt call and like actually being an LLM OS. You know? I think today it's mostly like LLM rails, you know?

Like there's no OS, but I think like actually helping people build things. That's why, you know, if you look at xAI 2b, it's like that's the OS, you know? Those are kind of-

Swyx1:01:40

Yeah

Alessio1:01:40

... like the OS primitives that you need around it.

Swyx1:01:43

Yeah. Okay. So I'll mention a, a couple things then. One layer I'm- I've been excited about publicly, but I haven't talked about it on this podcast, is memory databases-

Alessio1:01:52

Mm-hmm

Swyx1:01:52

... or memory layers of-- on top of vector databases. The, the vogue thing of last year was vector databases, right? Everybody had a vector database, uh, company. And I think the insight is that vector databases are too low level.

Like, they're, they're not very useful out of the box. They do cosine similarinar- similarity matching and retrieval, and that's about it. We'll briefly maybe mention here BM42, which, uh, was this whole debate between Vespa and who else? Uh-

Alessio1:02:14

Wand

Swyx1:02:15

... Quadrant, Qdrant, and, um, I think a couple other companies also chipped in. But it was ma-ma-ma-mainly a very, very public and ugly Twitter battle between, uh, benchmarking for, for databases. And the history of benchmarking for databases goes as far back as Larry Ellison and Oracle and, and all that.

It's just very cute to see it happening in the vector database space. Uh, nothing-- some things don't change. But on top of that, like, I think one of the reasons I put vector databases inside of these wars is in order to grow, the vector databases have to become more frameworks.

In order to grow, the ops companies have to become more frameworks, right? And then the framework companies have to become ops companies, which is what LangChain is. So one element of the vector database is growing. I've been looking for what the next direction of vector database is growing is, is memory.

Long conversation memory. I have on me this, um, uh, Bee, which is one of the personal AI wearables. I'm also getting the Limitless, uh, personal AI wearable, which is like I j- I just want it to record my whole conversation and just repeat back to me or let me, let me find-- augment my memory.

As a, as a-- I'm sure Character.AI has some version of this.

Alessio1:03:17

Yeah.

Swyx1:03:17

Like, everyone has conversation memory that is different from factual memory. And right now, vector database is very oriented towards factual memory, document retrieval, knowledge-based retrieval. But it's not the same thing as conversation retrieval, where I need to know what I've said to you, what I said to you yesterday, what I said to you a year ago or three years ago.

And that's, that's a different nature of retrieval, right? Um, so there's a, a-- at, at the conference that we ran, GraphRAG was a, was a lot of, uh, focus for people-

Alessio1:03:43

Mm-hmm

Swyx1:03:43

... the marriage of knowledge graphs and RAG. I think that this is commonly a trap in ML that people are like, they discover that graphs are a thing for the first time. They're like, "Oh yeah, everything's a graph.

Like, the future is graphs," and then nothing happens. Very, very common. This happened like three, four times in, in the industry's past as well. But maybe this time is different.

Alessio1:04:00

Maybe. Unless.

Swyx1:04:02

Unless. Unless. So this is a fun-- this is why I'm not an investor. Like, you have to get-

Alessio1:04:09

Yeah, yeah

Swyx1:04:09

... the time the, the this time is different because no ideas are really truly new.

Alessio1:04:14

Mm-hmm.

Swyx1:04:14

But some-sometimes this time is different.

Alessio1:04:18

Maybe. Maybe.

Swyx1:04:19

And so memory databases are one form of that, where, like, they're focused on the problem of long-form memory for, uh, for agents, for assistants, for chatbots and all that.

Alessio1:04:28

Yeah.

Swyx1:04:29

I, I, I definitely see that coming. There were some funding rounds that I can't really talk about in this sector, and I've seen that happen a lot. Um, yeah, I have one more category in LLMs, but, uh, any comments on memory?

Alessio1:04:38

Yeah, no, I, I think that makes sense to me that moving away from just semantic similarity I think is the most important because people use the same word with very different meanings, especially when talking, you know? When writing it's different, but yeah.

Swyx1:04:50

Yeah. The other direction that vector databases have gone into, which, um, LanceDB presented at my conference, was multimodality. So, uh, Lance, uh, Character.AI uses LanceDB for multimodal embeddings. That's just a minor difference. I don't think that's like a quantum leap in terms of what a vector database does for you.

The other thing that I see in LLMs world is, is mostly, um, the evolution of like, uh, just the, the ecosystem of agents, right? The, the, the agents talking to other agents and coordinating with other agents. So I interviewed Graham Newbig at, uh, I-Iclea, and he since announced that they are pivoting Open, Open Devin or broadening Open Devin into All Hands AI.

I'm not sure about that name, but, uh, it, it is the-- it is one of the three LLMOS startups that got funded in the past two months, uh, that I know about, and maybe you know more. They're, they're all building like this ecosystem of agents talk- working with other agents and all this, all this tooling, uh, for, for agents.

To me, makes more sense. It, it is probably the biggest thing I missed in, in doing the Four Wars, the need for startups to build this ecosystem thing up, right? So the big categories have been taken. Search, done.

Code interpreter, done. There's a long tail of others, right? So memory is emerging, then there's like other stuff.

Alessio1:06:00

Mm-hmm.

Swyx1:06:01

Uh, and so they're, they're focusing on that. Like, so to me, uh, browser is slightly different from search, and BrowserBase is, is another company I invested in that is focused on that, but they're not the only one in that category by, by, by any means.

Agent Ecosystem1:06:01

Swyx1:06:13

I used to tell people, go to the Devin demo and look at the four things that they offer, and just say each of the, each of those things is a startup.

Alessio1:06:18

Mm-hmm.

Swyx1:06:19

Devin, since then, they spoke at the conference as well. Scott, uh, was super nice to me and, and, uh, actually gave me some personal time as well. They have an updated chart of their plans. You look at their plans, they have like sixteen things.

Each of those things is a potential startup now.

Alessio1:06:31

Mm-hmm.

Swyx1:06:31

And that is the LLMOS. Everyone's building towards that direction because they need it to do what they need to do as a, as an agent. If you believe in the agent's future, you need all these things.

Alessio1:06:40

Yeah. You think the agent OS is its own company? Do you think it's an open standard? Do you think-

Swyx1:06:48

I would love it to be open standard. The reality is that people want to own that standard, so we have-- we actually wound down the AI Engineer Foundation w-where the, where the first project was the Agent Protocol, which E2B actually donated to the foundation because no one's interested.

Everyone wants to be VC-backed-

Alessio1:07:04

Right

Swyx1:07:04

... so then they wanna own it, right? So there's just-- it's too early to be open source. This-- the-- people will keep this proprietary, and more power to them. They need to make it work. They need to make revenue before all the other stuff can happen.

Alessio1:07:15

Yeah. I'm really curious. You know, we're investors in a bunch of agent companies. None of them really care about how to communicate with other agents. They're so focused internally, you know? But I think in the future, you know, I talked about this on-

Swyx1:07:27

I see. You're, you're talking about agent to a- other external agents.

Alessio1:07:30

Yeah. So I think-

Swyx1:07:30

I'm not talking about that yet

Alessio1:07:31

... yeah, I wonder when-- like, because that's where the future is going, right? So today it's like intra-agent-

Swyx1:07:37

I see, I see

Alessio1:07:37

... connectivity. You know? At some point it's like, well, it's not like somebody-- I'm selling into a company and the company already uses Agent X For that job, I need to talk to that agent, you know? But I think nobody really cares about that today, so I think that's usually-

Swyx1:07:51

Yeah. So, uh, I think that, that layer right now is OpenAPI. Just-

Alessio1:07:56

Yeah, yeah, yeah. Exactly

Swyx1:07:57

... give me a RESTful protocol-

Alessio1:07:58

Yep

Swyx1:07:58

... I can, I can interoperate with that. RESTful protocol only does request response, so then the next layer is something I have worked on, which is long-running request response, which is workflows, which is what Temporal was supposed to do before, let's just say, management issues.

Yeah, but like, you know, RPC or some kind of... You know, I, I think the, the dream is, uh, and this is one of the, the, my problems with the LLM OS concept, is that do we really need to rewrite every single thing for AI native use cases?

Shouldn't the AI just use the, these things, these tools the same way as humans use them? The reality is for now, yes, they, they, they need specialized APIs. In fu- in the distant future when these things cost nothing, then they-

Alessio1:08:38

Right. Yeah, yeah, yeah

Swyx1:08:38

... can use it the same way as humans does, but right now they need specialized interfaces. The layer between agents ideally should just be English, you know, like the same way that we, we talk. But like it-- English is too underspecified-

Alessio1:08:51

Yeah, yeah, yeah

Swyx1:08:51

... unstructured to, to make that happen, so.

Alessio1:08:52

Yeah.

Swyx1:08:53

So, yeah.

Alessio1:08:53

It's interesting because we talk to each other in English, but then we both use tools to do things-

Swyx1:08:59

Yes

Alessio1:08:59

... to then get the response back.

Swyx1:09:00

For those people who want to dive in a little bit more, I think AutoGen, I would, I would definitely recommend, uh, looking at that. CrewAI. There are established frameworks now that are working on inter agents communication layers to coordinate them, and not necessarily externally from company to company-

Alessio1:09:15

Yeah, yeah

Swyx1:09:15

... just internally as well.

Alessio1:09:16

Mm-hmm.

Swyx1:09:16

Between if you have multiple agents farming out work to do, do different things, you're gonna need this anyway. And, uh, I don't think it's that hard. It's, it's they are using English, they're using s- some, some mix of English and structured output.

And yeah, if you have a better idea than that, let us know.

Alessio1:09:31

Yeah, we're listening.

Swyx1:09:34

So that's the Four Wars discussion. I think I want to leave some discussion time open for miscellaneous trends that are happening in the industry that don't exactly fit in the Four Wars or, or are a layer above the Four Wars.

Commoditization1:09:34

Swyx1:09:46

So the first one to me is just this trend of open source. Obviously, this overlaps a lot with the GPU poor thing, but I want to really call out this depreciation thing that I've been working on. Like, I, I, I do think it's probably one of the, the bigger thesis that I've come-- that I've had in the past month, which is that we now have a rough idea of the dep- depreciation schedule of, uh, this, this sort of model spend.

And, yeah, I basically drew a chart. I'll link it in the show notes, but I drew, drew a chart of the price efficiency frontier of as of March, April 2024, and then I had listed all the models that list-- that sit within that frontier.

Haiku was, was the best cost per intelligence at that point in time. And then I did the same chart in July, two days ago, and the whole thing has moved, and Mistral is like deprecating their old models that, that used to be in the old frontier.

It is so shocking how predictive and tight this band is. Very, very tight band, and the whole industry is moving the same way. And it's roughly one order of magnitude drop in cost for the same level of intelligence every four months.

My previous number for this was one order of magnitude drop in cost every 12 months.

Alessio1:10:56

Mm-hmm.

Swyx1:10:56

But the timeline has accelerated because GPT-3 took about a year-

Alessio1:11:00

Yeah, yeah

Swyx1:11:01

... to, to drop, uh, order of magnitude. But now GPT-4, it's really crazy. I don't know what to say about that. But I just wanna-

Alessio1:11:07

Do, do you think GPT Next and Claude 4 push it back down because they're coming out with higher intelligence, higher cost? Or is it maybe like the-

Swyx1:11:17

It, it should-

Alessio1:11:18

... the timeline's going down because new frontier models are not really coming out at the same rate?

Swyx1:11:23

Interesting. I don't know. That's a, that's a really good question. Wow. I'm stumped. I, I don't have-

Alessio1:11:27

You're like, "Wow, you got a good question on this."

Swyx1:11:29

Yeah, I don't have a-- I don't, I don't, I don't have an answer. No, and I mean, you have a lot of good questions.

Alessio1:11:32

Yeah, yeah.

Swyx1:11:32

But like, uh, I thought I had solved this, and then now you came along with it. The first response is something I haven't thought about. Uh, yeah. Yeah, so there's, there's two directions here, right? When, when the cost of frontier of models are going up, potentially, like SB 1047 is gonna-

Alessio1:11:44

Right. Yeah, yeah, yeah

Swyx1:11:45

... make it illegal to train even larger models for as what I, I, I think the opposition has increased enough that it's not gonna be a real-

Alessio1:11:51

Mm-hmm

Swyx1:11:51

... real concern for people. But I think every lab basically needs a small, medium, large play, and like we said in the, the sort of model deployment framework, first you choose-- you pursue capability, then you pursue generalization, then you pursue efficiency.

And, and what we're talking about here is, is efficiency.

Alessio1:12:08

Yeah, yeah.

Swyx1:12:08

Right? Like that now we care about efficiency. The... This definitely one of the emerging stories of the year that has happened is efficiency matters for 4o, 4o Mini, and 3.5 Sonnet in a way that in January we-- nobody was talking about.

Alessio1:12:21

Mm-hmm.

Swyx1:12:22

And that's great.

Alessio1:12:23

Yeah.

Swyx1:12:23

Regardless of GPT Next and Claude 4 or whatever, or Gemini 2, we will still have efficiency frontiers to, to pursue. And it seems like doing the higher capable thing creates the synthetic data for us to do the efficiency, efficient thing, and that means lifting up the...

Like, I, I had this difference chart between Llama 3.0-8B, Llama 3.0-70B, versus their 3.1 differences.

Alessio1:12:48

Mm-hmm.

Swyx1:12:48

And the 8B had the most, uh, uplift across all the benchmarks. Right? It makes sense. You're, you're training from the 4 or 5B.

Alessio1:12:55

Yeah.

Swyx1:12:55

You're distilling from there, and, and it's gonna have the, the biggest lift up. So the best way to train more efficient models is to train the large model .

Alessio1:13:01

Right. Yeah, yeah.

Swyx1:13:02

And then you can distill, uh, down to the rest it. So this is fascinating on-- from the investor point of view. You're like, okay, you're worried about picks and shovels. You're worried about investing in foundation model labs. Um, and then, and that's a, that's a matter of opinion.

I, I do think that some foundation model labs are worth investing in because they, they do pay back very quickly. I think for engineers, the question is, what do you do when you know that your base cost is going down-

Alessio1:13:23

Mm-hmm

Swyx1:13:24

... an order of magnitude every four months? How do you, how do you make those assumptions? And I don't know the answer to that. I'm just posing the question. I'm calling attention to it.

Alessio1:13:31

Yeah.

Swyx1:13:31

Because I think that cognition burning like rumors is, I don't know n- nothing from Scott. I haven't talked to him at all about this, even though he's, he's very friendly. But they did that, they got the media attention, and now the, the cost of intelligence is going down, and it will be economically viable tomorrow.

In the meantime, they have a crap ton of value, uh, from, from user data, and a crap ton of value from some media exposure. And I think that the, the correct stunt to pull is to pull-- is to, like, make economically non-viable startups-

Alessio1:13:59

Right

Swyx1:13:59

... now and then wait .

Alessio1:14:01

Yeah.

Swyx1:14:02

But was, honestly, basically, I'm basically advocating for people to burn VC money .

Alessio1:14:05

Yeah. No, they, they can burn my money all they want if they're building something useful. I think the big problem, not a problem, but the price of the model comes out, and then people build on it. And then there's really no, uh...

The model providers don't really have a lot of leverage on, like, keeping the price high. You know? They just have to bring it down because the people downstream of them are not making that much money with them, you know?

And I wonder what's gonna be the model where it's like, "This model's so good, I'm not putting the price down," you know? Like, if GPT-4o was, like, amazing and was actually solving a lot of-- like, creating a lot of value downstream, people would be happy to pay.

I think people today are not that happy with the models, you know? Like, they're good, but, like, I'm not paying that much because I'm not really getting that much out of it. Like, we have, um, this AI center of excellence with a lot of the Fortune 500 groups, and there are people saving $10 million, $20 million a year, like, with these models doing boring stuff, you know, like document translation, things like that.

But nobody's making $100 million. Nobody's making $150 million. So, like, the prices just have to go down to match. But maybe that will change-

Swyx1:15:11

Yeah

Alessio1:15:11

... at some point. You know?

Swyx1:15:12

Yeah. I, I always mention temperature two use cases, right? Like, yeah, we're-- those are temperature zero use cases-

Alessio1:15:17

Exactly

Swyx1:15:17

... where you need precision, you need creativity. What are the cases where hallucinations is a feature, not a bug, right? So we're the first podcast to interview Websim-

Alessio1:15:24

Mm-hmm

Swyx1:15:24

... and I, I'm pretty, I, I'm still pretty positive about the generative part of AI. Like, we f- we took generative AI, and we used it to do reg.

Alessio1:15:31

Right.

Swyx1:15:31

You know, like

we have an infinite creativity engine. Let's go, let's go do, do more of that. Uh, yeah. So we'll hopefully do more episodes there. Uh, you have some stuff on, on agents you wanna...?

Alessio1:15:41

Yeah, no, I think this is something that we talked a lot about in, um... You know, we wrote this post months and months ago about, uh, shifting from software as a service to service as a software, and that's only more true now.

Service as Software1:15:41

Alessio1:15:52

I think, like, most companies that are buying AI tooling, they want the AI to do some sort of labor for them, and that's why the picks and shovels kind of disinterest maybe comes from a little bit. Most companies do not wanna buy tools to build AI.

They want the AI, and they also do not want to pay a lot of money for something that makes employees more productive because the productivity gains are not accruing to the companies. They're just accruing to the employees. You know, people work less, have longer lunch breaks because they get things done faster.

But most companies are not making a lot more money by making employees productive. That's not true for startups. So if you look at most startups today in AI, like, they're much smaller teams compared to before. Versus agents, we have companies like, you know, Brightwave, which we had on the podcast.

You're selling labor, which is something that people are used to paying on a certain pay scale. So when you're doing that, you know, if you ask Brightwave, they don't have it public, but, like, they charge a lot of money, more than you would expect because hedge funds and, like, investment banking, investment advisors, they're used to paying a lot of money for research.

It's like the labor. They don't even care that you use AI. They just want labor to be done.

Swyx1:16:57

I'll mention one, one pushback. But, uh-

Alessio1:16:59

Yeah, yeah

Swyx1:16:59

... as, as a hedge fund, we used to pay for analyst research out of our brokerage cost-

Alessio1:17:03

Mm-hmm

Swyx1:17:04

... and not read them. And that, that, to me, that's, that's my-

Alessio1:17:06

Yeah

Swyx1:17:07

... risk of, of Brightwave. But, but you know, I-

Alessio1:17:08

No, but I, I think the-

Swyx1:17:10

As a consumer of research, I'm like-

Alessio1:17:11

Well, if, if we wanna go down the- ... the rabbit hole, there's a lot of pressure on funds for, like, a OpEx efficiency, so th- there's not really capture, uh, researchers anymore in most funds, and, like, even the sellside research is like-

Swyx1:17:23

I see

Alessio1:17:23

... not that good.

Swyx1:17:23

So, so taking them u- from in-house to external thinking.

Alessio1:17:26

Yeah.

Swyx1:17:26

Yeah, yeah.

Alessio1:17:26

Exactly.

Swyx1:17:26

That makes sense. Yeah.

Alessio1:17:27

Um, so yeah, we s- you know, we had Dropzone that does security a-analysis. Same people are used to paying for managed security or, like, outsource SOC analysts. They don't wanna buy a AI tool to make the security team more productive.

So, um-

Swyx1:17:40

Okay. And what specifically does Dropzone do?

Alessio1:17:42

They do, uh, SOC analysis. So not SOC like the compliance, but it's like when you have security alerts, how do you investigate them? So large enterprises, they get like thousands of phishing email, and then they forward them to IT, and it's IT or security person, the tier zero has to go in and say, "That's a phishing email.

That isn't. That isn't." So they have an agent that does that. So the cost to do, like, for a human to do the analysis at the rate that they get paid, that's like $35 per alert. Dropzone is like $6 per alert.

Swyx1:18:10

Yeah.

Alessio1:18:10

So it's, it's a very basic economic analysis for the company whether or not they wanna buy it. It's not about, "Is my an-analyst gonna have more free time?" Like, "Is it more productive?" So selling the labor is, like, the story of the market right now.

Um-

Swyx1:18:25

My version of this is I should start a consulting services today-

Alessio1:18:28

Mm-hmm

Swyx1:18:28

... and then slowly automate myself, my, my employees out of a job, right?

Alessio1:18:32

Mm-hmm.

Swyx1:18:33

Is that fundable?

Alessio1:18:34

Is that fundable? That's a good question. I think whether or not depends how big you want it to make.

Swyx1:18:38

This is a services company, basically.

Alessio1:18:40

Yeah. That's, I mean, that's what ... I know now, now it's maybe not as good of an example, but Crowdstrike started as a-

Swyx1:18:46

As a services company

Alessio1:18:46

... um, security research. Um-

Swyx1:18:49

Yeah, I mean, it's still one of the most successful companies of all time.

Alessio1:18:51

Yeah. Yeah, yeah.

Swyx1:18:51

Maybe, uh... Yeah. It's in-interesting model. I'm always checking my, my biases there. Um, anything else on the, the agents, uh, side of things?

Alessio1:18:58

No, that's really something that people should spend more time on. It's like what's the end labor-

Swyx1:19:03

Yeah

Alessio1:19:03

... that I'm building? Because, you know, sometimes when you're being too generic and you wanna help people build things like Adept. Like Adept, you know, David was on the podcast, and he said they were sold out of things, but they're kinda like working-

Swyx1:19:14

And then he sold out himself.

Alessio1:19:16

Yeah. It's like they're working, they're working with each company, and the company has to invest the time to build with them-

Swyx1:19:22

Yeah

Alessio1:19:22

... the Adept agent.

Swyx1:19:23

You need more hands-off.

Alessio1:19:24

Exactly.

Swyx1:19:24

Yeah.

Alessio1:19:25

Um, so and that's more verticalized.

Swyx1:19:27

Yeah, yeah.

Alessio1:19:27

That's the way I think.

Swyx1:19:27

I'll sha- I'll shout out here Jason Liu. He was on, also on our podcast and spoke at the conference. Uh, he has a, this idea of like it's, it's reports, not RAG. You want to, you want things to produce reports 'cause reports can actually get consumed.

Alessio1:19:37

Yeah.

Swyx1:19:38

Uh, RAG is still too much work, still too much chat botting. I'll briefly mention that, uh, new benchmarks I'm, I'm thinking about. Um, I, I think you need to have, um, every-everyone in studying AI research Understanding the progress of AI and found- and foundation models needs to have in mind what is next after MMLU.

I have 10 proposals, most of them, uh, half of them come from the Hugging Face episode. So everyone's loving Clementine. Uh, I want her back on. She, she, she was amazing and very, very charismatic even though she made us take down the YouTube.

New Benchmarks1:19:54

Swyx1:20:06

Uh , but, uh, MUSR for multi-step reasoning-

Alessio1:20:09

Yeah

Swyx1:20:09

... math for math, IFEV for instruction following, BigBench Hard, and it, uh, code we're now getting to the area that the Hugging Face leaderboard does not have.

Alessio1:20:17

Mm-hmm.

Swyx1:20:18

And I'm, I'm considering making my own 'cause I, I, I care about this so much. So MBPP is the current one, uh, that is post-human eval, 'cause human eval is widely known to be saturated. And PsyCode is like the, the newest one that I would point people to.

Context utilization, we had Mark from Gradient on talk about Ruler, but also zero scores in Inf-Infinite Bench were the two that Llama 3 used, used instead of Ruler. But basically something that's a little bit more rigorous than needle in a haystack, uh, that is something that, that people need.

Then you have function calling. Here I think Gorilla, API Bank next is, uh, pretty, pretty consensus. I, I've got nothing there apart from, yeah, like all models need, need something like this. Vision now is like multi- like multimodality the, the vision is the most important.

Um, I think like ViVEval is actually the, the state-of-the-art here. I, you know, open to, open to being corrected and then multilinguality. So basically like these are the 10 directions, right? Uh, post-MMLU here are the frontier capabilities. If you're developing models or if you're, if you're encountering a new model, evaluate them on all these elements, and then you have a good sense of how state-of-the-art they are and what you need them for in terms of applying them to your use case.

So I just want to get that out there.

Alessio1:21:19

Yeah. And we had the RKGI thing. How do you think about benchmarking for, you know, everyday thing or like benchmarking for something that is maybe like a hard to reach goal?

Swyx1:21:30

Yeah, this has been a debate, uh, for, uh, that's obviously very important and probably more important for product usage, right? Here I'm talking about benchmarking for general model evals and then there's a, there's a schism in the AI engineering community or criticism in the AI engineering community that did not care about, enough about product evals.

So Hamul Hussein, uh, led that and I, I had a bit of disagreement with him, but I, I acknowledge that I, I think that is important. There was an oversight in my original AI engineer post. Uh, so the job of the AI engineer is to produce product specific evals for your use case and there's no way that these general academic benchmarks are going to do that because they don't know your use case.

Alessio1:22:06

Yeah.

Swyx1:22:06

It's not important. They will correlate with your use case and that is a good sign, right? These are very, very rigorous and thought through. So you want to look for correlates, then you want to look for specifics and that's something that only you can do.

So yeah, RKGI will correlate with IQ. It's, it's an IQ test.

Alessio1:22:22

Yeah, yeah, yeah.

Swyx1:22:22

Right? How well does IQ tests correlate to job performance? 5%, 10%.

Alessio1:22:27

Mm-hmm.

Swyx1:22:27

Not nothing, but not everything.

Alessio1:22:29

Yeah, yeah.

Swyx1:22:29

And so it's important.

Alessio1:22:30

Anything else?

Swyx1:22:31

Super intelligence. We, we can... You know, we, we try not to talk about safety. Uh, my, my favorite safety joke from our dinner is that, you know, if you're worried about agents taking over the world and you need a button to take them down, just install CrowdStrike-

Alessio1:22:43

Yeah

Swyx1:22:43

... on all, every agent and you have a button that has just been proved at the largest scale in the world to disable all agents, right? So that to save super intelligence you should just install CrowdStrike. That's what all AI subscribers should do.

Alessio1:22:57

That's funny, except for the CrowdStrike people. Awesome, man. This was great. I'm glad we did it. I'm sure we'll do it more regularly-

Swyx1:23:03

We should do more

Alessio1:23:04

... now that you're out of visa jail.

Swyx1:23:05

Yeah. Yeah. I, I think, uh, you know, AI News is surprisingly helpful for doing this. Like, uh-

Alessio1:23:10

Yeah

Swyx1:23:10

... yeah, I, I, I had no idea when I started. I just, I just thought I, I needed a thing to summarize Discords, but now it's becoming a proper media company. Like, yeah, 1,000 people sign up every month.

Uh, it's growing.

Alessio1:23:22

Cool. Thank you all for listening.

Swyx1:23:24

Yeah.

Alessio1:23:24

See you next time.

Swyx1:23:25

Bye.