# 2024 Year in Review: The Big Scaling Debate, the Four Wars of AI, Top Themes and the Rise of Agents

Latent Space · 2025-01-01

<https://addtry.com/a5079620-c52c-42be-813e-c6e6a78e8fda>

In their 100th episode, hosts Alessio and Swyx recap 2024 in AI, arguing that pre-training scaling has hit a wall—backed by Ilya Sutskever and others at NeurIPS—and that inference-time compute (o1, o3) is the new frontier. They dissect the "four wars": data quality (lawsuits vs. synthetic data), GPU haves vs. have-nots (with the middle class dying), multimodality (Sora, Veo 2, Gemini 2.0's native image output), and the LLM OS/agents stack (LangChain, E2B, memory). Market share shifted from OpenAI's 95% to 50-75% as Anthropic and Gemini gained ground; prices dropped ~3 orders of magnitude for same ELO. The episode predicts 2025 as the year agents finally enter production, driven by models like o1 and tools like Devin, and warns that AI will set the skill floor for roles.

## Questions this episode answers

### What was the 'big scaling debate' at NeurIPS 2024 and what did Ilya Sutskever say about pre-training?

At NeurIPS 2024, Ilya Sutskever declared that pre-training has hit a wall because data isn't scaling as fast as compute, shifting focus to inference-time compute. Swyx notes that other researchers like Noam Brown echoed this view. The new paradigm recasts scaling from larger models to spending more compute during inference, which is easier for big labs to monetize.

[7:01](https://addtry.com/a5079620-c52c-42be-813e-c6e6a78e8fda?t=421000)

### What are the 'Four Wars of AI' according to the Latent Space podcast?

Swyx and Alessio outline four AI wars: 1) the Data Quality War between IP owners (New York Times, Scarlett Johansson) and synthetic data; 2) the GPU Rich vs. GPU Poor war, where ultra-rich labs scale massive clusters while others go GPU-free; 3) the Multimodality War between specialized models (e.g., ElevenLabs) and 'god models' like Gemini; 4) the LLM Ops War over orchestration and memory.

[32:21](https://addtry.com/a5079620-c52c-42be-813e-c6e6a78e8fda?t=1941000)

### How did AI model market share change in 2024 between OpenAI, Anthropic, and Google?

According to Ramp data cited by Swyx, OpenAI's production traffic share dropped from ~95% in December 2023 to an estimated 50–75% by late 2024. Anthropic gained ground with Claude 3 and 3.5 Sonnet, while Gemini Flash captured about 50% of OpenRouter requests by being nearly free, sparking a price war at the low end.

[10:43](https://addtry.com/a5079620-c52c-42be-813e-c6e6a78e8fda?t=643000)

### How much did the cost of AI inference fall in 2024?

Swyx shows that the cost of equivalent intelligence (measured by LMSYS ELO) fell about three orders of magnitude in 2024. For instance, GPT-4-level quality cost $40 per million tokens at the start of 2024 but dropped to as low as $0.075 per million tokens with models like Amazon Nova, driven by competitive cuts from labs like Google and Anthropic.

[15:48](https://addtry.com/a5079620-c52c-42be-813e-c6e6a78e8fda?t=948000)

## Key moments

- **[0:00] Intro**
  - [1:42] Gartner places AI engineering at the peak of the hype curve, notes Swyx.
  - [3:15] The AI world is still research heavy but will invert to engineering heavy as it moves into production.
  - [4:05] "It's this very strange mix of keeping on top of research while not being a researcher, and then putting that research into production." — Swyx
- **[6:35] Scaling Walls**
  - [6:49] Swyx found John Frankel to debate against scaling; Frankel assumed a character to argue 'we've hit a wall' despite believing the opposite.
  - [7:46] Ilya, John Frankel, and Noam Brown all declared that pre-training scaling has hit a wall, pointing to inference time compute as the new frontier.
  - [10:11] Tier zero frontier AI models are a three-horse race between Gemini, Anthropic, and OpenAI, says Swyx.
  - [10:54] OpenAI held 95% market share of production AI traffic in December 2023, per Ramp estimates.
  - [12:10] Gemini Flash's free tier offers a billion tokens per day, capturing 50% of Open Router requests.
  - [14:24] Llama 405B is too slow and expensive to run, and the gap between open source and closed models is widening, not narrowing.
- **[32:21] Data War**
  - [32:41] Scarlett Johansson joined the list of parties suing AI companies, adding to the data quality war.
  - [33:04] Scale AI vs synthetic data community: Scale published a paper claiming synthetic data doesn't work; some see a conflict of interest.
  - [37:07] Cosine fine-tuned OpenAI's 4o model to beat o1, casting doubt on the moat of reasoning-specific models.
- **[38:21] GPU War**
  - [40:47] The GPU smiling curve: startups at the edges (close to hardware or customers) thrive, those in the middle fail.
  - [44:00] Suno reached $20M ARR as a GPU-poor startup running on Model; Bolt hit $20M as a cloud wrapper.
  - [45:12] Google acquired Character.AI for $2.7B, far above Swyx's earlier $1B price tag suggestion.
- **[45:31] Multimodality War**
  - [47:29] Gemini 2.0 introduced native image output and editing, promising a simpler workflow than chaining small models.
  - [48:49] Midjourney's product ease-of-use and workflow trumps raw pixel-level model quality, argues Alessio.
  - [51:53] DeepMind's Genie and Video Poet give them a years-long advantage in world modeling over OpenAI.
- **[52:36] LLMOS War**
  - [55:23] AutoGPT's initial promise of generality led to massive GitHub stars but low actual usage.
  - [59:47] Current AI memory products only provide explicit summarization, not implicit preference extraction.
  - [1:04:07] Anthropic's Model Context Protocol includes a 300-line memory implementation, undermining dedicated memory startups.
- **[1:05:33] Benchmarks**
  - [1:05:52] AI benchmarks shifted this year from MMLU to Sweetbench, LiveBench, AIME, and Frontier Math.
  - [1:07:07] GPQA benchmark declared dead at NeurIPS; Sweetbench reached 50% pass rate, with hopes for 80% next year.
- **[1:09:35] Capabilities**
  - [1:15:48] The inference price war drove a three orders of magnitude improvement in cost for equivalent intelligence over one year.
  - [1:16:59] For the same LMSYS ELO, model pricing dropped from $40 to $0.50 per million tokens in about a year.
- **[1:22:48] Monthly Rewind**
  - [1:25:47] Devin's launch in March created huge hype, but took 9 months to reach general availability at $500/month.
  - [1:27:34] GPT-4o's launch included a Sky voice resembling Scarlett Johansson, leading to controversy and its removal.
  - [1:30:39] Ilya Sutskever raised $1B for Safe Superintelligence (SSI); Dan Gross became CEO, focusing on a single product.
  - [1:33:26] OpenAI's o1 (Strawberry) launched in September, quickly rolling out to ChatGPT and the API.
  - [1:35:02] OpenAI Canvas launched in October, competing with Google Docs by offering an AI-integrated editing environment.
  - [1:36:50] DeepSeek R1 surprised as an open-source o1 competitor, launching without prior expectations.
- **[1:39:53] Wrap-Up**
  - [1:41:19] Prediction: 2025 will be the first year where AI sets the skill floor for a job role.
  - [1:42:05] Prediction from NeurIPS researchers: a foreign spy will be caught at a major AI lab in 2025.
  - [1:49:22] "Next year's the year of the agent in production." — Alessio

## Speakers

- **Alessio** (host)
- **Swyx** (host)

## Topics

Reasoning, Agents, Video Generation

## Mentioned

Anthropic (company), Google (company), Groq (company), Meta (company), NVIDIA (company), OpenAI (company), SSI (company), xAI (company), AWS (product), Blackwell (product), ChatGPT (product), Claude (product), Cursor (product), GPT (product), Gemini (product), H100 (product), Latent Space (product), Llama (product), Perplexity (product), o1 (product)

## Transcript

### Intro

**Alessio** [0:04]
Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, partner and CTO at Decibel Partners, and I'm joined by my co-host, Zwix, for the 100th time today.

**Swyx** [0:13]
Yay. Um, and we're so glad that, yeah, you know, everyone has, uh, followed us in this journey. How do you feel about it, 100 episodes?

**Alessio** [0:19]
Yeah, I know. Almost two years that we've been doing this. We've had four different studios.

**Swyx** [0:24]
Mm-hmm.

**Alessio** [0:25]
Uh, we've had a lot of changes. You know, we used to do this lightning round when we first started-

**Swyx** [0:29]
Knew that, yeah

**Alessio** [0:29]
... that we didn't like . And we tried to change the questions and like-

**Swyx** [0:32]
Because every answer was Cursor and Perplexity.

**Alessio** [0:34]
Yeah, exactly.

**Swyx** [0:35]
Like just-

**Alessio** [0:35]
I love Midjourney. It's like- ... do you really not like anything else? Like, what's the, what's the unique thing? And I think, yeah, we- we've also had a lot more research-driven content. You know, we had, like, Tri Dao, we had Jeremy Howard, we had more folks like that.

I think we wanna do more of that, too, in the new year, like having, uh, some of the Gemini folks, both on the research and the applied side. Yeah, but it's been a ton of fun. I think we both started, I wouldn't say as a joke.

We were kinda like, "Oh, we should do a podcast," and I think we kinda caught the right wave, obviously, and I think your Rise of the AI Engineer po- post just kinda gave people somewhere to congregate, and then the AI Engineer Summit.

And that's why when I look at our growth chart, it's kinda like a proxy for, like, the AI engineering, uh, industry as a whole.

**Swyx** [1:18]
Yeah.

**Alessio** [1:18]
Which is almost like, like, even if we don't do that much, we keep growing just because there's so many more AI engineers, so-

**Swyx** [1:24]
Yeah

**Alessio** [1:24]
... did you expect that growth, or did you expect it would take longer for, like, the AI engineer thing to kinda, like, become, you know-

**Swyx** [1:30]
Um-

**Alessio** [1:30]
... everybody talks about it today.

**Swyx** [1:32]
Yeah. My sign of that, that we have one is that Gartner puts it at the top of the hype c- curve right now. So Gartner has called the peak in A- AI engineering. I did not expect, um, to what level.

I- I knew that I was correct when I called it, because I did like two months of work going into that. But I didn't know, you know, how quickly it could happen, and obviously there's a chance that it would, it-- I could be wrong.

But I think, like, most people have come around to that concept. Hacker News hates it, which is a good sign. But there's enough people that have defined it, you know, GitHub, when they launched GitHub Models, which is the Hugging Face clone, they put AI engineers at the, in the banner, like above the fold, like in, in big letters.

So I think it's, like, kind of arrived as a, as a meaningful and useful definition, and I think people are trying to figure out where the boundaries are. Uh, I think that was a lot of the, quote-unquote, "drama" that happens behind the scenes at the World's Fair in June, because I think there's, there's a lot of doubt or questions about where ML engineering stops and AI engineering starts.

That's a useful debate to be had. In some sense, I actually anticipated that as well, so I intentionally did not put a firm definition there because most of the successful definitions are necessarily under-specified, and it's actually useful to have different perspectives.

**Alessio** [2:42]
Mm-hmm.

**Swyx** [2:43]
And, and, I mean, you don't have to specify everything from the outset.

**Alessio** [2:45]
Yeah. I was at, um, AWS reInvent, and the line to get into, like, the AI engineering talk, so to speak, which is, you know, applied AI and whatnot, was like, there were like hundreds of people just in line to go in.

I think that's kinda what enabled people, right? Which is what you kinda talked about is like, hey, look, you don't actually need a PhD, just-

**Swyx** [3:03]
Yeah

**Alessio** [3:04]
... just use the model. And then maybe we'll talk about some of the blind spots that you get as an AI engineer-

**Swyx** [3:09]
For sure

**Alessio** [3:09]
... with the earlier posts that we also had on, on the Substack. But yeah, it's been a heck of a, heck of a two years.

**Swyx** [3:15]
Yeah. You know, I always, I always kinda view the conference as like, so NeurIPS is, I think, like 16, 17,000 people, and the Latent Space Live event that we held there was 950 signups. I think the AI world, the ML world is still very much research heavy, and, and that's as it should be because ML is very much in a research phase.

But as we move this entire field into pr- production, I think that ratio inverts into becoming more engineering heavy. So at, at least I think engineering should be on the same level, even if it's never as prestigious.

**Alessio** [3:47]
Mm-hmm.

**Swyx** [3:47]
Like, it'll always be low status because, uh, at the end of the day, you're manipulating APIs or whatever.

**Alessio** [3:52]
Yeah.

**Swyx** [3:52]
But you're wrapping GPTs. But there's gonna be an increasing stack and, and an art to doing these, these things well, and I, I, uh, you know, and I think that's what we're focusing on for the podcast, the conference, and basically everything I do.

Seems to make sense.

**Alessio** [4:06]
Mm-hmm.

**Swyx** [4:06]
And I think we'll, we'll talk about the, the trends here that apply. It's, it's this very strange mix of, like, keeping on top of research while not being a researcher, and then putting that research into production. So, like, people always ask me, like, "Why are you covering NeurIPS?"

Like, this is a ML research conference. And I'm like, "Well, yeah, I mean, we're not going to, to, like, understand everything or re- reproduce all the, every single paper," but the stuff that is being found here is going to make its way-

**Alessio** [4:31]
Mm-hmm

**Swyx** [4:31]
... into production at some point, you hope. And then actually, like, when I talk to the researchers, they actually get very excited 'cause they're like, "Oh, you guys are actually caring about how this goes into production," and, and that's what they really, really want.

Their measure of success is previously just peer review, right?

**Alessio** [4:43]
Yeah.

**Swyx** [4:43]
Like getting, getting sevens and eights on their, um, academic review conferences and, and stuff. Like, citations is, is one metric, but money is a better metric.

**Alessio** [4:51]
Right. Yeah, and there were about 2,200 people on the live stream or something like that.

**Swyx** [4:58]
Yeah, yeah, 2,200 on the live stream.

**Alessio** [5:00]
So I tried my best to moderate, but it was a lot spicier in person with Jonathan and, and Dylan-

**Swyx** [5:05]
Yeah

**Alessio** [5:05]
... than it was in the chat on YouTube.

**Swyx** [5:07]
I would say that I actually also created Latent Space Live in order to address flaws that I perceived in academic conferences. This is not NeurIPS specific. It's ICML, it's ICLR, it's NeurIPS. Basically, it's very sort of oriented towards the sort of PhD student, uh, market, job market.

**Alessio** [5:23]
Mm-hmm.

**Swyx** [5:23]
Right? Like, literally all-- Basically everyone's there to advertise their research and skills and get jobs, and then obviously all the, the companies go there to hire them. And I think that's great for the individual researchers, but for people going there to get info is not great because you have to read between the lines, bring a ton of context in order to understand every single paper.

So what is missing is effectively what I ended up doing, which is domain by domain, go through and recap the best of the year. ... survey the field. And there are, like, NeurIPS had a, uh-- I think ICML had a, like, a position paper track.

NeurIPS added a benchmarks, uh, datasets track. These are ways in which to address that issue. Uh, and there's always workshops as well. Every, every conference has, you know, a last day of workshops and stuff that provide more of an overview.

But they're not specifically prompted to do so.

**Alessio** [6:09]
Mm-hmm.

**Swyx** [6:09]
And I think really organizing a conference is just about getting good speakers and giving them the correct prompts, and then they will just go and do their thing, and they do a very good job of it. So I think Sarah did a fantastic job of the, the startups prompt.

**Alessio** [6:22]
Mm-hmm.

**Swyx** [6:22]
I can't list everybody, but, uh, we did best of 2024 in startups, vision, open models, post transformers, synthetic data, small models, and agents. And then the last one was the-- Uh, and then we also did a, a quick one on reasoning with Nathan Lambert.

### Scaling Walls

**Swyx** [6:35]
And then the last one, obviously, was the debate that, uh, people were very hyped about. It was very awkward, and I'm really-- I'm really thankful for John Frankel, basically, who, who stepped up to challenge Dylan, 'cause Dylan was like, "Yeah, I'll do it," but he was pro-scaling.

**Alessio** [6:49]
Mm-hmm.

**Swyx** [6:49]
And I think everyone who is, like, in AI is pro-scaling.

**Alessio** [6:53]
Right.

**Swyx** [6:53]
So you need somebody who's ready to publicly say, "No, we've hit a wall." So that means you're saying Sam Altman's wrong. You're saying-

**Alessio** [7:01]
Right

**Swyx** [7:01]
... um, you know, everyone else is wrong. It helped that this was the day before Ilya went on-- went up on stage and then said, "Pre-training has hit a wall, and data has hit a wall." So actually, Jon- Jonathan ended up winning, and then Ilya supported that statement, and then Noam Brown on the last day further supported that statement as well.

So it's kind of interesting that I think the consensus kinda going in was that we're not done scaling, like you should believe in a bitter lesson.

**Alessio** [7:24]
Mm-hmm.

**Swyx** [7:25]
And then four straight days in a row, you had Seb Hochreiter, who is the, um, creator of the LSTM, along with, um, everyone's favorite OG in AI, which is, uh, Jürgen-

**Alessio** [7:34]
Mm-hmm

**Swyx** [7:34]
... um, Schmidhuber. He said that, um, we-- pre-training has hit a wall, or, like, we've, we've, we've run into a different kind of wall. And then we have, you know, John Frankel, Ilya, and then Noam Brown, like, all saying variations of the same thing.

**Alessio** [7:46]
Mm-hmm.

**Swyx** [7:46]
That we, we have hit some kind of wall in the status quo of what pre-trained scaling, scaling large pre-trained models has looked like, and we need a new thing. And obviously, the new thing for people is some make-- Either people are calling it inference time compute or test time compute.

I think the collective terminology has been inference time.

**Alessio** [8:04]
Okay.

**Swyx** [8:04]
And I think that makes sense because test time, calling it test meaning has a very pre-trained bias, meaning that the only reason for running inference at all is to test your model.

**Alessio** [8:11]
Hmm.

**Swyx** [8:11]
Uh, that is not true. Like

**Alessio** [8:13]
Right. Yeah.

**Swyx** [8:14]
Yeah. So, so I quite agree that OpenAI seems to have adopted, or the community seems to have adopted this terminology of ITC instead of TTC, and that, that makes a lot of sense because, like, now we care about inference.

**Alessio** [8:25]
Yeah.

**Swyx** [8:25]
Even right down to compute optimality. Like, I actually interviewed this author who re-recovered or reviewed the Chinchilla paper.

**Alessio** [8:32]
Hmm.

**Swyx** [8:32]
Chinchilla paper is compute optimal training, but what is not stated in there is it's pre-trained compute optimal training. And once you care-- start caring about inference compute optimal training, you have a different scaling law and in a way that we did not know last year.

**Alessio** [8:45]
I wonder because John is-- he's also on the side of attention is all you need, like he had the bet with Sasha. So I'm curious, like, he doesn't believe in scaling, but he thinks the transformer. I wonder if he's still-

**Swyx** [8:55]
So, so-

**Alessio** [8:56]
... of that camp

**Swyx** [8:56]
... so he-- obviously, everything is nuanced, and, you know-

**Alessio** [8:58]
Yeah, yeah

**Swyx** [8:59]
... I told him to play a character for this debate, right? So he actually does-- Yeah, he still, he still believes that we can scale more. Uh, he just assumed the character-

**Alessio** [9:06]
Ah. Mm-hmm

**Swyx** [9:07]
... to be very game for, for playing this debate. So even more kudos to him that he assumed a position that he didn't believe in and still won the debate.

**Alessio** [9:16]
Get wrecked, Dylan. Um, do you just wanna quickly run through some of these things, like, uh, Sarah's-

**Swyx** [9:22]
Yeah. I, I, um-

**Alessio** [9:22]
... presentation, just the highlights or-

**Swyx** [9:24]
You know, not so-- Yeah, I c-- we can't go through everyone's slides, but I pulled out some things as a factor of, like, stuff that we were gonna talk about.

**Alessio** [9:30]
Yeah.

**Swyx** [9:30]
Um-

**Alessio** [9:30]
And we'll publish the rest.

**Swyx** [9:32]
Yeah, we'll publish on this feed the best of 2024 in those, those domains, and hopefully people can benefit from the work that, uh, our speakers have done. But I think it's, uh-- these are just good slides, and I've been, I've been looking for sort of end of year recaps from, from people.

The field has progressed a lot. You know, I think the max ELO in 2023 on LMSYS used to be two, uh, twelve hundred for LMSYS ELOs, and now everyone is at least at 1275-

**Alessio** [9:56]
Mm-hmm

**Swyx** [9:56]
... in their ELOs. And this is across Gemini, ChatGPT, Groq, o1 AI, which with their E-large model, and, uh, Anthopic, of course. It's a very, very competitive race. There are multiple frontier labs all racing. Um, but there is a clear tier zero frontier.

**Alessio** [10:11]
Mm-hmm.

**Swyx** [10:11]
And then there's, like, a tier one.

**Alessio** [10:13]
Yeah, yeah.

**Swyx** [10:13]
And it's like a wish everything else. And tier zero is extremely competitive. It, it's effectively now a three-horse race between Gemini, Anthropic, and OpenAI. I would say that people are still holding out a candle for XAI. XAI, I think for some reason, because their API was very slow to roll out, it's not included in these-

**Alessio** [10:32]
Oh

**Swyx** [10:32]
... like, uh, metrics. So it's actually quite hard to put on there. Like a- as someone who also does charts, XAI is continually snubbed because they don't work well with the benchmarking people.

**Alessio** [10:42]
Yeah, yeah.

**Swyx** [10:43]
So anyway, there's a little trivia for why XAI always gets ignored. The other thing is, uh, market share. So these are slides from Sarah, and we have it up on, on the screen. It has gone from very heavily OpenAI.

Uh, so we have some numbers and estimates. These, these are from Ramp, uh, estimates of OpenAI market share in December 2023, and this is basically, what is it? GPT-3.5 and GT-4 being ninety-five percent of-

**Alessio** [11:07]
Yeah, ninety percent. Mm-hmm

**Swyx** [11:08]
... production traffic. And I think if you correlate that with stuff that we asked Harrison Chase on, on the LangChain episode-

**Alessio** [11:14]
Mm-hmm

**Swyx** [11:14]
... it was true. And then Claude 3 launched mid-middle of this year. Or I think Claude 3 launched in March. Claude, uh, 3.5 Sonnet was in June-ish. And you can start seeing the market share shift to-towards opening-

**Alessio** [11:26]
Mm-hmm

**Swyx** [11:26]
... uh, towards Anthopic, uh, very, very aggressively. The more recent one is Gemini. So if you-- if I scroll down a little bit, this is a even more recent data set. So Ramp's data set ends in September two-- 2024.

**Alessio** [11:37]
Mm-hmm.

**Swyx** [11:38]
Gemini has basically launched a price war at the low end, um, with-

**Alessio** [11:41]
Yeah

**Swyx** [11:41]
... Gemini Flash, uh, being basically free for personal use. Like, I think people don't understand the free tier. It's something like a billion tokens per day. Unless you're trying to abuse it, you c- you cannot really exhaust your free tier on Gemini.

They're, they're really trying to get you to use it. They know they're in, like, third place-

**Alessio** [11:56]
Mm-hmm, yeah

**Swyx** [11:57]
... um, fourth place, depending how you, how you count, and so they're going after the lower tier first, and then, you know, maybe the up- upper tier later. But yeah, Gemini Flash, according to Open Router, is now 50% of their Open Router requests.

**Alessio** [12:10]
Mm-hmm.

**Swyx** [12:10]
Obviously, these are the small requests.

**Alessio** [12:12]
Yeah, yeah, yeah.

**Swyx** [12:12]
These are small, cheap requests that net, uh, were, uh, mathematically going to be more. The smart ones obviously are still going to OpenAI. But, you know, it's a very, very big shift in, in the market. Like, basically over the course of 2023 t- going into 2024, OpenAI has gone from 95 market share to-

**Alessio** [12:26]
Yeah

**Swyx** [12:26]
... reasonably somewhere between 50 to 75 market share.

**Alessio** [12:29]
Yeah, I'm really curious how ramped up's the attribution to the model if it's API because I think it's all-

**Swyx** [12:34]
Credit card spend.

**Alessio** [12:35]
Well, but it's, uh the credit card doesn't say. Maybe, maybe they'll have-- Maybe when they do expenses they upload the PDF. But yeah, the, the Gemini I think makes sense. I think that was one of my main 2024 takeaways, that, like, the best small model companies are the large labs, which is not something I would have thought that the open source kinda, like, long tail would be like the small model.

**Swyx** [12:54]
Yeah. Different sizes of small models we're talking about here, right? Like, so small model here for Gemini is AB, right? Uh, Mini we don't know what the small-

**Alessio** [13:02]
Mm-hmm

**Swyx** [13:02]
... model size is. But yeah, it's probably in the double digits or maybe single digits, but probably double digits. The open source community has kind of focused on the 1 to 3B size.

**Alessio** [13:12]
Mm-hmm, yeah.

**Swyx** [13:12]
Maybe zer- maybe 0.5B. Uh, that's Moondream. And that is small for you, then, then that's great. It makes sense that we, we have a range for small now, which is, like, may- maybe 1 to 5B.

**Alessio** [13:22]
Yeah.

**Swyx** [13:22]
I'll even put that at, at the high end. And so this includes Gemma from Gemini as well, but it also includes the Apple Foundation models-

**Alessio** [13:29]
Mm-hmm, yeah

**Swyx** [13:30]
... um, which are, which are all s- like I think Apple Foundation is 3B.

**Alessio** [13:33]
Yeah. No, that's great. I mean, I think in the start small just meant cheap. I think today small is actually a more nuanced, um-

**Swyx** [13:40]
Yeah

**Alessio** [13:40]
... discussion, you know-

**Swyx** [13:41]
Yeah

**Alessio** [13:41]
... that people didn't, weren't really having before.

**Swyx** [13:44]
Yeah.

**Alessio** [13:44]
Um-

**Swyx** [13:44]
Uh, we can keep going. This is a slide that I smiley disagree with Sarah. She's pointing to the C- uh, the Scale CEW leaderboard.

**Alessio** [13:51]
Mm-hmm.

**Swyx** [13:51]
Uh, uh, I think the researchers that I talked with at NeurIPS were kind of positive on this because basically you need private test sets to prevent contamination, and Scale is one of maybe three or four people this year that has, like, really made an effort in, in-

**Alessio** [14:08]
Yeah

**Swyx** [14:08]
... doing a, a credible private test set leaderboard. Llama 4O 5B does well compared to Gemini and GPT-4o, and I think that's, that's good. I would say that, you know, it's good to have an open model that is that big that does well on those metrics.

But anyone putting 4O5 B in production will tell you-

**Alessio** [14:27]
Mm-hmm

**Swyx** [14:27]
... if you scroll down a little bit to the artificial analysis numbers, that it is very slow and very expensive to infer. Um, it doesn't even fit on, like, one node of a, of H100s. Cerebras will be happy to tell you they can serve 4O5B

**Alessio** [14:41]
Mm-hmm, right

**Swyx** [14:42]
... on their, on their super large chips. But, um, you know, if you need to, to do anything custom to it, you are s- you're still, you're still k- kinda constrained. So is 4O5B really that relevant? Like, I think m- most people are basically saying that they only use 4O5B as a teacher model-

**Alessio** [14:55]
Mm-hmm

**Swyx** [14:55]
... to distill down t- to something. Even Meta is doing it. So when Llama 3.3 launched, they only launched the 70B because they used 4O5B to distill the 70B. So I, I don't know if, like, open source is keeping up.

I think they're-- the, the open source industrial complex is very invested in telp- telling you that the, uh, the gap is narrowing. I kind of disagree.

**Alessio** [15:15]
Mm-hmm. Yeah, yeah.

**Swyx** [15:16]
I, I think that the gap is widening, uh, with o1. I think th- there are very, very smart people trying to narrow that gap, and they should. I really wish them success. But you cannot use a chart that is nearing 100 in your saturation chart and look, the, the, the distance between open source and closed source is narrowing.

Of course, it's gonna narrow it 'cause you're, you're near 100.

**Alessio** [15:33]
No, yeah.

**Swyx** [15:33]
Um this is stupid. Um, but in metrics that matter, um, is open source narrowing? Probably not for o1 for a while.

**Alessio** [15:41]
Yeah.

**Swyx** [15:41]
Uh, and it, it, it's really up to the open source guys to figure out if they can match o1 or not.

**Alessio** [15:47]
I think inference time compute is bad for open source just because, you know, Doc can donate the flops at training time, but he cannot donate the flops at inference time.

**Swyx** [15:55]
Mm.

**Alessio** [15:56]
So it's really hard to, like, actually keep up on that axis.

**Swyx** [15:59]
Big, big business model shift.

**Alessio** [16:01]
You know? So I don't know what that means for the GPU clouds. I don't know what that means for the hyperscalers. But obviously the big labs have a lot of advantage because, like-

**Swyx** [16:10]
Yeah

**Alessio** [16:10]
... it's not a static artifact that you're putting the compute in. You're kinda doing that still, but then you're putting a lot of compute at inference too.

**Swyx** [16:17]
Yeah, yeah, yeah. Um, I mean, Llama 4 will be reasoning oriented.

**Alessio** [16:21]
Mm-hmm.

**Swyx** [16:21]
We talked with Thomas Shalom.

**Alessio** [16:22]
Yeah.

**Swyx** [16:22]
Um, kudos for getting that episode together. That was really nice. Good, well timed. Actually, I connected with the AI Meta guy, uh, at NeurIPS, and, um, yeah, we're, we're gonna coordinate something for Llama 4.

**Alessio** [16:33]
Yeah, yeah. And our friend Clara Shi-

**Swyx** [16:35]
Mm-hmm

**Alessio** [16:35]
... just joined to lead the-

**Swyx** [16:36]
Yes

**Alessio** [16:36]
... business agent side, so I'm sure we'll have her on-

**Swyx** [16:38]
Awesome

**Alessio** [16:38]
... in the new year.

**Swyx** [16:39]
Yeah. So, um, my comment on, on the business model shift, this is super interesting. Apparently it is wide knowledge that OpenAI wanted more than $6.6 billion for their fund raise.

**Alessio** [16:48]
Hm.

**Swyx** [16:49]
They wanted to raise higher, and they did not. And what that means is basically, like, it's very convenient that we're not getting GPT-5 whi- which would have been a larger pre-train, which would have a lot of upfront money.

In- instead we're, we're fi- we're converting fixed costs into variable costs-

**Alessio** [17:03]
Mm-hmm

**Swyx** [17:03]
... right? And passing it on effectively to the customer.

**Alessio** [17:07]
Yep.

**Swyx** [17:07]
And it's so much easier to take margin there because you can directly attribute it to, like, "Oh, you're using this more, therefore you, you pay more of the cost and I'll, I'll just slap a margin in there." So, like, that lets you control your gross margin and, like, tie your, your s- your spend or your, your sort of inference spend accordingly.

And it's just, it's just really interesting to... that, that, that this change in the sort of inference paradigm has arrived exactly at the same time that the, the funding environment for pre-training is effectively drying up-

**Alessio** [17:35]
Mm-hmm

**Swyx** [17:36]
... kind of. I feel like maybe the VCs are very in tune with research anyway, so, like, they would have noticed this, but, um, it's just interesting.

**Alessio** [17:43]
Yeah. Yeah. And I was looking back at our yearly recap of last year, and the big thing was, like, the mixed trial price fights, you know? And I think now it's almost like there's nowhere to go. Like, you know, Gemini Flash is, like, basically giving it away for free, so-

**Swyx** [17:56]
Yeah

**Alessio** [17:56]
... I think this is a good way for the labs to generate more revenue and pass down-

**Swyx** [18:00]
It's great

**Alessio** [18:01]
... some of the compute to-

**Swyx** [18:02]
Yeah

**Alessio** [18:02]
... to the customer.

**Swyx** [18:02]
I think they're gonna keep going. I think that, uh, $2,000 ChatGPT will come.

**Alessio** [18:06]
Yeah, yeah. No, totally. I mean, next year the first thing I'm doing is signing up for Devin, signing up for the-

**Swyx** [18:12]
Really?

**Alessio** [18:12]
... Pro ChatGPT just to try. I just wanna see what does it look like to spend $1,000 a month.

**Swyx** [18:17]
Yes

**Alessio** [18:17]
... on AI?

**Swyx** [18:18]
Yes. I think-

**Alessio** [18:19]
Wow

**Swyx** [18:19]
... if your, if your, your job is a at least AI content creator or VC or, you know, someone who- whose job it is to stay on, stay on top of things, you should already be spending like $1,000 a month on, on stuff.

And then obviously, easy to spend, hard to use.

**Alessio** [18:31]
Yeah.

**Swyx** [18:32]
You have to actually use. The good thing is that actually Google lets you do a lot of stuff for free now.

**Alessio** [18:37]
Mm-hmm.

**Swyx** [18:37]
Um, so like Deep Research that they just launched uses a ton of inference and it's, it's free while it's in preview.

**Alessio** [18:45]
Yeah.

**Swyx** [18:46]
So you should use it.

**Alessio** [18:47]
Yeah. They need to put that in Lindy. I've been using Lindy lately.

**Swyx** [18:50]
Yeah.

**Alessio** [18:50]
I've been a- built a bunch of things once we had Flow because I like the new thing. It's pretty good. I even did a phone call assistant, um-

**Swyx** [18:57]
Yeah, they just launched the voice

**Alessio** [18:59]
... so you can do-- Yeah. I think once they get advanced voice mode-like capability. Today it's still like speech-to-text. You can kinda tell. Um, but it's good for like reservations and things like that, so I have a meeting prepper thing.

It's, uh-

**Swyx** [19:13]
Yeah

**Alessio** [19:13]
... it's good.

**Swyx** [19:14]
Okay. I feel like we've, we've covered a lot of stuff. Uh, yeah, you know, I think we will go over the individual, uh, talks in a separate episode.

**Alessio** [19:23]
Yeah.

**Swyx** [19:23]
Uh, I don't wanna take too much time with, uh, this stuff.

**Alessio** [19:25]
Mm-hmm.

**Swyx** [19:25]
But that it- suffice to say that there is a lot of progress in each field. Uh, we covered vision. Basically, this is all like the audience voting for what they wanted.

**Alessio** [19:33]
Yeah.

**Swyx** [19:33]
And then I just invited the best speaker I could find in, in each audience. Especially agents, um, Graham, who I talked to at ICML in Vienna, he is currently still number one. It's very hard to stay on top of Sweet Bench.

He start- he, uh-- OpenHand is currently still number one on Sweet Bench full, which is the hardest one. He had very good thoughts on agents, which I, which I'll highlight for people. Everyone is saying 2025 is the year of agents-

**Alessio** [19:55]
Mm-hmm

**Swyx** [19:56]
... just like they said last year. And, uh, but he had thoughts on like eight parts of what are the frontier problems to solve in agents, and so I'll highlight that talk as well.

**Alessio** [20:06]
Yeah. The number six, which is the how can agents learn more about the environment, has been super interesting to us as well just to think through. Because yeah, how, how do you put an agent in an enterprise where most things in an enterprise have never been public?

You know-

**Swyx** [20:21]
Mm-hmm

**Alessio** [20:21]
... a lot of the tooling, like the code bases and things like that, so yeah, there's not really-

**Swyx** [20:24]
So just indexing and RAG?

**Alessio** [20:26]
Well, yeah, but it's more like you can't really RAG things that are not documented, but people know them based on how they've been doing it, you know? So I think there's almost this like-

**Swyx** [20:35]
Oh, institutional knowledge

**Alessio** [20:35]
... you know-- The, yeah, the boring word is kinda like a business process extraction. It's like-

**Swyx** [20:39]
Yeah, yeah, I see

**Alessio** [20:39]
... how do you actually understand how these things are done?

**Swyx** [20:41]
I see.

**Alessio** [20:42]
Um, and I think today the, yeah, the agents are-- that most people are building are good at following instruction but are not as good at like extracting them from you.

**Swyx** [20:50]
Yeah.

**Alessio** [20:50]
Um, so I think that will be a big unlock.

**Swyx** [20:53]
Cool.

**Alessio** [20:53]
Just to touch quickly on the Jeff Dean thing. I thought it was pretty-- I mean, we'll link it in the, in the things, but I think the main focus was like, how do you use ML to optimize the systems instead of just focusing on ML to do something else?

Yeah, I think speculative decoding we had, you know, Eugene from RWKV on the podcast before. Like, he's doing a lot of that with Federalless AI.

**Swyx** [21:13]
Everyone is. It's-- I would say it's the norm. I'm a little bit uncomfortable with how much it costs, uh, because it-

**Alessio** [21:18]
Mm

**Swyx** [21:18]
... it does use more of the GPU per call. Uh, but because everyone is so keen on fast inference, then yeah, makes sense.

**Alessio** [21:25]
Exactly. Um, yeah, but we'll link that, um, obviously Jeff-

**Swyx** [21:29]
Yeah

**Alessio** [21:30]
... is great.

**Swyx** [21:30]
So the-- Jeff is-- Jeff's talk was more-- It wasn't focused on Gemini.

**Alessio** [21:34]
Mm-hmm.

**Swyx** [21:34]
I think people, uh, got the wrong impression from my tweet. It is more about how Google approaches ML and uses ML to design systems and like-- and then, and then systems feed back into the ML. And I think this ties in with Lubna's talk on synthetic data, where it's basically the story of bootstrapping of humans and AI in AI research or AI in production.

So her talk was on synthetic data where like how much synthetic data has grown in 2024 or in the pre-training side, the post-training side, and the eval side. And I think Jeff then also extended it basically to chips, uh, to chip design.

So he spent a lot of time talking about Alpha Chip, and most of us in the audience are like, "We're not working in hardware, man. Like you guys are great. TPU is great. Okay, we'll buy TPUs."

**Alessio** [22:14]
And then there was the Ilya talk.

**Swyx** [22:16]
Yeah.

**Alessio** [22:17]
But and then we have a essay tied to it, what Ilya saw.

**Swyx** [22:21]
Mm-hmm.

**Alessio** [22:21]
I don't know if we're calling them essays. What are we calling these? But posts.

**Swyx** [22:24]
Uh, for, for me it's just like bonus for Latent Space supporters because I feel like they haven't been getting anything.

**Alessio** [22:29]
Yeah.

**Swyx** [22:29]
And then, uh-

**Alessio** [22:30]
It's fair

**Swyx** [22:30]
... I wanted a more high-frequency way to write stuff. Like that one I wrote in an afternoon. I think basically we now have an answer to what Ilya saw.

**Alessio** [22:39]
Mm-hmm.

**Swyx** [22:39]
It's one year since the blip, and we know what he saw in 2014, uh, 14. We know what he saw in 2024. We think we know what he sees in 2024. He, he gave some hints. And then we have vague indications of what, what he saw in 2023.

**Alessio** [22:54]
Mm-hmm.

**Swyx** [22:54]
So that was the-

**Alessio** [22:56]
Yeah.

**Swyx** [22:56]
Oh, and then 2016 as well. Because of this lawsuit with Elon, OpenAI is publishing emails-

**Alessio** [23:02]
Mm-hmm

**Swyx** [23:02]
... from Sam's-- like his personal text messages to Shivon Zilis or whatever. So like we have emails from Ilya saying, "This is what we're seeing in OpenAI, and this is why we need to scale up GPUs." And I think it's very prescient in 2016 to write that.

And so like it is exactly like basically his insights. It's him and Greg basically just kinda driving the scaling up of, of OpenAI while they're still playing Dota. They are like, "No," like-

**Alessio** [23:28]
Yeah, yeah, yeah

**Swyx** [23:29]
... like-

**Alessio** [23:29]
Yeah

**Swyx** [23:29]
... like, "We, we see the path here."

**Alessio** [23:31]
Yeah, and it's funny. Yeah, they even mention, you know, we can only train on 1v1 Dota. We need to train on 5v5, and that takes too many GPUs and-

**Swyx** [23:38]
Yeah, a- yeah, and at least for me, I, I can speak for myself, like I didn't see the path from Dota to where we are today. I think even maybe if you ask them, like they, they wouldn't necessarily draw a straight line.

**Alessio** [23:46]
Mm-hmm.

**Swyx** [23:46]
But, um-

**Alessio** [23:47]
Yeah. Yeah, no, I definitely, no. But I think like that was like the whole idea of almost like the RL, and we talked about this with Nathan on his podcast. It's like with RL you can get very good at specific things, but then you can't really like generalize as much.

And I think the language models are like the opposite, which is like you're gonna throw all this data at them and scale them up, but then you really need to drive them home-

**Swyx** [24:07]
Yeah

**Alessio** [24:07]
... on a specific task later on.

**Swyx** [24:09]
Yeah.

**Alessio** [24:09]
And we'll talk about the OpenAI reinforcement fine-tuning, um-

**Swyx** [24:12]
Yeah

**Alessio** [24:12]
... announcement too and all of that. But yeah, I think like scale is all you need. That's kinda what Ilya will be. Remembered for

**Swyx** [24:19]
It'll be remembered for, yeah.

**Alessio** [24:21]
Um, and I think just maybe to clarify on, like, the pre-training is over thing that people love to tweet, I think the point of the talk was like everybody-- we're scaling these chips, we're scaling the compute, but, like, the second ingredient, which is data, is not scaling at the same rate.

So it's not necessarily pre-training is over. It's kinda like what got us here won't get us there. In his email, he predicted, like, 10X growth every two years or something like that, and I think maybe now it's like, you know, you can 10X the chips again, but you can only-

**Swyx** [24:49]
I think it's 10X per year. W-was it? I, I don't know.

**Alessio** [24:53]
Exactly. And Moore's Law is like 2X.

**Swyx** [24:55]
Yeah.

**Alessio** [24:55]
So it's like, you know, mu-mu-much faster than that. And yeah, like the fossil fuel of AI analogy is kinda like, you know, the little background tokens-

**Swyx** [25:03]
Yeah

**Alessio** [25:03]
... thing. And so the OpenAI reinforcement fine-tuning is basically like instead of fine-tuning on data, you fine-tune on a reward model. So it's basically like instead of being data-driven, it's like task-driven. And I think people have tasks to do.

They don't really have a lot of data. So I'm curious to see how that changes how many people fine-tune because I think this is what people run into is like, "Oh, you can fine-tune Llama," and it's like, "Okay, where do I get the data to fine-tune it on?"

You know? So it's great that we're moving the thing. And then I really like he had this chart where, like, you know, the brain mass and the body mass thing. It's basically like mammals have scaled linearly by brain and body size, and then humans kinda like broke off the slope.

So it's-

**Swyx** [25:44]
Yeah

**Alessio** [25:44]
... almost like maybe the mammal slope is like the pre-training slope, and then the- ... post-training slope is like the, the human one.

**Swyx** [25:50]
Yeah. I wonder what the-- I mean, we'll know in 10 years, but, uh-

**Alessio** [25:53]
Yeah

**Swyx** [25:53]
... I wonder what the Y-axis is for, for Ilya's SSI. We'll try to get them on.

**Alessio** [25:58]
Ilya- ... if you're listening, you're welcome here. Yeah, and then he had, you know, what comes next, like agent synthetic data and inference compute.

**Swyx** [26:05]
Yeah. I don't think-

**Alessio** [26:05]
I thought all of that was like-

**Swyx** [26:06]
I don't think he was dropping any alpha there.

**Alessio** [26:07]
Yeah, yeah, yeah.

**Swyx** [26:08]
Yeah.

**Alessio** [26:08]
Any other NeurIPS highlights or...?

**Swyx** [26:11]
I think that there was comparatively a lot more work-- Oh, by the way, I, I need to plug that, uh, my friend Yi made this, like, little nice card here-

**Alessio** [26:20]
Yeah, it was really nice

**Swyx** [26:21]
... uh, of, uh, of like all the... He's-- She called it Must-Read Papers of 2024.

**Alessio** [26:26]
Mm-hmm.

**Swyx** [26:26]
So I laid out some of these at NeurIPS, and it was just gone. Like, everyone just picked it up 'cause people are dying for, like, little guidance and visualizations-

**Alessio** [26:34]
Yeah

**Swyx** [26:35]
... of, of each paper. And so, uh, I thought it's really super nice-

**Alessio** [26:38]
Mm-hmm

**Swyx** [26:38]
... uh, that, that we got that.

**Alessio** [26:39]
Should we do a latent space book-

**Swyx** [26:41]
Uh-

**Alessio** [26:41]
... for each year?

**Swyx** [26:42]
I thought about it. I thought about it.

**Alessio** [26:42]
For each year we should-

**Swyx** [26:43]
Coffee, coffee table book.

**Alessio** [26:43]
Yeah.

**Swyx** [26:44]
Yeah. Uh-

**Alessio** [26:45]
Okay. Put it in the will. Um, hi, Will. By the way, we haven't introduced you. He's our new, you know-

**Swyx** [26:51]
We need to pull up-

**Alessio** [26:51]
Juro organist Jamie. We are Will.

**Swyx** [26:52]
We need to pull up more things. One thing I saw that, um... Okay, there, there-

**Alessio** [26:56]
What this folks see.

**Swyx** [26:58]
Okay, one, one fun one and then one more, more general one.

**Alessio** [27:01]
Mm-hmm.

**Swyx** [27:01]
So the fun one is, uh, this paper on agent collusion. This is a paper on steganography. Uh, this is Secret Collusion among AI Agents: Multi-Agent Deception via Steganography. I try to go to NeurIPS in order to find these kinds of papers because the real reason-- Like, NeurIPS this year is-- has a lottery system.

A lot of people actually even go and don't buy tickets, um-

**Alessio** [27:20]
Mm-hmm

**Swyx** [27:20]
... 'cause they just go and attend the side events.

**Alessio** [27:22]
Yeah.

**Swyx** [27:22]
And then also the people who go and end up crowding around the most popular papers which you already know-

**Alessio** [27:27]
Yeah

**Swyx** [27:27]
... and already read them before you showed up to NeurIPS. So the only reason you go there is to, to talk to the paper authors. But there's like something like 10,000 other papers out there that, you know, are just people's work that they, that they did on the year and they, they failed to get attention for one reason or another, and this was one of them.

Uh, it was like all the way at the back. And this is a DeepMind paper that actually focuses on collusion between AI agents, uh, by hiding messages in the text that they generate.

**Alessio** [27:53]
Mm.

**Swyx** [27:53]
Uh, so that's what steganography is. So a very simple example would be the first letter of every word. If you pick that out, you, you know, and decode, there's a different message than, than, than that. But something I've always emphasized is, uh, to LLMs, we read left to right.

LLMs can read up, down, sideways, you know, in random character order, and it, it's the same to them as it is to us. So if we were ever to get, you know, self-motivated, unaligned LLMs that were trying to collaborate to take over the planet, this would be how they do it.

They, they spread, they spread messages among us in the messages that we generate. And he developed a scaling law for that.

**Alessio** [28:27]
Hmm.

**Swyx** [28:27]
Um, so he, he marked, uh... I'm showing it on screen right now, how the emergence of this phenomenon, uh, basically for, for example, for cipher encoding, GPT-2, Llama 2, Mixtral, GPT-3.5, zero capabilities and its sudden emergence of GPT-4.

**Alessio** [28:40]
Hmm.

**Swyx** [28:40]
And this is the kind of Jason Wei-type emergence-

**Alessio** [28:42]
Yeah, yeah

**Swyx** [28:43]
... properties that, that people kinda look for. I think the-- what pe- made this paper stand out as well, so he developed a benchmark for steganography collusion, and he also focused on shelling point collusion, which is very-

**Alessio** [28:54]
Mm

**Swyx** [28:54]
... low coordination. Like, for agreeing on a decoding, encoding format, you kind of need to have some agreement on that, but, but shelling point means, like, very, very low or almost no coordination. So for example, if I, if I ask someone if the only message I give you is meet me in New York, and you-- uh, you have no idea where or when, you would probably meet me at Grand Central Station.

**Alessio** [29:14]
Mm-hmm.

**Swyx** [29:15]
That is the sh-- Grand Central Station-

**Alessio** [29:16]
Yeah

**Swyx** [29:16]
... is a shelling point and, and probably somewhere, somewhere during the day. That is Gran-- The, the shelling point of New York is Grand Central. To that extent, shelling points for steganography are t- are things like the, the, the common decoding methods that we talked about.

It will be interesting at some point in the future when we are worried about alignment.

**Alessio** [29:31]
Yeah.

**Swyx** [29:31]
It is not interesting today, but it's interesting that DeepMind is already thinking about this.

**Alessio** [29:35]
Mm-hmm. Interesting. I think that's, like, one of the hardest things about NeurIPS is, like, the long tail-

**Swyx** [29:40]
Long-

**Alessio** [29:40]
... papers

**Swyx** [29:41]
... very long tail.

**Alessio** [29:41]
It's like-

**Swyx** [29:42]
I, I found a pricing guy. I'm gonna feature him on the podcast. Basically, uh, it's guy-- Some-- this guy from NVIDIA worked out the optimal pricing for language models.

**Alessio** [29:51]
Hmm.

**Swyx** [29:51]
It was basically an econometrics paper at NeurIPS where, like, everyone else is talking about GPUs.

**Alessio** [29:56]
And the guy with the GPUs is talking about economics instead.

**Swyx** [29:59]
About, uh, pricing. Yeah.

**Alessio** [29:59]
Yeah.

**Swyx** [30:00]
That was the sort of fun one. The broader focus I saw is that model papers at NeurIPS are kinda dead. No one really presents models anymore. It's just datasets because it's all the grad students are working on. So, like, there was a datasets track, and then I, I was looking around like, "You don't need a datasets track because every paper is datasets paper."

And so, uh, uh, data sets and benchmarks-

**Alessio** [30:25]
Yeah

**Swyx** [30:25]
... they're kind of like flip sides of the same thing. So yeah, if you're a grad student, you're a GPU poor, you kind of work on that, and then the, the sort of big model that people walk around and pick the ones that they like, and then they use it in their models.

**Alessio** [30:36]
Mm-hmm.

**Swyx** [30:37]
And, you know, that's, that's kind of how it develops, I, I, I feel like. Um, like, like you didn't-- last year you had people like Hao Tian who worked on Llava-

**Alessio** [30:45]
Mm-hmm

**Swyx** [30:45]
... uh, which was take Llama and add Vision and, and obviously xAI hired him, and he added, uh, Vision to Grok. Now he's the Vision Grok guy. This year I don't think there was any of those.

**Alessio** [30:56]
Yeah. What were the most popular, like, orals? Last year it was like the Mixed Monarch, I think was like the most attended-

**Swyx** [31:03]
Yeah

**Alessio** [31:04]
... one.

**Swyx** [31:04]
Uh, I need to look it up.

**Alessio** [31:06]
Yeah, I mean, if nothing comes to mind, that's also kind of like a, an answer in a way. But I think last year there was a lot of interest in like-

**Swyx** [31:13]
Yeah

**Alessio** [31:13]
... furthering models and like different architectures and, and all that.

**Swyx** [31:17]
I will say that I feel, I felt the orals, oral picks this year were not very good.

**Alessio** [31:21]
Hmm.

**Swyx** [31:21]
Either that or maybe it's just a highlight of how I have changed in terms of how I view papers.

**Alessio** [31:29]
Mm-hmm. Yeah, yeah.

**Swyx** [31:29]
So like, in my estimation, two of the best papers in this year for data sets were Data Comp and Refine Web or Find Web.

**Alessio** [31:37]
Mm-hmm. Yep.

**Swyx** [31:38]
These are two actually, actually industrially used papers not highlighted for oral. I think DCLM got the spotlight. Find Web didn't even get a spotlight. So like, it's just the, the picks were different.

**Alessio** [31:48]
Mm-hmm.

**Swyx** [31:49]
Um, but one thing that does get a lot of play that a lot of people are debating is the role that's scheduled. This is the Schedule-Free Optimizer paper from Meta, from Ed- Aaron Defazio. And this year in the ML community, there's been a lot of chat about shampoo, soap, all the bathroom amenities for optimizing your learning rates.

And, um, most people at the big labs, uh, who I asked about this, um, say that it's cute, but it's not-

**Alessio** [32:15]
Hmm

**Swyx** [32:15]
... something that matters. I don't know. But it's something that was, that was discussed-

**Alessio** [32:18]
Yeah

**Swyx** [32:18]
... and been very, very popular.

**Alessio** [32:20]
Four Wars of AI recap, maybe-

### Data War

**Swyx** [32:21]
Yeah

**Alessio** [32:21]
... just quickly. Um, where do you wanna start?

**Swyx** [32:24]
Uh.

**Alessio** [32:24]
Data, I guess?

**Swyx** [32:26]
Yeah. So to remind people, this is the Four Wars piece that we did as, uh, one of our earlier recaps of this year, and the belligerents are on the left, journalists, writers, artists, uh, anyone who know- owns IP, basically, New York Times, Stack Overflow, Reddit, Getty, Sarah Silverman, George R.R.

Martin. Yeah. And I think this year we can add Scarlett Johansson to that-

**Alessio** [32:45]
Mm-hmm

**Swyx** [32:45]
... to that side of the fence. So anyone suing OpenAI, basically. I, I actually wanted to get a snapshot of like all the lawsuits.

**Alessio** [32:52]
Hmm.

**Swyx** [32:52]
Um, I'm sure some lawyer can, can do it. That's the data quality war. On the right-hand side, we have the synthetic data people, uh, and I think we talked about Luminous Talk, you know, really showing how much synthetic data has, has come along this year.

I think there was a bit of a fight between Scale AI and the synthetic data community 'cause Scale published a paper saying that synthetic data doesn't work. Surprise, surprise, Scale is the leading vendor of non-synthetic data .

**Alessio** [33:17]
Only cage, cage-free annotated data is useful.

**Swyx** [33:22]
So I think there's, there's some debate going on there, but I don't think there's much debate anymore that at least synthetic data for the reasons that are blessed in, uh, in Luminous Talk makes sense.

**Alessio** [33:32]
Mm-hmm.

**Swyx** [33:33]
I, I don't know if you have any perspectives there.

**Alessio** [33:34]
I think, again, going back to the reinforcement fine-tuning, I think that will change a little bit how people think about it. I think today people m- mostly use synthetic data, yeah, for distillation and kind of like fine-tuning a smaller model from like a larger model.

I'm not super aware of how the Frontier labs use it outside of like the rephrase the web thing that Apple also did. But yeah, I think it'll be useful. I think like whether or not that gets us the big next step, I think that's maybe like TBD, you know.

I think people love talking about data because it's like a GPU poor thing, you know. I think, uh-

**Swyx** [34:08]
Mm-hmm

**Alessio** [34:09]
... synthetic data is like something that people can do, you know?

**Swyx** [34:12]
Yeah.

**Alessio** [34:12]
So they feel more opinionated about it compared to, yeah, the optimizers stuff, which is like they don't really work on.

**Swyx** [34:19]
I think that there is an angle to the reasoning synthetic data. So this year we covered in the paper club the s- um, star series of papers, so that's star, QStar, VStar. It basically helps you to synthesize reasoning steps or at least distill reasoning steps from a verifier.

And if you look at the OpenAI RFT API that they, that they released or that they announced, basically they're asking, they're asking you to submit graders, or they choose from a preset list of graders. Basically, it feels like a way to create valid synthetic data for them to fine-tune their reasoning paths on.

Um, so I think that is another angle where, um, it starts to make sense. And so, like, it's very funny that basically all the data quality wars between, let's say, the music industry or like the newspaper publishing industry or like the textbooks industry on the big labs, it's all of the pre-training era.

**Alessio** [35:13]
Mm-hmm.

**Swyx** [35:13]
And then like the, the new era, like the reasoning era, like nobody has any problem with-

**Alessio** [35:17]
Right

**Swyx** [35:17]
... all the reasoning, uh, fine-tunes be- uh, especially because it's all like sort of-

**Alessio** [35:20]
Yeah

**Swyx** [35:20]
... math and science oriented with, with very reasonable graders. I think the, the more interesting next step is how does it generalize beyond STEM? We've been using o1 for AI news for a while, and I, I would say like for summarization and creative writing and instruction following, I think it's underrated.

I started using o1 in our intro songs before we killed the intro songs, but it's very good at writing lyrics. You know, it can actually say like... I think one of the o1 Pro demos that Noam was showing was that, you know, you can write an entire paragraph or three paragraphs without using the letter A, right?

**Alessio** [35:54]
Mm.

**Swyx** [35:54]
So like, like literally just anything sort of token, like not even token level, character level manipulation and counting and instruction following-

**Alessio** [36:02]
Yeah

**Swyx** [36:02]
... it's, uh, it's very, very strong at. So no surprises when I ask it to rhyme, uh, and to, to create song lyrics, it's gonna do that very much better-

**Alessio** [36:08]
Mm

**Swyx** [36:08]
... than in, than in previous models. So I think it's underrated for creative writing.

**Alessio** [36:12]
Yeah. What do you think is the rationale that they're gonna have in court when they don't show you the thinking traces of o1, but then they want us to... Like, they're getting sued for using other publishers' data, you know?

**Swyx** [36:25]
I see.

**Alessio** [36:25]
But then on their end, they're like, "Well, you shouldn't be using my data to then train your model."

**Swyx** [36:29]
Mm-hmm.

**Alessio** [36:29]
So I'm curious to see how that- Kind of comes-

**Swyx** [36:31]
Um, yeah

**Alessio** [36:32]
... OpenAI has-

**Swyx** [36:32]
I mean, OpenAI has many ways to publish- to punish people without bringing- taking them to court. Already banned ByteDance for distilling-

**Alessio** [36:38]
Mm-hmm

**Swyx** [36:38]
... their, their info. And so anyone caught distilling the chain of thought will be just disallowed to continue on, on, on the API, and it's fine. It's no big deal. Like, I, I don't even think that's an issue at all just because the chain of thoughts are pretty well hidden.

Like, you have to work very, very hard-

**Alessio** [36:54]
Yeah

**Swyx** [36:54]
... to, to get it to leak, and then even when it leaks the chain of thought, you don't know if it's the real one.

**Alessio** [36:59]
Mm-hmm.

**Swyx** [36:59]
So there's much less concern here.

**Alessio** [37:01]
Yeah, yeah, yeah.

**Swyx** [37:02]
The bigger concern is actually that there's not that much IP hiding behind it.

**Alessio** [37:07]
Mm.

**Swyx** [37:07]
That, like, Cosine, which we talked about, that we, we talked to him on Dev Day, can just fine-tune 4o to beat o1, that Cloud Sonnet so far is beating o1, uh, on coding tasks without at least o1 preview, wi-without being a reasoning model, and same for Gemini Pro or Gemini 2.0.

**Alessio** [37:25]
Mm-hmm.

**Swyx** [37:26]
Um, so, like, how much is reasoning important? How much of a moat is there in this, like, proprietary sort of training data that they've, uh, presumably ac-accomplished? Because, like, even DeepSeek was able to-

**Alessio** [37:37]
Yeah

**Swyx** [37:37]
... to do it, and they had, you know, two months notice to do this, to do R1. So it's actually unclear how much moat there is. Um, obviously, you know, if you talk to the Strawberry team, they'll, they'll be like: "Yeah, I mean, we spent the last two years doing this."

So we don't know. Um, and it's going to be interesting because there'll be a lot of noise from people who say they have inference time-

**Alessio** [37:56]
Mm-hmm

**Swyx** [37:57]
... compute and actually don't because they just have fancy chain of thought. And then there's other people who actually do have very good chain of thought, and you will not see them on the same level as OpenAI because OpenAI has invested a lot in building up the mythology of their team.

**Alessio** [38:10]
Mm-hmm.

**Swyx** [38:11]
Um, which makes sense. Like, the real answer is somewhere in between.

**Alessio** [38:13]
Yeah. I think that-then that's kinda like the main data war story developing. GPU poor versus GPU rich.

### GPU War

**Swyx** [38:21]
Yeah.

**Alessio** [38:22]
Where do you think we are? I think there was, again, going back to, like, the small model thing, there was, like, a time in which the GPU poor were kinda like the rebel faction working on, like, these models that were, like, open and small and cheap.

And I think today people don't really care as much about GPUs anymore. You also see it in the price of the GPUs. Like, you know, that market has kinda like plummeted because there's-- people don't wanna be-- they wanna g- be GPU free.

They don't even wanna be poor. They just wanna be, you know, completely without them. Yeah, how do you think about this war developing?

**Swyx** [38:52]
You can tell me about this, but, like, I feel like the, the appetite for GPU-rich startups, like the, you know, the, the funding plan is we will raise sixty million, and we'll give fifty of that to Nvidia.

**Alessio** [39:01]
Mm-hmm.

**Swyx** [39:02]
That is gone, right? Like no one's-

**Alessio** [39:03]
Yeah

**Swyx** [39:03]
... no one's pitching that. This was literally the plan- the exact plan of, like, I can name like four or five startups-

**Alessio** [39:09]
Mm-hmm

**Swyx** [39:09]
... that, you know, this time last year. So yeah, GPU-rich startup's gone. Um, but I think, like, the GPU ultra rich, The GPU ultra high net worth is still going. So, um, now we're, you know, we had Leopold's essay on the trillion-dollar cluster.

We're not quite there yet. We have multiple labs, um, you know, xAI very famously, you know, Jensen Huang praising them for being best boy number one in, uh, spinning up a hundred thousand GPU cluster in like twelve days or something.

**Alessio** [39:36]
Mm-hmm.

**Swyx** [39:36]
So likewise at Meta, likewise at OpenAI, likewise at the other labs as well. So, like, the GPU ultra rich are going to keep doing that because I think partially it's an article of faith now that you just need it.

Like you d- you don't even know what it's going-- what you're gonna use it for. You just, you just need it. And it makes sense that if-- especially if we're going into more research-y territory than we are. So let's say 2020 into 2023 was let's scale big models territory-

**Alessio** [40:02]
Mm-hmm

**Swyx** [40:02]
... 'cause we had GPT-3 in 2020, and we're like, okay, we go from 125B to 1.8B, uh, uh, 1.8T, and that was GPT-3 to GPT-4. Okay, now that's done. Like, and as far as everyone is concerned, Claude, you know, Opus 3.5 is not coming out.

GPT-4.5 is not coming out. And Gemini 2, like we, we don't have Pro, whatever. We've hit that wall, whatever-

**Alessio** [40:24]
Yeah

**Swyx** [40:24]
... that wall is. Maybe I'll call it like the two trillion parameter wall. Like we're not going to 10 trillion.

**Alessio** [40:29]
Mm-hmm.

**Swyx** [40:29]
Like it's just n- like no one thinks it's a good idea, at least from training costs, from, uh, amount of data, or at least the inference para-- Like would you pay 10X-

**Alessio** [40:38]
Right

**Swyx** [40:38]
... the price of-

**Alessio** [40:39]
Yeah, yeah, yeah

**Swyx** [40:39]
... GPT-4? Probably not. Like

**Alessio** [40:41]
Mm-hmm.

**Swyx** [40:42]
Like you want something else that, that is at least more useful. So it makes sense that people are pivoting in, in terms of the inference paradigm. And so when it's more research-y, then you actually need more just general purpose compute to mess around with, uh, at the exact same time that production deployments of the o- the previous paradigm are still ramping up-

**Alessio** [40:58]
Mm-hmm

**Swyx** [40:58]
... um, uh, pretty aggressively. So it, it makes sense that the GPU rich are, are growing. We have now interviewed both Together and Fireworks and Replicate. Uh, we haven't done any scale yet, but I, I think Amazon maybe kind of a sleeper one, Amazon-

**Alessio** [41:11]
Mm

**Swyx** [41:11]
... in a sense of like they at Reinvent, I wasn't expecting them to do so well, but they are now a foundation model lab.

**Alessio** [41:18]
Mm-hmm.

**Swyx** [41:18]
It's kind of interesting. Um, I think, uh, you know, David went over there and, like, started just training models.

**Alessio** [41:26]
Yeah. I mean, that's the power of prepaid contracts. I think like a lot of AWS customers, you know, they do these big reserve instance contracts, and now they gotta use their money.

**Swyx** [41:35]
Mm.

**Alessio** [41:35]
Um, that's why so many startups get bought through the AWS marketplace, so they can kind of bundle them together-

**Swyx** [41:41]
Yeah

**Alessio** [41:41]
... in preferred pricing.

**Swyx** [41:42]
Okay, so maybe GPU super rich doing very well, GPU middle class dead, and then-

**Alessio** [41:48]
Yeah

**Swyx** [41:48]
... GPU poor.

**Alessio** [41:49]
I mean, my thinking is like everybody should just be GPU rich.

**Swyx** [41:52]
Huh?

**Alessio** [41:53]
There shouldn't really be... Even the GPU poor is like does it really make sense to be GPU poor? Like if you're GPU poor, you should just use the cloud models.

**Swyx** [42:01]
Yes.

**Alessio** [42:01]
You know? And I think there might be a future once we kinda like figure out what the size and shape of these models is, where like the tiny box and these things come to fruition, where like you can be GPU poor at home.

But I think today it's like why are you working so hard to like get these models to run on like very small clusters where it's like it's so cheap to run the

**Swyx** [42:22]
Yeah. Yeah.

**Alessio** [42:23]
You know?

**Swyx** [42:23]
Yeah, yeah. I think mostly p- people, people think it's cool. People think it's a-

**Alessio** [42:27]
Yeah

**Swyx** [42:27]
... it's a stepping stone to scaling up. So they, they aspire to be GPU rich one day, and they, they're working on new methods. Like, news research, like, probably the, the most deep tech thing they've done this year is, um, distro or whatever the new name is.

**Alessio** [42:39]
Yeah, yeah.

**Swyx** [42:39]
There's a lot of interest in heterogeneous computing, distributed computing. I tend generally to de-emphasize that historically, but it may be coming to a time where it is starting to be relevant. I don't know. You know, SF Compute launched their compute marketplace this year, and, like, who's really using that?

Like, it's a bunch of dis- small-

**Alessio** [42:55]
Disparate

**Swyx** [42:56]
... clusters-

**Alessio** [42:56]
Yeah

**Swyx** [42:56]
... disparate, uh, types of compute. And if you can make that useful, then that, that will be very beneficial to the broader community, but maybe still not the source of frontier models.

**Alessio** [43:08]
Yeah.

**Swyx** [43:08]
Right? It is just going to be a second tier of compute that is unlocked for people.

**Alessio** [43:12]
Yeah.

**Swyx** [43:12]
And that's fine. But yeah. I mean, I, I think this year I would say a lot more on device. We are-- W- I, I now have Apple Intelligence on my phone. Doesn't do anything apart from summarize my notifications, but still not bad.

Like, it's multimodal.

**Alessio** [43:26]
Yeah. The notification summaries are so and so-

**Swyx** [43:29]
Sometimes they're funny

**Alessio** [43:29]
... in my experience.

**Swyx** [43:30]
Yeah, but they, they add, they add juice to life. And then, um, Chrome Nano, uh, Gemini Nano is coming out in Chrome.

**Alessio** [43:35]
Yeah.

**Swyx** [43:35]
Uh, they're, they're still feature flagged, but you can s- you can try it now if you, if you use the, uh, the alpha. And so, like, I, I think, like, we- we're, we're getting the sort of GPU poor version of, of a lot of these things coming out, and I, I think it's, like, quite useful.

Like, Windows as, as well, rolling out RWKV in sort of every Windows deployment is, is super cool. Um, and then I think the last thing I-- that I never put in this GPU poor war that I think I should now-

**Alessio** [44:00]
Mm

**Swyx** [44:00]
... is the number of startups that are GPU poor but still scaling very well as sort of wrappers on top of either a foundation model lab or a GPU cloud. A GPU cloud, it would be Suno. Suno Ramp has rated as one of the top-ranked gr- fastest growing startups of the year.

Um, I think the last public number is, like, zero to twenty million this year in ARR, and Suno runs on Model.

**Alessio** [44:24]
Mm.

**Swyx** [44:24]
So Suno itself is not GPU rich, but they're just doing their training on, on Model, uh, w- who we've also talked to on, on the podcast. The other one would be Bolt, which is a straight cloud wrapper.

And, and, uh, a- again, another-- uh, now they've announced twenty million ARR, which is another up-

**Alessio** [44:43]
Yeah

**Swyx** [44:43]
... step up from our eight million-

**Alessio** [44:44]
Yeah

**Swyx** [44:44]
... that we, that we put on the, the pod- the title. So yeah. I mean, it's crazy that all these GPU poors are finding a way while the GPU riches are also finding a way, and then the, the only failures...

I, I kinda call this the GPU smiling curve-

**Alessio** [44:56]
Mm-hmm

**Swyx** [44:56]
... uh, where the, the, the edges do well 'cause you're either close to the machines and you're, you're, like, number one on the machines, or you're, like, close to the customers and you're number one on customer side.

**Alessio** [45:03]
Mm-hmm.

**Swyx** [45:04]
And the people who are in the middle, inflection, um, character-

**Alessio** [45:08]
Yeah

**Swyx** [45:08]
... didn't, didn't do that great. I think character did the best out of all, all of them.

**Alessio** [45:11]
Mm-hmm.

**Swyx** [45:12]
Um, like, you have a note in here that we apparently said that character's price tag was one B.

**Alessio** [45:16]
Yeah.

**Swyx** [45:16]
Did I say that?

**Alessio** [45:17]
Yeah.

**Swyx** [45:17]
Oh.

**Alessio** [45:17]
You said Google should just buy them for one B.

**Swyx** [45:19]
Yeah, yeah.

**Alessio** [45:20]
I thought it was a crazy number.

**Swyx** [45:20]
Yeah.

**Alessio** [45:20]
Then they paid two point seven billion.

**Swyx** [45:22]
I mean, for, like... Yeah, what do you pay for now? Like, I, I don't know what-

**Alessio** [45:24]
Yeah

**Swyx** [45:24]
... the bidding war was like. Maybe the starting price was one B.

**Alessio** [45:28]
I mean, whatever it was, it worked out- ... for everybody involved. Multimodality war. In this one we never had text to video in the first version-

### Multimodality War

**Swyx** [45:35]
Huh

**Alessio** [45:36]
... which now is the hottest.

**Swyx** [45:38]
I-- yeah, I would say it's a subset of image, but yes.

**Alessio** [45:40]
Yeah. Well, but I think at the time it wasn't really-

**Swyx** [45:43]
Mm-hmm

**Alessio** [45:43]
... something people were doing, and now we had Veo two just came out yesterday.

**Swyx** [45:46]
Mm-hmm.

**Alessio** [45:47]
Uh, Sora was released last month-

**Swyx** [45:49]
Yeah

**Alessio** [45:49]
... last week.

**Swyx** [45:49]
Have you tried Sora?

**Alessio** [45:50]
I've not tried Sora because the, the, the day that I tried-

**Swyx** [45:52]
Yeah, it's-

**Alessio** [45:53]
It wasn't-

**Swyx** [45:54]
We wasn't-

**Alessio** [45:54]
Yeah.

**Swyx** [45:54]
I think it's generally available now. You can tr- you can go to sora.com and try it.

**Alessio** [45:57]
Yeah. And then, yeah-

**Swyx** [45:58]
Um-

**Alessio** [45:58]
... yeah, the, they had the outage, which I think also-

**Swyx** [46:01]
Yeah

**Alessio** [46:01]
... played a part into it. Um-

**Swyx** [46:03]
Small things. Just-

**Alessio** [46:04]
Yeah, yeah

**Swyx** [46:04]
... minor point.

**Alessio** [46:05]
What's the other model that you posted today that was on Replicate video o1-live?

**Swyx** [46:09]
Yeah. Very, very-

**Alessio** [46:10]
There's just a lot

**Swyx** [46:10]
... nondescript, uh, name-

**Alessio** [46:12]
Yeah

**Swyx** [46:12]
... but it is from Minimax, which I think is a Chinese lab. The Chinese labs do surprisingly well at the, at the video models. I, I'm not sure it's actually Chinese. I, I, I don't-- don't hold me up to that.

Yep, China.

**Alessio** [46:27]
Well, no, it's good.

**Swyx** [46:28]
So Hailuo. Yeah. Chinese love video. I-- what can I say? They, they have a lot of training data for video. Or a, a more, more relaxed-

**Alessio** [46:35]
That's true

**Swyx** [46:35]
... uh, regulatory environment.

**Alessio** [46:37]
Uh, well sure, in some way. Yeah. I don't think there's much else there. I think, like, you know, on the image side, I think-

**Swyx** [46:45]
Um-

**Alessio** [46:45]
... it's still open-

**Swyx** [46:46]
Yeah, I mean-

**Alessio** [46:46]
... open war

**Swyx** [46:47]
... uh, uh, Eleven Labs now a unicorn.

**Alessio** [46:48]
Yeah.

**Swyx** [46:49]
Uh, so, so basically what is multimodality war? Multimodality war is do you specialize in a single modality? Right? Or do you have God model that does all the modalities, right? So this is definitely still going-

**Alessio** [47:01]
Yeah

**Swyx** [47:01]
... in a sense of Eleven Labs, uh, you know, now a unicorn. Pika Labs doing, doing well. They launched Pika two point oh recently.

**Alessio** [47:08]
Yeah.

**Swyx** [47:08]
HeyGen, I think, I think has reached a hundred million ARR. Assembly, I don't know, but they, they have billboards all over the place, uh, so I assume they're, they're doing very, very well. So these are all specialist mo- specialist models and specialist startups, and, uh-

**Alessio** [47:19]
And, yeah

**Swyx** [47:19]
... they seem to be doing well.

**Alessio** [47:20]
And product especially.

**Swyx** [47:21]
Yeah. And then there's the big labs who are doing the, the sort of all-in-one play. And then here I would highlight Gemini two for having native image output. Have you seen the demos? Um-

**Alessio** [47:32]
No.

**Swyx** [47:32]
Yeah, it's, it's hard to keep up. Uh-

**Alessio** [47:33]
Yeah

**Swyx** [47:33]
... literally they launched this last week. And, uh, shout-out to Paige Bailey, who came to the Latent Space event to demo on the day-

**Alessio** [47:39]
Nice

**Swyx** [47:39]
... of launch. And she wasn't prepared. She was just like, "I'm just gonna show you." So they have voice. They have, you know, obviously image input, and then they obviously can code gen and all that. But the new one that OpenAI and Meta both have, but they haven't launched yet-

**Alessio** [47:54]
Mm

**Swyx** [47:55]
... is image output. So you can literally, um... A- I think their demo video was that you put in an image of a car and you ask for minor modifications to that car, they can generate you that modification-

**Alessio** [48:05]
Mm

**Swyx** [48:05]
... exactly as you asked. So there's no need for the stable diffusion or comfy UI workflow-

**Alessio** [48:11]
Yeah, yeah, yeah

**Swyx** [48:11]
... of like mask here, and then like infill there, inpaint there, and all that, all that stuff. This is small model nonsense.

**Alessio** [48:17]
Yeah.

**Swyx** [48:17]
Big model people are like, "Ha. We got you in, in as everything in the transformer." This is the multimodality war, which is do you-

**Alessio** [48:24]
Mm

**Swyx** [48:24]
... do you bet on the god model or do you string together a whole bunch of small models like a, like a chump?

**Alessio** [48:29]
Yeah. I don't know, man. Yeah, that would be interesting. I mean, obviously, I use Midjourney for all of our thumbnails. Um-

**Swyx** [48:35]
Yes. Still-

**Alessio** [48:36]
They launch-

**Swyx** [48:36]
Still Sora.

**Alessio** [48:37]
They've been doing a ton on the product, I would say. They launched a new Midjourney editor thing. They've been doing a ton because I think, yeah, the model is kinda like maybe, you know, people say Black Forest, the Black Forest models are better than Midjourney on a pixel-by-pixel basis, but I think when you put it, put it together, um-

**Swyx** [48:53]
Have you tried the same prompts on Black Forest?

**Alessio** [48:55]
Yes, but the problem is just like, you know, on Black Forest it generates one image, and then it's like you gotta regenerate. You don't have all these like UI, UI thing. Like, what I do-

**Swyx** [49:03]
Skill issue, bro.

**Alessio** [49:04]
No, but i- it's like time issue, you know? It's like-

**Swyx** [49:06]
Yeah, so-

**Alessio** [49:07]
On Midjourney-

**Swyx** [49:07]
Call the API four times

**Alessio** [49:09]
... no, but then there's no, like, variate. Like- ... the, the good thing about Midjourney is, like, you just go in there-

**Swyx** [49:15]
Yeah

**Alessio** [49:15]
... and you're cooking. There's a lot of stuff that just makes it really, really easy.

**Swyx** [49:18]
Okay, yeah, yeah.

**Alessio** [49:18]
And I think people underestimate that. Like, it's not really a skill issue because I'm paying Midjourney, so it's a-

**Swyx** [49:23]
Yeah, yeah

**Alessio** [49:23]
... Black Forest skill issue because I'm not paying them.

**Swyx** [49:24]
Correct.

**Alessio** [49:25]
You know? It's a-

**Swyx** [49:25]
Yeah. So okay. So, uh, this is a UX thing, right? Like, you, you, you, you understand that at least, uh, we think that Black Forest should be able to do all that stuff.

**Alessio** [49:33]
Yeah.

**Swyx** [49:33]
I would also shout out Recraft has come out, uh, on top of the image arena that, uh, Artificial Analysis has done. Has apparently taken Flux's place. Is this still true? So Artificial Analysis is now a company. Uh, we-- I've, I highlighted them I think in, in one of the early AI News' of the year, and they have a whole-- launched a whole bunch of a- arenas.

So they're trying to take on LM Arena, Anastasios and crew, and they have an image arena. Oh yeah, Recraft V3 has now, uh, beaten Flux 1.1, which is very surprising 'cause Flux and Black Forest Labs are the old Stable Diffusion crew who left Stability after, um, the management issues.

So Recraft has come from nowhere to be the top image model. Uh, very, very strange. I would also highlight that Groq has now launched Aurora, which is-- It's very interesting dynamics between Groq and Black Forest Labs because Groq's images were originally launched, uh, in partnership with Black Forest Labs as a-

**Alessio** [50:24]
Yeah

**Swyx** [50:24]
... as a thin wrapper, and then Groq was like, "No, we, we'll make our own." Um, and so they, they've, they've made their own. I, I don't know. There are no, there are no APIs or benchmarks about it.

They just announced it. So yeah, that's the multimodality war. I would say that so far the small model-- the dedicated model people are winning because they are just focused on their tasks. But the big model people are always catching up.

And the moment I saw the Gemini 2 demo of image editing where I can put in an image and just request it and it does, like, that's how AI should work, not like a whole bunch of complicated steps.

So it really is something. And, uh, I think one frontier that we haven't seen this year, like obviously video has done very well and it'll continue to grow. You know, we, when we ha- have the release of Sora Turbo today, but at some point we'll get full Sora or at least the Hollywood labs will get full Sora.

We haven't seen video to audio or video synced to audio, and so the researchers that I talk to are already starting to talk about that as the next frontier. But there's still maybe like, uh, five more years of video left to-

**Alessio** [51:22]
Mm

**Swyx** [51:22]
... actually be Sora. I would say that Gemini's approach compared to OpenAI, Gemini seems-- or DeepMind's approach to video seems a lot fully-- more fully fledged than OpenAI because if you look at the ICML recap that I published that so far nobody has listened to

that people have listened to it. It's just a different-- definitely a different audience.

**Alessio** [51:43]
It's only a seven-hour slot.

**Swyx** [51:44]
It's only seven hours.

**Alessio** [51:44]
Why are people not listening?

**Swyx** [51:46]
It's like everything in, in one-

**Alessio** [51:47]
Please.

**Swyx** [51:48]
Uh, so, so DeepMind has-- is working on Genie. They also launched Genie 2 and Video Poet.

**Alessio** [51:53]
Yeah.

**Swyx** [51:53]
So, like, they have maybe four years advantage on world modeling that OpenAI does not have.

**Alessio** [51:59]
Yeah, yeah.

**Swyx** [51:59]
'Cause OpenAI basically only started Diffusion Transformers last year when they hired, uh, Bill Peoples. So DeepMind has, has a bit of an advantage here-

**Alessio** [52:06]
Yeah

**Swyx** [52:06]
... I, I would say in, in, in, in showing. Like, the reason that Veo 2, well, one, they, they cherry-pick their videos, so obviously it looks better than Sora. But the reason I would believe that Veo 2, uh, when it's fully launched will do very well is because they have all this background work in-

**Alessio** [52:21]
Yeah

**Swyx** [52:21]
... video that they've done for years. Like, like last year's NeurIPS, I already in-- was interviewing some of their vid- video people. I forget the model name. But for, for people who are dedicated fans they can go to Neuro 2023 and see, see that paper.

**Alessio** [52:32]
Mm-hmm. And then last but not least, the LLMOS slash-

**Swyx** [52:36]
Yeah. We renamed this one

### LLMOS War

**Alessio** [52:37]
... Rag- RagOps.

**Swyx** [52:38]
Yeah, yeah.

**Alessio** [52:38]
Formerly known as RagOps 4.

**Swyx** [52:41]
Yeah. I put the latest, uh, chart on the Brain Trust episode. I, I think I'm gonna separate the-- these essays from the episode notes.

**Alessio** [52:48]
Yep.

**Swyx** [52:48]
So the reason I used to do that, by the way, is 'cause I wanted to show up on Hacker News. I want the podcast to show up on Hacker News.

**Alessio** [52:52]
Mm-hmm.

**Swyx** [52:53]
Right? So I always put a essay inside of there because Hacker News people-

**Alessio** [52:56]
Mm-hmm

**Swyx** [52:56]
... like to read and not listen.

**Alessio** [52:57]
Yeah.

**Swyx** [52:58]
But-

**Alessio** [52:58]
So episode essays-

**Swyx** [52:59]
Yeah.

**Alessio** [53:00]
... I don't know. We're just doing them separately.

**Swyx** [53:01]
You say LangChain, LlamaIndex still growing.

**Alessio** [53:03]
Yeah. So I looked at the PyPi stats. You know? I don't care about stars. Um, on PyPi you see LangChain-

**Swyx** [53:10]
Do you wanna share your screen?

**Alessio** [53:11]
Um, yes. I prefer to look at actual downloads, not at, um, stars on GitHub. So if you look at, you know, LangChain still growing. These are the last, last six months. LlamaIndex still growing. What I've basically seen is like things that, one, obviously these things have a commercial product, so there's like people buying this and sticking with it versus kinda hopping in between things versus, you know, for example, CrewAI not really growing as much.

The, the stars are growing. If you look on GitHub, like the stars are growing, but kinda like the usage is-

**Swyx** [53:43]
Mm

**Alessio** [53:43]
... kinda like flat in the last six months. So-

**Swyx** [53:45]
Have they done some kind of a reorg where they, they did like a split of packages and now it's like a bundle of packages? Sometimes that happens, you know.

**Alessio** [53:53]
I didn't see that.

**Swyx** [53:55]
I can see both. I can, I can see both happening.

**Alessio** [53:57]
Oh.

**Swyx** [53:57]
That CrewAI is, is very loud but, but not used. And then-

**Alessio** [54:01]
Yeah, but an- anyway, to me it just like-

**Swyx** [54:03]
Yeah, there's no split

**Alessio** [54:04]
... the split-

**Swyx** [54:04]
There's a split

**Alessio** [54:04]
... I mean, Auto-- similar with Auto-GPT, it's like there's still a wait list for Auto-GPT to be used. Um, like the, the cloud-

**Swyx** [54:11]
Yeah. They, they, they're still, they're still kicking. They, they-

**Alessio** [54:13]
Yeah

**Swyx** [54:13]
... they announced some stuff recently. But, um-

**Alessio** [54:15]
But I think that's another one where it's the fastest growing project in the history of GitHub. But I think, you know, when, when you maybe like run the numbers on like the value of the stars and like the value of the hype, I think in AI you see this a lot, which is like a lot of stars, a lot of interest at a rate that you didn't really see in the past in open source where nobody's running to star- ...

uh, you know, a NoSQL database.

**Swyx** [54:35]
Yeah.

**Alessio** [54:35]
It's kinda like just the people that actually use it.

**Swyx** [54:37]
Yeah.

**Alessio** [54:38]
Um-

**Swyx** [54:38]
I think one thing that's interesting here, and one obviously is that in AI you kinda get paid to promise things.

**Alessio** [54:44]
Yeah.

**Swyx** [54:44]
And then you-- to deliver them, you know, people have a lot of patience. I think that patience has come down o- over time. One example here is Devin, right, this year, where a lot of promise in March and then, and then it took nine months to get to GA.

Uh, but I think people are still coming around now on Devin. Devin's product has improved a little bit, and even you're gonna be a paying customer. So I think something Devin-like will work. I don't know if it's Devin itself.

The AutoGPT has an interesting second layer in terms of what I think is the dynamics going on here, which is a very AI-specific layer.

**Alessio** [55:16]
Mm.

**Swyx** [55:16]
Over-promising, under-delivering is a-- applies to any startup. But for AI specifically, there's this promise of generality-

**Alessio** [55:23]
Yeah

**Swyx** [55:23]
... that I can do anything, right? So AutoGPT's initial problem was making money. Like increase my net worth. And I think that means that there's a lot of broad interest from a lot of different people who are trying to do all different things on this one project.

So that's why there's, there's concentrates a lot of stars. And then obviously, because it does too much maybe or it's not focused enough, then it does-- it, it fails to deploy. So that would be my explanation for why the interest-to-usage ratio is so low.

And the, the second one is obviously pure execution, like the, the team needs to have a vision and execute, like half the core team left right after-

**Alessio** [55:57]
Yeah

**Swyx** [55:57]
... uh, AI Engineer Summit last year.

Um, that would be my explanation as to why-

**Alessio** [56:03]
Mm-hmm

**Swyx** [56:03]
... like this promise of generality works basically only for ChatGPT.

**Alessio** [56:07]
Right.

**Swyx** [56:07]
And maybe for this year NotebookLM.

**Alessio** [56:09]
Mm.

**Swyx** [56:10]
Like sticking anything in there, it'll mostly be correct.

**Alessio** [56:13]
Yeah.

**Swyx** [56:13]
And then for basically everyone else, it's like, y- you know, we will help you complete code. We will help you with your PR reviews.

**Alessio** [56:19]
Mm-hmm.

**Swyx** [56:19]
Like small things.

**Alessio** [56:20]
Yeah, yeah. All right, code interpreting, we talked about a bunch of times. We soft announced the E2B fundraising, uh, on this podcast. CodeSandbox got acquired by Together AI last week-

**Swyx** [56:34]
Mm-hmm

**Alessio** [56:34]
... um, which they're now also gonna offer as an API. So, uh, more and more activity, which is great. Yeah, and then, uh, in the last ep-- two episodes ago with Bolt, we talked about the web container stuff they've been working on.

I think like there's maybe the spectrum of code interpreting, which is like, you know, dedicated SDK. There's like, yeah, the models of the world, which is like, hey, we got a sandbox. Now you just kinda run the commands and orchestrate all of that.

I think this is one of the-- I mean, E2B's growth has just been crazy just because, I mean, everybody needs to run code, right? And I think now all the products and the-- everybody's graduating to like, okay, it's not enough to just do chat.

So Perplexity, which is a E2B customers, they do all these nice charts for like finance and all these different things. It's like the products are maturing, and I think this is becoming more and more of kinda like a hair-on-fire problem, so to speak.

So yeah, excited to see more. And this was one that really wasn't on the radar when we first wrote the four wars.

**Swyx** [57:34]
Yeah. I think mostly because I was trying to limit it to RAG and ops.

**Alessio** [57:38]
Yeah.

**Swyx** [57:38]
But I think now that the frontier has expanded in terms of the core set of tools. Core set of tools would include code interpreting, like, like tools that every agent needs, right?

**Alessio** [57:50]
Mm-hmm.

**Swyx** [57:50]
And Gra-Graham in his, uh, State of Agents talk had this as well, which is kind of interesting, uh, for me, because like everyone finds the same set of things. So it's, it's basically like someone-- everyone needs web browsing, everyone needs code interpreting, and then everyone needs some kind of memory or planning or, um, whatever-

**Alessio** [58:08]
Mm

**Swyx** [58:08]
... whatever that is. We'll discover this more over time, but I think this is what we've discovered so far.

**Alessio** [58:12]
Yep.

**Swyx** [58:12]
Um, I will also call out, uh, Morphe labs for launching a time travel VM.

**Alessio** [58:17]
Mm-hmm.

**Swyx** [58:17]
I think that basically the statefulness of these things needs to be locked down a lot. Basically, you can't just spin up a VM, uh, run code on it, and then kill it. Because sometimes you might need to un-re-- time travel back, like unwind or fork to explore different paths for sort of like a tree search approach to your agent development.

I would call out the, the newer ones, the, the new implementations as the emerging frontier in terms of like what people kinda are going to need for agents to do very fan-out approaches to, to all these sort of code execution.

Um, and then I'll also call out that I think ChatGPT Canvas with, uh, what they launched in the 12 days of Shipmas that they announced, has surprisingly superseded Code Interpreter. Like Code Interpreter was last year's thing, and now Canvas can also write code, and they also run code and do more than Code Interpreter used to do.

So right now it has not killed it. So there's, there's a toggle box for Canvas and for Code Interpreter when you create a new custom GPTs. You know, my, my whole thesis that custom GPTs is your roadmap for investing-

**Alessio** [59:16]
Mm-hmm

**Swyx** [59:16]
... 'cause it's, it's what everyone needs. So now there's a new box called Canvas that everyone can-- has access to. But basically, there's no reason why you should use Code Interpreter over Canvas. Like Canvas has-- incorporates the diff mode that both Anthropic and OpenAI and Fireworks has now shipped.

That I think is going to be the norm for next year, uh, that e-everyone needs some kind of diff mode Code Interpreter thing.

**Alessio** [59:39]
Mm-hmm.

**Swyx** [59:39]
Like Aider was also very early to this. Like the Aider benchmarks-

**Alessio** [59:42]
Yeah

**Swyx** [59:42]
... were also all, all based on diffs and Cursor as well.

**Alessio** [59:45]
Yep. You wanna talk about memory?

**Swyx** [59:47]
Memory. Uh, you think it's not real.

**Alessio** [59:49]
Yeah, I just don't. I, I think most memory product today just like a summarization and extraction.

**Swyx** [59:57]
Yeah.

**Alessio** [59:57]
I don't think there is-

**Swyx** [59:57]
They're very immature.

**Alessio** [59:59]
Yeah. There's no implicit memory, you know. It's all explicit memory of what you've read in. There's no implicit extraction of like, oh, you said no to this, you said no to this 10 times, so you don't like going on hikes at 6 a.m.

Like it doesn't-- none of the memory products do that. They'll summarize what you say explicitly. Um-

**Swyx** [1:00:18]
When you say memory products, you mean the, the startups that are more offering memory as a service?

**Alessio** [1:00:22]
Yeah. Or even like, you know, Linde has like memories, you know. It's like based on what I say, it remembers it. So it's less about making an actual memory of my preference.

**Swyx** [1:00:32]
Yeah.

**Alessio** [1:00:32]
It's more about what I explicitly said. Um, and I'm trying to figure out at what level that gets solved. You know, like is it-- do these memory products like the memGPTs of the world create a better way to implicitly extract preference or can that be done very well, you know?

I think that's why I don't think... It's not that I don't think memory is real, I just don't think that like the approaches today are like actually memory are what you need a system to have.

**Swyx** [1:00:58]
Yeah. I would actually agree with that. Um, but I would just point it to it being immature rather than, uh, not needed. Like it's clearly-

**Alessio** [1:01:05]
Mm-hmm

**Swyx** [1:01:05]
... something that we will want at some point. And so the people developing it now are, you know, not very good at it. And I would definitely predict that next year will be better, and the year after that will be better than that.

I definitely think that last time we had the Shunyu pod with Harrison as guest host-

**Alessio** [1:01:22]
Mm-hmm

**Swyx** [1:01:22]
... I over-focused on LangMem as, as a separate product. He has now rolled it into LangGraph as, as a memory service, the same API. And I, I, I mean, I, I think that everyone will need some kind of memory, and I think that this is-- has distinguished itself now as a separate need from a normal rag vector database.

Like you will need a memory layer, whether it's on top of a vector database or not, it's up to you. A memory database and a vector database are kind of two different things. Like I've had to justify this so much actually that I have a draft post in the, in the In Space dashboard that, uh, basically says like, what is the difference between memory and knowledge?

**Alessio** [1:01:54]
Mm-hmm.

**Swyx** [1:01:54]
And to me it's very clear. It's like knowledge is about the world around you and like there's knowledge that you have, which is the, the rag corpus-

**Alessio** [1:02:00]
Yep

**Swyx** [1:02:00]
... that your- maybe your company docs or whatever. And then there's external knowledge, which is the stuff that you Google, so like you use something like EXA, whatever. And then there's memory, which is my interactions with you over time.

Both can be represented by vector databases or knowledge graphs, doesn't really matter. Time is a specifically important one in memory-

**Alessio** [1:02:19]
Mm-hmm

**Swyx** [1:02:19]
... because you need a decay function, and then you also need like a review, uh, function.

**Alessio** [1:02:22]
Yeah.

**Swyx** [1:02:23]
A lot of pe- people are implementing this as sleep. Like when you sleep, you like literally you sort of process the day's memories and you-

**Alessio** [1:02:29]
Mm

**Swyx** [1:02:29]
... come up with new insights that, that you then persist and bring into context in the future. So I feel like this is being developed. LangGraph has a version of this. Zepp is another one that's based on Neo4j's Knowledge Graph that, that has a version of this.

Uh, memGPT used to have this, but I think I feel like Letta, since it was funded by Quiet Capital, has broadened out into more of a sort of general LLMOS type startup, which I feel like there's a bunch of those now.

There's, there's all hands in all this.

**Alessio** [1:02:56]
Do you think this is a LLMOS product or should it be a consumer product?

**Swyx** [1:03:00]
I think it's a building block. I think every-- I think there should-- The, uh-- Just like every consumer product is going to have a-- going to eventually want a, a gateway, you know, for, for managing their requests and ops tool, you know, that kind of stuff.

Um, code interpreter for maybe not exposing the code, but executing code under the hood for sure. So it's gonna want memory. It's gonna want long live memory. So as a consumer, let's say you are a new dot computer who, um, you know, they've, they've, they've launched their own, um, little agents, or if you're a friend.com, you're going to want to invest in memory at some point.

Maybe it's not today. Maybe you can push it off a lot further with like a million token context. But at some point you need to compress your memory and to selectively retrieve it.

**Alessio** [1:03:43]
Mm.

**Swyx** [1:03:43]
And then what are you gonna do? You, you ha- you have to reinvent the whole memory stack, and these guys have been doing it for a year now.

**Alessio** [1:03:50]
Yeah. To me it's more like I wanna bring the memories. Like it's almost like they're my memories, right? So why-

**Swyx** [1:03:56]
So you selectively, selectively choose the memory you want to bring in.

**Alessio** [1:03:58]
Yeah. Why does, why does every time that I go to a new product it needs to relearn everything about me?

**Swyx** [1:04:02]
Okay. You want portable memories.

**Alessio** [1:04:03]
Yeah. Is it like a protocol? Like how does that work?

**Swyx** [1:04:07]
Uh, speaking of protocols, Anthropic's model context protocol that they launched has a 300 line of code memory implementation.

**Alessio** [1:04:13]
Mm.

**Swyx** [1:04:13]
Very simple. Very bad news for all the memory startups.

But that's all you need. And, uh, yeah, it would be nice to have a portable memory of you to ship to everyone else. Simple answer is there's no standardization for a while because everyone will experiment with their own stuff.

And I think Anthropic's success with MCP s- suggests that basically no one else but the big labs can do it because no one else has the, the sway-

**Alessio** [1:04:37]
Yeah, yeah, yeah

**Swyx** [1:04:37]
... to do this. Then that's, that's how it's gonna be. Uh, like unless you have something silly like... Okay. Some-- One form of standardization basically came from Georgi Gerganov with Llama CPP.

**Alessio** [1:04:50]
Yeah.

**Swyx** [1:04:50]
Right? And that was p- completely open source, completely bottoms up, and that's because there's just this infinite amount of work that needed to be done there, and then people built up from there. Another form of standardization is Comfy UI from Comfy Anonymous.

So like that kind of standardization can be done. So someone basically has to create that for the role play community-

**Alessio** [1:05:09]
Mm

**Swyx** [1:05:09]
... because those are the people with the longest memories. Right now, the role play community, as far as I understand, I've looked at Silly Tavern, I've looked at Cobalt, they only share character cards, and there's like four or five different standardized standard versions of this or these character cards, but nobody has e- exportable memory yet.

If there was anyone that developed memory first that became a standard, it would be those guys.

**Alessio** [1:05:29]
Cool. I'm excited-

**Swyx** [1:05:30]
Yeah

**Alessio** [1:05:30]
... to see what people build. Benchmarks.

**Swyx** [1:05:33]
Okay.

### Benchmarks

**Alessio** [1:05:33]
One of our favorite pet topics.

**Swyx** [1:05:35]
Uh, yeah, yeah. Um, so basically I just wanted to mention this briefly. Like, um, I think that in a year end of year review, it's useful to remind everybody where we were. So we talked about how in LMSis ELO everyone has gone up, and it's a very close race, and I think benchmarks as well.

I was looking at the OpenAI live stream today when they introduced o1 API with structured output and everything, and the benchmarks they're talking about are like completely different than the benchmarks that we were talking about, uh, this time last year.

This time last year, we were, we were still talking about MMLU, little bit of distill like GSM8K. There's stuff that's, uh, basically in V1 of the Hugging Face open models leaderboard, right? We talked to Clementine, uh, about the decisions that she made to upgrade to V2.

I would s- also say LMSys, now Ella Marina, also has emerged this year as, as a, as the leading like battlegrounds between the, the big frontier labs. But also we have also seen like the emergence of Swebench, LiveBench, MMLU Pro, and AIME.

AIME specifically for o1. It will be interesting to see like top most cited benchmarks of the year from 2020 to 2021, 2, 3, 4, and then going to 5, and you can see what has been saturated and solved and what people care about now.

And so now people care a lot about frontier math coding, right? There's literally a benchmark called frontier math, which I, I spent a bit of time talking about at NeurIPS. There's Amy, uh, there's LifeBench, there's MMLU Pro, and there's SweetBench.

I, I f- I feel like this is good. And then, um, there was another one this time last year it was GPQA. I'll, I'll put math and GPQA here, uh, as, as sort of top benchmarks of, of last year.

And NeurIPS GPQA was declared dead, which is very sad. Uh, people are still talking about GPQA Diamond. Uh, so literally the, the name of GPQA is called Google Proof Question Answering, so it's supposed to be resistant to saturation for a while.

Uh, and No Brown said that GPQA was dead. So now we only care about SweetBench, LifeBench, MMLU Pro, AIME. And even SweetBench, we don't care about SweetBench proper, we care about- Yeah ... SweetBench Verified. Verified. Uh, we s- we care about the SweetBench multimodal, and then we also care about the new Kawinsky Prize from Andy Kawinsky, which is the guy that we talked to yesterday, who has launched a similar sort of Arc AGI attempt on a SweetBench type- Hmm ...

metric, which arguably is a bit more useful. OpenAI also has MLEbench, uh, which is m- more tracking sort of, um, ML research and bootstrapping, which arguably like this is the key metric that is most relevant for the frontier labs, which is when the researchers can automate their own jobs.

So that is a kink in the acceleration curve, uh, if we were ever to reach that. Yeah, that makes sense. I mean, I'm curious, I think Dylan at the debate, he said SweetBench 80% was like his hope for end of next year as a kinda like, you know- Yeah ...

watermark that the models- Yeah ... are still improving, so. Yeah. I mean, keeping in mind we started the year at 13%. Yeah, exactly. And so now we're about 50. Um, OpenHands is around there. And yeah, 80 sounds fine.

Uh, Kawinsky Prize is 90. Yeah. And then as we get to 100, then the open source catches up . Oh, yeah. Oh, yeah. Once again . Magically gonna to close the gap between the closed source and open source.

So basically, I, I think my advice to people is, uh, keep track of the slow cooking of- Yeah ... benchmark language, because the labs that are not that frontier will, will keep measuring themselves on- Yeah ... last year's benchmarks, and then the labs that are actually frontier will tell you about benchmarks you've never heard of, and you'll be like, "Oh," like, "okay, there's, there's new, there's new territory to, to, to go on."

That would be the quick tip there. Yeah. Uh, maybe, maybe I won't, uh, belabor this point too much. I would also say maybe VO has introduced some new video benchmarks, right? Mm. Like basically every new frontier capabilities, and this is the next section that we're gonna going to go into, introduces new benchmarks.

We also briefly talked about Ruler as like the, the new- Yeah. Yeah, yeah ... sort of, uh, you know, last year we was like needle in a haystack, and Ruler- Yeah ... is basically n- mul- multidimensional needle in a haystack.

Yeah. We'll link all the episodes in the- Yeah ... description of the podcast ... this is like a review of all the episodes that we've done- Yeah ... which I have in my head. This is one of the slides that I did at my DevDay talk.

### Capabilities

**Swyx** [1:09:35]
So, so we're moving on from benchmarks to capabilities, and I think I have a useful categorization that I've been kinda sell... I, I'd be curious on your feedback or edits. I think there's basically f- like s- I, I kinda like the thought spot model of, um, what's mature, what's emerging, what's frontier, what's niche.

Mm. So mature is like stuff that you can just rely on in production. It's solved. Everyone has it. So what's solved is general knowledge, MMLU, and then what's solved is kinda long context. Yeah. Everyone has 100, 128K. Today, o1 announced 200K, um, which is very expensive.

I don't, I don't know- Yeah ... what the price is there. What's solved, kinda solved is RAG. Um, there's like 18 different kinds of RAG, but it's mostly solved. Batch transcription, I would say Whisper is, is- Mm ...

uh, something that you should be using on a, on a as much as possible, and then code generation and kind of solve. Uh, there's, there's different tiers of code generation, and I really need to split out- Yeah, yeah, yeah ...

single line of code- Mega ... complete- Yeah. Yeah ... versus multi-file, uh, generation. I think that is, that is definitely emerging. So on the emerging side, tool use, I would still consider emerging maybe, maybe more mature- Mm ...

already, but they only launched function output this year. Yeah, yeah, yeah. So, so next year- I think emerging is fine ... more mature. Uh, vision language models. Everyone has vision now, I think. In- yeah, including o1. Mm-hmm. So this is clear.

A, a subset of vision is PDF parsing. Um- Mm ... and I, I think the, uh, the community's very excited about the work being done with CodePaelli and CodeQuen. Yeah. What's for you the breakpoint for vision to go to mature?

I think, I think it's basically now . This is, uh, this is maybe two months old. Yeah, yeah, yeah. That's right. NVIDIA, most valuable company in the world. Uh, also I think this was, this was in June. Then also they surprised a lot on, on the upside for their, um, Q3 earnings.

I think the quote that I highlighted in AI News was that it is the best... like, like Blackwell is the best-selling series the, in, in the history of the company, and they're sold-- I mean, obviously they're almost sold out.

Yeah. But for him to make that statement, I think it's a, it's another indication that the transition from the H to the B series is gonna go very well. Yeah, the... I mean, if you had just bought NVIDIA when ChatGPT came out, that would be- Yeah ...

insane. Uh, you know, which one more, you know, NVIDIA or Bitcoin? I think, I think NVIDIA. I think in gains, yeah. Well, I think the question is, like, people ask me, like, is there-- what's the reason to not invest in NVIDIA?

I think it's really just like the, they have committed to this. They went from a two-year cycle to one-year cycle, right? Mm-hmm. And so it takes one misstep to delay. You know, like, there have been delays in the past and, like, when delays happen, they're- Yeah ...

typically very good buying opportunities. Anyway . Hey, this is Fix from the editing room. I actually am, uh, just realized that we lost about 15 minutes of audio and video that, uh, was in the episode that we shipped, and I'm just cutting it back in and re-recording.

We don't have time to re-record by, before the end of the year. It's December 31st already. So I'm just going to do my best to recover what we have and then sort of segue you in nicely to the end.

Uh, so our plan was basically to cover, like, what we felt was emerging capabilities, frontier capabilities, and niche capabilities. So emerging would be tool use, vision language models, which you just heard, real-time transcription, which, uh, I have, uh, on, uh, one of our upcoming episodes to be, as well as, uh, you, you can try it in Whisper Web GPU, which is amazing.

Uh, I think diarization capabilities are also maturing as well, but still way too hard to do properly. Like, we, we had to do a lot of stuff for the latent space transcripts to, to come out right. Um, I think maybe, you know, Dwarkesh recently has been talking about how he's using Gemini 2.0 Flash to do it, and I think that might be a good effort, a good way to do it.

And especially if there's crosstalk involved, that might be really good. But, uh, there might be other reasons to use normal diarization models as well. Specifically Pionote. Text and image we talked about a lot, so I'm just gonna skip.

And then, then we go to frontier, which, uh, you know, I think like basically I would say is on the horizon, but not quite ready for broad usage. Like it's, it's, you know, interesting to, to show off to people, but like we haven't really figured out how like the daily use, the, the, the large amount of money is gonna be made on long inference, on real-time interruptive, sort of real-time API voice mode things, on on-device models, as well as all the, all other modalities.

And then niche models, uh, niche capabilities. I always say like base models are very underrated. People always love talking to base models as well. Um, and we're increasingly getting less access to them. Uh, it's quite possible, I think, you know, Sam Altman for 2025 was like asking about what he should want-- what people want him to ship or what people want him to open source, and people really want GPT-3 base.

Uh, we may get it. We may get it. It's just for historical interest. Um, but, uh, you know, at this point, but there's... We, we, we may get it. Like it, it's definitely not, uh, a significant IP anymore for him.

So we'll see. Um, you know, I think OpenAI has a lot more things to worry about than shipping base models, but it would be a very, very nice things to do for the community. Um, state-space models as well.

I would say like the hype for state-space models this year, even though, um, you know, the post transformers talk at Latent Space Live was extremely hyped, uh, and very well attended and watched. Um, I would say like it feels like a step down this year.

I don't know why. Um, it seems like things are scaling out in state-space models, uh, and RWKVs. And so Cartesia I think is doing extremely well. Uh, we use them for a bunch of stuff, especially for small talks and, and some of our sort of NotebookLM podcast clones.

I, I think they're a real challenger to ElevenLabs as well. Um, and RB- RWKV, of course, is rolling out on Windows. So, um, I, I, I'll still, I'll still say these, these are niche. We've been talking about them as the future for a long time, and I mean, we live technically in a year in the future from last year, and we're still saying the exact same things as we were saying last year.

So what's changed? I don't know. Um, I do think the L- xLSTM paper, which we will cover when we cover the sort of NeurIPS papers, um, is worth a look. Um, I, I think they, they are very clear-eyed as to, um, how they want to fix LSTM.

Okay, so and then we al- also wanna cover a little bit, uh, like the major themes of the year, um, and then we, we wanted to go month by month. So I'll bridge you into back to the recording, which, uh, we still have the audio of.

So the ma- one of the major themes is sort of the inference race to the bottom. We started this, uh, last year, uh, this time last year with the Mistral price war of 2023, um, with, uh, Mis- Mistral going from $1.80 for per token down to $1.27, uh, in the, in the span of like a couple weeks.

And, um, you know, I think this, uh, a lot of people are also interested in the price war and sort of the price intelligence curve for this year as well. Um, I started tracking it, I think, roundabout in March of 2024 with, uh, Haiku's launch.

And so this is, uh, if you're watching the YouTube, this is what I initially charted out as like, here's the frontier. Like everyone's kind of like in a pretty tight range of LMSYS ELO versus the model pricing. You can pay more for more intelligence and you, and it'll be cheaper to get less intelligence, but roughly it correlates to, um, a- aligned, uh, and a tr- a trend line.

And then I could update it again in July and see that everything had kinda shifted right. Um, so for the same amount of ELO, let's say GPT-4 2023, uh, would, would be about sort of 1175 in ELO. You could...

And you, you used to get that for like $40 a, per token, uh, per million tokens. And now you get Claude 3 Haiku, which is about the same ELO, uh, for 50 cents. And so that's a, a two orders of magnitude improvement, um, in about two years.

Uh, sorry, in about a year. Um, but more, more importantly, I think, uh, you can see the more recent launches like Claude 3 Opus, which launched in March this year, um, now basically superseded completely, completely dominated by Gemini 1.5 Pro, which is both cheaper, $5 a month, uh, $5 per million, as well as smarter.

Uh, so it's about slightly higher than 1250 in ELO. Um, so the March frontier and shift to the July frontier is roughly one order of magnitude improvement per, uh, sort of ISO ELO. Um, and I think what you're starting to see now, uh, in July is the emergence of 4o Mini and DeepSeek V2 as outliers to the July frontier, where July frontier used to be maintained by 4o, Llama 405, G- Gemini 1.5 Flash, and Mistral Nemo.

Uh, these things kinda break the frontier, and then if you update it like a month later, uh, I think if I go back a month here, um, y- you update it, you can see start, you can see more items start to appear, uh, here as well with the August frontier, with Gemini 1.5 Flash coming out, uh, with an August update as, as compared to the June update, um, being a lot cheaper, uh, and roughly the same ELO.

And then, uh, we update for September, um, and that this is one of those things where, um, it really started to s- to, uh, we really started to understand the pricing curves being real instead of something that some random person on the internet drew, drew on a chart, because Gemini 1.5 cut their prices, and cut their prices exactly in line with where everyone else is in terms of their ELO price charts.

So, um, if you plot by, by, by September, we had the o1-preview and pricing and, and costs and ELOs. Um, so the, the frontier was o1-preview, uh, GPT-4o, o1-mini, uh, 4o Mini, and then Gemini Flash at the low end.

Uh, that was the frontier as of September. Gemini 1.5 Pro was not on that frontier. But then they cut their prices, uh, they halved their prices, and suddenly they were on the frontier. Um, and so like it's a very, very tight and predictive line, which I thought it was, uh, really in- interesting and entertaining as well.

Um, and, uh, I thought that that was kinda cool. Um, in November, we had 3.5 Haiku New Um, and obviously we had Sonnet as well. Uh, Sonnet is, uh, is not-- I don't know where there's Sonnet on this chart, but, um, Haiku New, uh, basically, uh, was 4X the price of old Haiku.

Or the... Sorry, 3.5 Haiku was 4X the price of 3 Haiku, and people were kind of unhappy about that. Um, there's a reasonable assumption, to be honest, uh, that it's not a price hike, it's a just a bigger model, so it costs more.

Uh, but we just don't know that. There was no transparency on that, so we are m- left to draw our own conclusions on what that means. Um, that's just is what it is. Um, so yeah, that, that would be the sort of price ELO chart.

Uh, I would say the, the main update for this one, if you go to my LLM pricing chart, which is public, um, you can ask me for it or I've shared it online as well. The most recent one is Amazon Nova, which we briefly, briefly talked about on the pod, where, um, that they've really sort of come in and, you know, basically offered Amazon Basics LLM, uh, where Amazon Pro, Nova Pro, Nova Lite, and Nova Micro are the efficient frontier for, uh, their intelligence levels of 1200 to 1300.

Um, you wanna get beyond 1300, you have to pay up for the o1s of the world and the 4os of the world and the Gemini 1.5 Pros of the world. Um, but, uh, 2 Flash is not on here, and is probably a good deal high-higher.

Flash Thinking is not on here, as well as all the other QWQs, R1s and, and all the other sort of thinking models. So I'm gonna have to update this chart. It's, uh, it's always a struggle to keep up to date.

But I wanna give you the idea that basically for, uh, through the month-- through the, through the, through, through 2024, for the same amount of ELO, what you used to p-pay at the start of 2024, um, you know, let's say, you know, 50...

four, $40 to $50 per million tokens, uh, now is available, uh, approximately at, with Amazon Nova, uh, approximately at, I don't know, $0.075, uh, dollars per token. So like seven, 7, 7.5 cents. Um, so that is a couple orders of magnitude at least.

Uh, actually almost three orders of magnitude improvement in a year. And I used to say that intelligence, the cost intelligence was coming down, um, one order of magnitude per year, like 10X. Um, you know, uh, that is already faster than Moore's Law, but coming down three times this year, um, is something that I think not enough people are talking about.

And so even though people understand that intelligence has become cheaper, I don't think people are appreciating how much more accelerated this year has been. And obviously, I think a lot of people are speculating how much more next year will be with H200s becoming commodity, Blackwells coming out.

We-- It's very hard to predict, and obviously there are a lot of factors beyond just the GPUs. So that is the sort of thematic overview, and then we went into the sort of the a- the, the annual overview.

This is basically, um, us going through the AI News, uh, releases of the, of, uh, of the year and just picking out favorites. Um, I had Will, our new research assistant, uh, help out with the research, but you can go onto AI News and check out, um, all the, all the sort of top news of, of the day.

Uh, but we had a little bit of an AI rewind thing, which, uh, I'll briefly bridge you in back to the recording that we had. So January, we had the first round of the year for Perplexity. Um, and for me, it was notable that Jeff Bezos backed it.

### Monthly Rewind

**Swyx** [1:22:57]
Um, Jeff doesn't invest in a whole lot of companies, but when he does, um, you know, he backed Google back in the day, and now he's backing the new Google, which is kinda cool. Perplexity is now worth $9 billion.

They-- I think they have four rounds this year. Um a, uh, Will also picked out s- uh, that Sam was talking about GPT-5 soon. Uh, this was back when, uh, he was, I think, at one of the sort of global summit type things, Davos.

And, um, yeah, no GPT-5. It's actually we got o1 and o3. In February, uh, you know, people were, uh... we were still sort of thinking about last year's DevDay, and this is three months on from DevDay. Uh, people were kind of losing confidence in GPTs, um, and I, I've, I feel like that hasn't super recovered yet.

I hear from people that, uh, there are still stuff in the works and you should not give up on them and, and they're actually underrated now, um, which is good. So I think people are taking a stab at the problem.

I think it's a thing that should exist, and we just need to keep iterating on them. Honestly, um, any marketplace is hard. It's very hard to judge given all the other stuff they've shipped. Um, ChatGPT also released Memory in February, which we talked about a little bit.

We also had Gemini's diversity drama, which, uh, we don't tend to talk a ton about in this podcast because we try to keep it technical. But we also had started seeing, uh, context window size blow out. Uh, so we...

This year, I mean, it was, it was Gemini with 1 million tokens. Um, but also, I think there's 2 million tokens talked about. We had a podcast with Gradients talking about how to fine-tune for 1 million tokens. It's not just like what you declare to be your token w- to- token context, but you also have to use it well.

And increasingly, I think people are looking at not just Ruler, which is sort of multi-needle in a haystack we talked about, but also Muser and like reasoning over long context, not just being able to retrieve over con- long context.

Um, so that's what I would call out there. Uh, specifically, I think Magic.dev as well made a lot of waves for their 100 million token model, which was kinda teased last year, but whatever it was, they made s- they made some noise about it.

Um, still not released, so we don't know, but we'll try to get them on, on the podcast. In March, Claude 3 came out, which huge, huge, huge for Anthropic. This basically mar- started to mark the shift of market share that we talked about on the, uh, earlier in the pod, where, uh, most production traffic was on OpenAI, and now Anthropic, um, had a decent frontier model family that people could shift to.

And obviously, now we know that Sonnet is, is kind of the workhorse, um, just like 4o is the workhorse of, of OpenAI Devin, um, came out in March, and that was a very, very big launch. It was probably one of the most well-executed PR campaigns, um, maybe in tech, maybe this decade.

Um, and, and then I think, you know, there was a lot of backlash as to, like, what specifically was real in the v- in the videos that they launched with, and obv- and then they took nine months to, to ship to GA.

And, uh, now you can buy it for $500 a month and form your own opinion. I think some people are happy, some people are less so, but, um, it's very hard to live up to the promises that they made.

And, um, the fact that they-- some of them-- for some of them they do, which is interesting. Um, I think, uh, the, the main thing I would caution now for Devin, and I think people call me a Devin shill sometimes 'cause I say nice things.

Like, one nice thing doesn't mean I'm not- I'm a shill. Um, basically it in- is that, like, a lot of the ideas can be copied, and this is the- always the threat of quote-unquote GPT wrappers, that you achieve product market fit with one feature, it's gonna be copied by 100 other people.

So of course you gotta compete with branding and better product and better engineering and all that sort of stuff, which Devin has in spades. So we'll see. April, uh, we actually talked to Udio and Suno. Um, we talked to Suno specifically, but Udio I, I also got a beta access to, and like, um, AI music generation.

We, we played with that on the podcast. I loved it. Some of our friends at the pod like play it in their cars. Like I've rode in their cars while they played our Suno intro songs, and I freaking loved using o1 to craft the lyrics and Suno to, and Udio to, to, to make the, the, the songs.

Uh, but ultimately, like a lot of people, you know, some people were skipping them. I don't know what exact percentages, but those, you know, 10% of you that skipped it, you're, you're the reason why we cut the intro songs.

Um, we also had Llama 3 release, so you know, I think people always want to see, uh, you know, like a, a good frontier, uh, open source model, and Llama 3 obviously delivered on that with the 8B and 70B.

The 400B came later. Then, um, May, GPT-4o released. Um, we, uh... And, and it was like kind of a model efficiency thing, but also I think just a really good demo of all the, uh, the things that 4o is capable of.

Like this is where the messaging of omnimodel really started kicking in. You know, pre- previously 4 and 4 Turbo were all text, um, and not natively, uh, sort of vision. I mean, they had vision, but not, not natively voice and, you know, that, uh, I think everyone was- fell in love immediately with the Sky voice, and Sky voice got taken away, um, before the public release.

And, um, I think it's probably self-inflicted. Um, I, I think the, the, the version of events that has Sam Altman basically putting a foot in his mouth with a three-letter tweet, um, causing decent grounds for a lawsuit in- when- where there was no grounds to be had because they actually just used a voice actress that sounded like Scarlett Johansson, um, uh, is unfortunate because we could've had it and we, we don't.

So that's what it is, and that's what the consensus seems to be from the people I talk to. Uh, people be pining for the Scarlett Johansson voice. In June, Apple announced Apple Intelligence at WWDC. Um, and um, we have it, and most of us, if you update your phones, uh, have it now if you're, if you're on an iPhone.

And I would say it's like decent. You know, like I think it wasn't the game changer, the thing that caused the Apple stock to rise like 20%. And just 'cause everyone was like gonna update- upgrade their iPhones just to get Apple Intelligence.

Um, it did not become that. But, um, it, it is the, uh, probably the largest scale rollout of transformers yet, um, after Google rolled out BERT for search. And, um, and people are using it, and it's a 3B, you know, foundation model that's running locally on your phone with LoRAs that are hot swaps, and we have papers for it.

Honestly, Apple did a fantastic job of doing the best that they can. They, they're not the most transparent company in the world and, and nobody expects them to be, but, um, they gave us more than I think we normally get for Apple tech, and, uh, that's very nice for the research community as well.

Nvidia, um, uh, I think we continue to talk about... I think, um, I was at the Taiwanese trade show Com- Computex, um, and saw, saw him signing, you know, women body parts, and I think that was maybe a sign of the times, maybe a sign that things have peaked.

But, uh, things have clearly not peaked 'cause they continued going. Um, Ilya... And, and then, and then that bridges us back into the episode recording. Uh, I'm gonna stop now and stop yapping. But, uh, yeah, we've you know, we recorded a whole bunch of stuff.

We lost it, and we're scrambling to re-re-record it, uh, for you. But also t- we're trying to close the chapter on 2024. So, uh, now I'm gonna cut back to the recording where we talk about the rest of June, July, August, September, and the second half of 2024's news, and we'll end the episode there.

Ilya came out from the woodwork.

**Alessio** [1:30:41]
Yeah. Raised, raised a billion.

**Swyx** [1:30:43]
Uh, saw, saw a term sheet. Raised a billion dollars. Dan Gross seems to have now become full-time CEO of, of the company, which is interesting. Uh, I, I thought he was gonna be an investor for life.

**Alessio** [1:30:52]
Yeah.

**Swyx** [1:30:52]
Um, but now he's operating. Uh-

**Alessio** [1:30:54]
He was ambassador for a short amount.

**Swyx** [1:30:55]
For, for-

**Alessio** [1:30:56]
Very short amount

**Swyx** [1:30:57]
... two years. What, what else can we say about Ilya? I mean, like, I think this, this idea that you only ship one product and it's a straight shot to super intelligence seems like a really good focusing mission.

**Alessio** [1:31:07]
Yeah.

**Swyx** [1:31:07]
But then, like it runs counter to basically both Tesla and OpenAI in terms of the ship intermediate products that get you to that vision.

**Alessio** [1:31:16]
Well, I think the question is, like, OpenAI now needs then more money because they need to support those products, and I think maybe their bet is like with 1 billion we can get to the thing. Like we don't wanna have to have intermediate steps.

**Swyx** [1:31:28]
Right.

**Alessio** [1:31:29]
Like we're just making it clear that like this is-

**Swyx** [1:31:31]
Yeah

**Alessio** [1:31:31]
... what it's about.

**Swyx** [1:31:32]
But then like where do you get your data? Like, you know, where you

**Alessio** [1:31:35]
Yeah, totally. Well, that's the-

**Swyx** [1:31:37]
Um, so, so I think that's the, that's the question. I think, uh, we can also use this as, as part of a general theme of the safety wing of OpenAI leaving.

**Alessio** [1:31:45]
Yeah.

**Swyx** [1:31:46]
Uh, it's fair to say that, you know, Jan Leike also left, and, like, basically the entire sa- safe, uh-

**Alessio** [1:31:51]
Yep

**Swyx** [1:31:51]
... uh, super alignment team left.

**Alessio** [1:31:53]
Yeah, then there was Artifacts, kinda like the ChatGPT Canvas, uh, equivalent-

**Swyx** [1:31:57]
Yeah

**Alessio** [1:31:57]
... that came out.

**Swyx** [1:31:57]
Code interpreter. I think, I think more, more code-oriented.

**Alessio** [1:31:59]
Yeah.

**Swyx** [1:32:00]
No one has a Canvas clone yet apart from-

**Alessio** [1:32:02]
Yeah. Yeah, yeah, yeah

**Swyx** [1:32:03]
... OpenAI. Uh, interestingly, I think the same person responsible for Artifacts and Canvas, Karina-

**Alessio** [1:32:09]
Mm-hmm

**Swyx** [1:32:09]
... um, so she left Anthropic after this to join OpenAI-

**Alessio** [1:32:12]
Yeah

**Swyx** [1:32:12]
... one of the rare, uh, reverse moves.

**Alessio** [1:32:14]
Yeah, and then we had AI engineer World's Fair in June. That was over 2,000 people-

**Swyx** [1:32:19]
Sounds like a good conference

**Alessio** [1:32:20]
... not including us. I would love to attend the next one.

**Swyx** [1:32:25]
If only we can get tickets. Um-

**Alessio** [1:32:27]
Oh, man

**Swyx** [1:32:27]
... yeah, uh, but I, I think a really good demo. We now have it and deployed for everybody. Uh, and Gemini actually kind of beat them to the GA release, which is, uh, kind of interesting.

**Alessio** [1:32:36]
Yeah.

**Swyx** [1:32:36]
I think that everyone should basically always have this on, as long as you're comfortable with the privacy settings because then you have a second person kind of looking over your shoulder.

**Alessio** [1:32:43]
Yeah.

**Swyx** [1:32:44]
And, like, this time next year I would be willing to bet that I would just have this running on my machine. And, uh, you know, I, I, I think that assistance always on that you can talk to with Vision that sees what you're seeing, I think that is where that leads one of our software experience to go.

**Alessio** [1:32:58]
Yeah.

**Swyx** [1:32:58]
Then it will be another few years for that to happen in real life in-

**Alessio** [1:33:02]
Mm-hmm

**Swyx** [1:33:02]
... outside of the screen.

**Alessio** [1:33:03]
Yep.

**Swyx** [1:33:03]
But for screen experiences, I think it's basically here but not even- evenly distributed. And, and, you know, we've just seen the GA of this capability-

**Alessio** [1:33:11]
Mm-hmm

**Swyx** [1:33:11]
... that was demoed in June.

**Alessio** [1:33:13]
And then July was Llama 3.1 which, you know, we've done a whole podcast on.

**Swyx** [1:33:17]
Yep.

**Alessio** [1:33:17]
Um, but that was, that was great.

**Swyx** [1:33:19]
July, August, kinda quiet. Yeah.

**Alessio** [1:33:20]
Yeah.

**Swyx** [1:33:20]
August was structure outputs.

**Alessio** [1:33:21]
Yeah, structure outputs. We also did a full podcast on that. And then September we got o1.

**Swyx** [1:33:26]
Yes.

**Alessio** [1:33:26]
Strawberry, AKA QStar-

**Swyx** [1:33:29]
Yeah

**Alessio** [1:33:29]
... AKA we had a nice party with strawberry glasses.

**Swyx** [1:33:31]
Yes. I think very underrated. Like, this is basically from the first internal demo of Q- of Strawberry was, let's say November 2023.

**Alessio** [1:33:41]
Mm-hmm.

**Swyx** [1:33:42]
So between November to September, like the, the whole red teaming and everything, honestly a very good ship rate. Like I, I don't know if, like, people are giving OpenAI enough credit-

**Alessio** [1:33:52]
Yeah

**Swyx** [1:33:52]
... for like this all being available in ChatGPT and then shortly after in, in API. I think maybe on the same day. I, I, I don't know. I don't remember the exact sequence already. But like this is like a frontier model that was like rolled out very, very quickly to the, to the whole world.

**Alessio** [1:34:06]
Mm-hmm.

**Swyx** [1:34:06]
And then we immediately got used to it, immediately said it was shit because-

**Alessio** [1:34:09]
Right. Yeah

**Swyx** [1:34:09]
... I'm still using Sonnet, whatever. But like still very good. And then obviously now we have, uh, o1 Pro and o1 Full. I think like in terms of like biggest ships of the year-

**Alessio** [1:34:18]
Mm-hmm

**Swyx** [1:34:18]
... I think this is it, right?

**Alessio** [1:34:19]
Yeah. Yeah, totally. Yeah. And I think it now opens a whole new Pandora bo- Pandora's box for like the inference time compute and then-

**Swyx** [1:34:25]
Yeah. Yeah, yeah

**Alessio** [1:34:25]
... and all that, so.

**Swyx** [1:34:26]
Yeah. It's funny because, like, it could have been done by anyone else be- before.

**Alessio** [1:34:30]
Yeah.

**Swyx** [1:34:30]
Literally this is an open secret they were working on it ever since they hired Noam. Um, but they-- no one else did.

**Alessio** [1:34:36]
Yeah.

**Swyx** [1:34:36]
Another discovery, I think, um, Ilya actually worked on a previous version called GPT Zero in 2021. Same exact idea and it failed.

**Alessio** [1:34:46]
Yeah.

**Swyx** [1:34:46]
Whatever that means. Yeah.

**Alessio** [1:34:47]
Timing.

**Swyx** [1:34:48]
Voice mode also released.

**Alessio** [1:34:49]
Voice mode, yeah. Um, yeah, I think most people have tried it by now- ... because it's generally available.

**Swyx** [1:34:54]
Yeah. Um, I think your, your wife also likes it.

**Alessio** [1:34:56]
Yeah. Yeah, yeah. She talks to it all the time. Um-

**Swyx** [1:34:59]
Okay

**Alessio** [1:35:00]
... Canvas in October.

**Swyx** [1:35:02]
Okay.

**Alessio** [1:35:02]
Another big release. Compute use.

**Swyx** [1:35:03]
Have you used it much?

**Alessio** [1:35:05]
Not really, honestly.

**Swyx** [1:35:07]
Yeah. I, I use it a lot.

**Alessio** [1:35:07]
What do you use it for mostly?

**Swyx** [1:35:09]
Drafting anything. I think that people don't see where all this is heading. Like OpenAI is really competing with Google in everything. Canvas is Google Docs, and like it's, it's a full document editing environment with a, with a auto assistor thing at the side that is arguably better than Google Docs, um, at least for some editing use cases, right?

'Cause it, it has a much better AI-

**Alessio** [1:35:28]
Yep

**Swyx** [1:35:28]
... integration than Google Docs with Gemini on the side. And so OpenAI is taking on Google and Google Docs. It's also taking on- taking it on in search. They, you know, they launched their, their little, uh, Chrome extension thing to, to be the default search.

And I, I think like piece by piece it's, it's kinda really tackling on Google in a very smart way that-

**Alessio** [1:35:47]
Yeah

**Swyx** [1:35:47]
... I think is additive to workflow, and people should start using it as intended because this is a peek into the future. Maybe they're not successful, but at least they're trying.

**Alessio** [1:35:55]
Yeah.

**Swyx** [1:35:55]
And I think Google has gone without competition for so long that anyone trying will be, will be-- will at least receive some attention from me.

**Alessio** [1:36:02]
Yeah. And then, yeah, computer use also came out, um-

**Swyx** [1:36:06]
Yeah

**Alessio** [1:36:06]
... that month. That was-- Yeah, that was a busy... It's been a busy-

**Swyx** [1:36:08]
Yeah

**Alessio** [1:36:09]
... couple months.

**Swyx** [1:36:10]
Busy couple months. Uh, I would say that computer use was one of the most upvoted demos, uh, on Hacker News of the year. But then comparatively I don't see people using it as much.

**Alessio** [1:36:20]
Yeah. Yeah, yeah.

**Swyx** [1:36:21]
This is how you feel the difference between a mature capability and an emerging capability. Maybe this is why Vision is emerging.

**Alessio** [1:36:27]
Mm-hmm.

**Swyx** [1:36:28]
Because I launched computer use, you're not using it today, but you use everything else in the mature category. And i- it's mostly because it's not precise enough, or it's too slow, or it's too expensive, and those will be the, the main-

**Alessio** [1:36:38]
Yeah

**Swyx** [1:36:38]
... criticisms.

**Alessio** [1:36:39]
Yeah. That makes sense. It's also just like overall uneasiness about just letting it go crazy-

**Swyx** [1:36:45]
I don't care

**Alessio** [1:36:45]
... on your computer. Yeah. No, no, totally. But I think a lot of people do.

**Swyx** [1:36:49]
November.

**Alessio** [1:36:50]
The R1, so that was kinda like the open source o1 competitor by DeepSeek.

**Swyx** [1:36:54]
This was a surprise. Yeah, nobody knew it was coming.

**Alessio** [1:36:55]
Yep.

**Swyx** [1:36:55]
Um, everyone knew like F1 we had a preview at the Fireworks HQ, and then I think some, some other labs did it. But I think R1 and QW- uh, QWQ, Quill-

**Alessio** [1:37:06]
Yeah

**Swyx** [1:37:06]
... um, from the Qwen team, both Alibaba affiliated-

**Alessio** [1:37:09]
Mm-hmm

**Swyx** [1:37:09]
... I think are the leading contenders on, on that front.

**Alessio** [1:37:11]
Yeah.

**Swyx** [1:37:12]
And, um, we'll, we'll see. We'll see.

**Alessio** [1:37:14]
What else to highlight? I, I think the Stripe Agent toolkit-

**Swyx** [1:37:17]
You like that one?

**Alessio** [1:37:18]
I mean, it's a small-- It's a small thing, but it just like people are like agents are not real. It's like when you have, you know, companies like Stripe and like start to build things to support it-

**Swyx** [1:37:25]
Yeah

**Alessio** [1:37:25]
... it might not be real today, but obviously they don't have to do it because they don't-- they're not an AI company. But the fact that they do it ... shows that there's one demand-

**Swyx** [1:37:34]
Yeah

**Alessio** [1:37:34]
... and so there's belief from-

**Swyx** [1:37:35]
Yeah

**Alessio** [1:37:35]
... on their end.

**Swyx** [1:37:36]
This is a broader thing about r- a broader thesis for me that I'm exploring around do we need special SDKs for agents?

**Alessio** [1:37:43]
Mm-hmm.

**Swyx** [1:37:43]
Why can't normal SDKs for humans do the same thing? Stripe Agent Toolkits happens to be a wrapper on the Stripe SDK. It's fine. It's just like a nice little DX layer. But like, it's still unclear to me. Uh, I think, um, I have been asked my opinion on this before, and I said, I think I said it on a podcast, which is like the, the main layer that you need is the separate auth roles so that you don't assume it's a human-

**Alessio** [1:38:04]
Yeah, yeah

**Swyx** [1:38:04]
... um, uh, doing the- doing these things, and you can lock things down much quicker, or you can identify whether it is, it is an agent acting on your behalf or actually you.

**Alessio** [1:38:13]
You. Mm-hmm.

**Swyx** [1:38:13]
Um, and that, that is something that you need. Um, I had my ElevenLabs key pwned because I lost my laptop, and, uh, I saw a whole bunch of API calls, and I was like, "Oh, is that me or is that, is that someone?"

And, and it turned out to be a key that had, that, that had committed, uh, on-

**Alessio** [1:38:28]
Mm-hmm

**Swyx** [1:38:28]
... onto GitHub and, and that I didn't scrape. And so sourcing of where API usage is coming from, I think, um, you know, you should attribute it to agents and build for that world. But other than that, I think SDKs, I would see it as a failure of dev tech-

**Alessio** [1:38:43]
Yeah

**Swyx** [1:38:43]
... and, and AI that we need e- every single thing needs to be reinvented for agents.

**Alessio** [1:38:48]
I agree in some ways. I think in other ways we've also, like, not always made things super explicit. There's kinda like a lot of defaults that people do when they design APIs. But like, um, I think if you were to redesign them in a world in which the person or the agent using them has like almost infinite memory and context-

**Swyx** [1:39:06]
Yeah

**Alessio** [1:39:06]
... like you would maybe do things differently.

**Swyx** [1:39:08]
Yeah.

**Alessio** [1:39:08]
But I don't know. I think to me the most interesting is like REST and GraphQL is almost more interesting in the world of agents because agents could come up with so many different things to query. Versus like before I always thought GraphQL was kinda like not really necessary because like you know what you need, just build the REST endpoint for it.

So yeah, I'm curious to see what else changes. And then yeah, the search wars. I think that was, you know, SearchGPT, Perplexity-

**Swyx** [1:39:33]
Dropbox

**Alessio** [1:39:33]
... you know, Dropbox Dash. Yeah. We had Drew on the pod, and then we had it at the Pioneer Summit. The fact that Dropbox has a Google Drive integration is just like if you told somebody five years ago, it's like-

**Swyx** [1:39:44]
Ooh

**Alessio** [1:39:45]
... Dropbox doesn't really care about-

**Swyx** [1:39:46]
Yeah. We've talked about this

**Alessio** [1:39:46]
... hosting your files, you know? It's like that doesn't compute. So yeah, I'm curious to see where, where that goes-

**Swyx** [1:39:52]
Cool

**Alessio** [1:39:53]
... this whole space.

### Wrap-Up

**Swyx** [1:39:54]
And that brings us up to December. Uh, still, still developing. I'm curious what the last day of, um, OpenAI's shipments will, will bring.

**Alessio** [1:40:00]
Yeah.

**Swyx** [1:40:00]
I think everyone's expecting something big there. I think so far has been a very eventful year. Definitely has grown a lot. Uh, we were asked by Will actually like whether we made predictions. I don't think we, we did, but maybe-

**Alessio** [1:40:10]
Not really

**Swyx** [1:40:11]
... maybe we should.

**Alessio** [1:40:11]
I, I think, well, I think we definitely talked about agents.

**Swyx** [1:40:14]
Yes.

**Alessio** [1:40:15]
And I don't know if we said it was the year of the agents, but we said-

**Swyx** [1:40:19]
But next year is the year.

**Alessio** [1:40:20]
No, no, but- ... well, you know, the anatomy of autonomy, that was April 2023.

**Swyx** [1:40:24]
Mm-hmm.

**Alessio** [1:40:25]
You know? So obviously there's been belief for a while.

**Swyx** [1:40:28]
Yes.

**Alessio** [1:40:28]
Um, but I think now the models are c- I would say maybe the last, yeah, two months-

**Swyx** [1:40:33]
Yeah

**Alessio** [1:40:33]
... have made a big push in like capability with like 3.6, o1.

**Swyx** [1:40:36]
Yeah. I mean, Ilya saying the word agentic on stage at NeurIPS-

**Alessio** [1:40:39]
Yeah

**Swyx** [1:40:39]
... it's, it's a, it's a big deal. Uh, Satya I think also saying that a lot-

**Alessio** [1:40:43]
Yeah

**Swyx** [1:40:43]
... these days. I mean, Sam has been saying that for a while now. So DeepMind w- when they announced Gemini 2.0, they announced Deep Research, but also Project Mariner, which is a browser agent-

**Alessio** [1:40:53]
Mm

**Swyx** [1:40:53]
... which is their computer use type thing, as well as Jules, which is their code agent, and I think that basically complements with whatever OpenAI is shipping next year, which is codename Operator, uh, which is their agent thing.

It makes sense that you pr- i- if it actually replaces a junior employee, they will charge $2,000 for it.

**Alessio** [1:41:09]
Yeah. I think that's, uh, my whole... I, I did this post that it's pinned on my Twitter, so you can find it easily, but about skill floor and skill ceiling in jobs, and I think the skill floor more and more.

I think 2025 will be the first year where the AI sets the skill floor of a role, you know? I don't think that has been true in the past, but yeah, I think now we're like, you know, if Devin works, if all, all these customer support agents are working, so now to be a customer support person you need to be better than an agent because the economics just don't work.

I think the same is gonna happen, too, in software engineering, which I think the skill floor is very low, you know? Like there's a lot of people doing software engineering that are really not that good. So I, I'm curious to see at the next yearly recap what other jobs are gonna have that change.

**Swyx** [1:41:53]
Yeah. Every NeurIPS that I go I have some chats with researchers, and, uh, the, the-- I'll, I'll, I'll just highlight the best prediction-

**Alessio** [1:41:59]
Yeah

**Swyx** [1:41:59]
... from that group, and then we'll move on to end of year recap in terms of we'll just go down the list of top five podcasts and then, and then we'll end it.

**Alessio** [1:42:05]
Yep.

**Swyx** [1:42:05]
So the, the best rec- uh, best prediction was that there will be a foreign spy caught in, uh, at one of the major labs. So this is part of the consciousness-

**Alessio** [1:42:17]
Interesting

**Swyx** [1:42:17]
... already that, uh-

**Alessio** [1:42:17]
Yeah

**Swyx** [1:42:17]
... you know, like, you know, whenever you see someone who is like too attractive in a- ... in a San Francisco party where like the ratio is like 100 guys to one girl, and like suddenly the girl is like super interested in you, like, you know, it may not be your looks.

Um, so there's a lot of like state-level secrets that are kept in these labs and not that much security. I think, uh, if anything, the situ- situational awareness essay did to, to raise awareness of it. Like I think it, it was directionally correct even if not precisely correct.

**Alessio** [1:42:44]
Yep.

**Swyx** [1:42:44]
We should start caring a lot about this. Um, OpenAI has hired a CISO this year, and I think like the, the security space in general-- Oh.

**Alessio** [1:42:51]
Yeah, yeah.

**Swyx** [1:42:51]
I remember what I was gonna say about Apple foundation model before we, we, we cut for break. They announced Apple Secure Cloud, Cloud Compute.

**Alessio** [1:42:56]
Yeah.

**Swyx** [1:42:57]
And I think, um, we are also interested in investing in areas-

**Alessio** [1:43:00]
Mm-hmm

**Swyx** [1:43:00]
... that are basically secure cloud LLM inference for everybody. I think like what we have today is not secure enough-

**Alessio** [1:43:06]
Yeah

**Swyx** [1:43:06]
... because it's, it's like normal security when like this is l- literally of state-level interest.

**Alessio** [1:43:11]
Mm-hmm. Agreed.

**Swyx** [1:43:12]
Top episodes?

**Alessio** [1:43:12]
Yeah. So I'm just going through the Substack. Number one, the David Wen one.

**Swyx** [1:43:18]
Yeah.

**Alessio** [1:43:18]
It's the most popular of 2024.

**Swyx** [1:43:20]
Uh-

**Alessio** [1:43:20]
"Why Google Failed to Make GPT-3".

**Swyx** [1:43:22]
I will take a little bit of credit for that, for the naming of that one because I think that was the Ha- that was the Hacker News thing. It's very funny because like actually, obviously he wants to talk about Adept-

**Alessio** [1:43:29]
Yeah.

**Swyx** [1:43:29]
... but then he spent half the episode talking about his time at, at OpenAI. But I think it was a very useful insight that I'm still using today. Even in like the Ilya post I was still referring to what he said.

And then when I, when we do podcast episodes, I try to look for that.

**Alessio** [1:43:42]
Mm-hmm.

**Swyx** [1:43:42]
I, I try to look for things that we'll still be referencing in the future.

**Alessio** [1:43:46]
Yeah.

**Swyx** [1:43:46]
And that concentrated bet-ness, David talked about the brain-compute marketplace, and then Il- Ilya in his emails that we, that I covered in the "What Ilya Saw"-

**Alessio** [1:43:56]
Mm-hmm

**Swyx** [1:43:56]
... essay had the OpenAI side of this where they were like Um, one big training run is much, much more valuable than the hundred equivalent small training runs.

**Alessio** [1:44:05]
Yep.

**Swyx** [1:44:05]
So we need to go big, and we need to concentrate that-

**Alessio** [1:44:08]
Mm-hmm

**Swyx** [1:44:08]
... to not spread them.

**Alessio** [1:44:09]
Number two, how NotebookLM was made.

**Swyx** [1:44:12]
Yeah. Um-

**Alessio** [1:44:13]
That was-

**Swyx** [1:44:13]
That was-

**Alessio** [1:44:13]
Yeah

**Swyx** [1:44:14]
... that was fun.

**Alessio** [1:44:14]
Yeah. And everybody-- I mean, I think that's, like, a great example of, like, just timeliness. You know? I think it was top of mind for everybody. They were great guests. Um, it just made the rounds on social media.

**Swyx** [1:44:25]
Yeah. Um, and that one, I would say, Risa is obviously a star. She's-- but she's been on every episode, every podcast. But Usama, I think, you know, actually, actually being the guy who worked on the audio model, being able to talk to him I think was, was a, was a great gift for us, and I think f- uh, people should listen back to how they trained-

**Alessio** [1:44:40]
Yep

**Swyx** [1:44:41]
... the NotebookLM model because I think you put that level of attention on any model, you will make it Sora.

**Alessio** [1:44:47]
Yeah. No, that's true.

**Swyx** [1:44:48]
And it's specifically like, uh, they didn't have evals. They just-

**Alessio** [1:44:52]
Vibes

**Swyx** [1:44:53]
... they had, yeah, a group session with vibes.

**Alessio** [1:44:55]
The ultimate guide to prompting.

**Swyx** [1:44:57]
Yeah.

**Alessio** [1:44:57]
That was number three. I think all these episodes that are, like, summarizing things that people care about but they're disparate I think always do very well.

**Swyx** [1:45:05]
This helps us save on a lot of smaller prompting episodes, right?

**Alessio** [1:45:09]
Yeah.

**Swyx** [1:45:09]
If we interviewed individual paper authors with, like, a 10-page paper that is just a different prompt, like, not as useful as, like, an overview survey thing.

**Alessio** [1:45:17]
Yeah.

**Swyx** [1:45:18]
I think the question is what to do from here. People have-- Actually, I, I would say I've been surprised by how well-received that was. Should we do ultimate guide to other things, and then should we do Prompting 201, right?

**Alessio** [1:45:30]
Yeah.

**Swyx** [1:45:30]
Those, those are, those are the two lessons that we can learn from the success of this one.

**Alessio** [1:45:32]
I think if somebody does the work for us, that was the good thing about Sander. Like, he had done all the work for us.

**Swyx** [1:45:37]
Yeah, yeah. Sander's very, very-

**Alessio** [1:45:38]
So-

**Swyx** [1:45:38]
... um, fastidious about this.

**Alessio** [1:45:40]
Yeah.

**Swyx** [1:45:40]
So he, he did a lot of work on that. Uh, a-and, you know, definitely keen to have him-

**Alessio** [1:45:43]
Yeah

**Swyx** [1:45:43]
... on next year to, to talk more prompting. Okay.

**Alessio** [1:45:46]
So-

**Swyx** [1:45:46]
Then, then the next one is the not safe for work one.

**Alessio** [1:45:48]
No.

**Swyx** [1:45:49]
Or structured outputs.

**Alessio** [1:45:51]
That actually was Pinterest.

**Swyx** [1:45:52]
Really?

**Alessio** [1:45:53]
Yeah.

**Swyx** [1:45:54]
Okay, we have a different list then, but yeah.

**Alessio** [1:45:56]
I'm just going on the Substack.

**Swyx** [1:45:58]
I see. I see. So that includes the number of likes, but, uh, I was, I was going by downloads.

**Alessio** [1:46:03]
Mm.

**Swyx** [1:46:04]
It's fine.

**Alessio** [1:46:04]
I see.

**Swyx** [1:46:05]
Number of interest.

**Alessio** [1:46:06]
I would say this is almost recency bias in the way that, like, the audience keeps growing, and then, like-

**Swyx** [1:46:11]
I see

**Alessio** [1:46:11]
... the most recent episodes get more views.

**Swyx** [1:46:13]
I see.

**Alessio** [1:46:14]
So I, I would say definitely, like, the NSFW one was very popular. What people were telling me they really liked-

**Swyx** [1:46:21]
Yeah

**Alessio** [1:46:21]
... because it was something people don't, don't cover.

**Swyx** [1:46:24]
Yeah.

**Alessio** [1:46:24]
Um, yeah, structured outputs, I think people liked that one. The-- I mean, the SIM one, yeah, I, I think that's, like, something I refer to all the time. I think that's one of the most interesting areas-

**Swyx** [1:46:33]
The structured output area?

**Alessio** [1:46:34]
... for the new year. No, the simulation, um-

**Swyx** [1:46:36]
Oh, WebSim, WebSim. Really?

**Alessio** [1:46:37]
Yeah. Not, not that use case, but, like, how do you use that for, like, model-

**Swyx** [1:46:42]
Yeah

**Alessio** [1:46:42]
... um, training-

**Swyx** [1:46:43]
Yeah, yeah, yeah

**Alessio** [1:46:43]
... and, like, agents learning and, and all of that.

**Swyx** [1:46:45]
Yeah. Um, so I would definitely point to our newest seven-hour long episode on simulative environments because it is the, let's say, the scaled-up-

**Alessio** [1:46:55]
Yeah

**Swyx** [1:46:55]
... uh, very serious AGI lab version of WebSim and WorldSim. If you take it very, very seriously, you get Genie 2, which is exactly what you need to then build Sora and-

**Alessio** [1:47:04]
Mm

**Swyx** [1:47:04]
... everything else. Um, so yeah, I think s- uh, simulative AI, still in summer. Still-

**Alessio** [1:47:09]
Yeah. Still in summer

**Swyx** [1:47:10]
... still, still coming. And I was actually reflecting on this. Like, would you, would you say that the AI winter has, like, coming on or, like, was never even here? 'Cause we did a winds of AI winter episode and I, you know, I was, like, trying to look for signs.

I think that's kind of gone now.

**Alessio** [1:47:24]
Yeah. I would say it was here in the vibes but not really in the reality. You know, when you look back at the yearly recap, it's like every month there was, like, progress. There wasn't really a winter.

**Swyx** [1:47:34]
Yeah, yeah.

**Alessio** [1:47:34]
There was maybe, like, a hype winter, but I don't know if the-

**Swyx** [1:47:37]
Yeah

**Alessio** [1:47:37]
... that counts as a real winter.

**Swyx** [1:47:39]
I, I think the scaling has hit a wall thing has been a big driving discussion for 2024.

**Alessio** [1:47:44]
Yeah.

**Swyx** [1:47:44]
And, you know, with some amount of conclusion on, in, in NEURIPS that we were also kind of pointing to in the winds of AI winter episode. But, like, it's not a winter by any means.

**Alessio** [1:47:54]
Yeah.

**Swyx** [1:47:55]
We know what winter feels like, and this is not winter. Uh, so I think things are, things are going well. I think every time that people think that there's, like, not much happening in AI, just think back to this time last year-

**Alessio** [1:48:06]
Right. Yeah

**Swyx** [1:48:06]
... and understand how much has changed from benchmarks to frontier models to market share between OpenAI and the rest. And then also cover, like, you know, the, the various coverage areas that we've marked out, how the discussion has, has evolved a lot and what we take for granted now versus what we did not have-

**Alessio** [1:48:21]
Yep

**Swyx** [1:48:21]
... a year ago.

**Alessio** [1:48:22]
Yeah. And then just to, like, throw that out there, there have been 133 funding rounds over 100 million in AI-

**Swyx** [1:48:29]
Mm-hmm

**Alessio** [1:48:29]
... this year.

**Swyx** [1:48:30]
Does that include Databricks, the, the largest venture round-

**Alessio** [1:48:32]
It does

**Swyx** [1:48:32]
... in history?

**Alessio** [1:48:32]
$10 billion. Sheesh. Well, that Mosaic now has been bought for two something billion-

**Swyx** [1:48:40]
Uh-huh

**Alessio** [1:48:40]
... because it was mostly stock, you know, so price goes up-

**Swyx** [1:48:43]
I see

**Alessio** [1:48:44]
... theoretically, the-

**Swyx** [1:48:45]
I see. So it was bought at a valuation of 40, right?

**Alessio** [1:48:47]
Yeah.

**Swyx** [1:48:48]
It was like 43 or something like that.

**Alessio** [1:48:49]
And, and at, at the time-- I remember at the time there was a question about whether or not that valuation was real, right?

**Swyx** [1:48:54]
Yeah. Well, that, that's why everybody-

**Alessio** [1:48:55]
'Cause Snowflake was down.

**Swyx** [1:48:56]
Yep.

**Alessio** [1:48:57]
Uh, and, like, Databricks was a private valuation that was, like, two years old.

**Swyx** [1:49:00]
Yep.

**Alessio** [1:49:00]
And it's like, who knows what, what this thing's worth. Now it's worth 60 billion.

**Swyx** [1:49:03]
It's worth more. It's worth more. That's what it's worth. It's worth more than what you thought. Um, yeah, it's been a crazy year, but I'm excited for next year. I feel like this is almost like, you know, now the agent thing needs to happen, and I think that's really the unlock.

**Alessio** [1:49:17]
Yeah. No, I think it needs to happen.

**Swyx** [1:49:18]
You know, I think that's, uh-- I mean, I need-

**Alessio** [1:49:19]
I, I have to agree with you. Agent-- next year's the year of the agent in production.

**Swyx** [1:49:22]
Yeah. I don't- I don't... You know, it's almost like I'm not 100% sure it will happen, but, like, it needs to happen. Otherwise, it's definitely the, the winter next, next year. Any other parting thoughts?

**Alessio** [1:49:34]
Uh, I'm very grateful for you. Uh-

**Swyx** [1:49:35]
Oh

**Alessio** [1:49:35]
... I think that, I think you've been, uh, the, the-- a dream partner to, to build Latent Space with. And, uh, and also the Discord community, the Paper Club people have been beyond my wildest dreams, like, uh, so supportive and, and successful.

Like, w- it's amazing that, you know, the, the community has, you know, grown so much, and, like, the, the vibe has not sh- changed.

**Swyx** [1:49:54]
Yep. Yeah, that is true.

**Alessio** [1:49:55]
Uh, you know, we started this-

**Swyx** [1:49:56]
We're almost at 5,000 people.

**Alessio** [1:49:57]
Yeah. We started this Discord, like, four years ago.

**Swyx** [1:49:58]
Yeah.

**Alessio** [1:49:59]
And, uh, still, like, people get it when they join. They're like, "You post news here, and then you discuss it in threads," and, you know, you, you try not to self-promote too much.

**Swyx** [1:50:06]
Mm-hmm.

**Alessio** [1:50:06]
And, and mostly people obey the rules-

**Swyx** [1:50:09]
Yeah, yeah

**Alessio** [1:50:09]
... and sometimes you smack them down a little bit, but that's okay.

**Swyx** [1:50:11]
We rarely have to ban people-

**Alessio** [1:50:13]
Yeah, yeah

**Swyx** [1:50:13]
... which is, uh, which is great. But yeah, man, it's been awesome, man. I think we both started not knowing where this was gonna go.

**Alessio** [1:50:20]
Yeah.

**Swyx** [1:50:20]
And now we've done 100 episodes. It's easy to see how we're gonna get to 200.

**Alessio** [1:50:23]
Mm-hmm.

**Swyx** [1:50:24]
I think maybe when we started it wasn't easy to see how we would get to 100.

**Alessio** [1:50:27]
Yeah.

**Swyx** [1:50:27]
You know? Yeah, excited for more. Subscribe on YouTube-

**Alessio** [1:50:31]
Yeah. We need YouTube

**Swyx** [1:50:31]
... because I swear to God we're doing so much work-

**Alessio** [1:50:33]
We need YouTube

**Swyx** [1:50:33]
... to make that work, so-

**Alessio** [1:50:34]
Yeah

**Swyx** [1:50:34]
... follow us there

**Alessio** [1:50:35]
It's very expensive for, for, uh, un-unclear payoff as to, like, what we're actually gonna get out of it, but, um, hopefully people discover us more there. Like-

**Swyx** [1:50:42]
Yeah

**Alessio** [1:50:42]
... I, I, I do believe in YouTube as a podcasting platform much more so than Spotify.

**Swyx** [1:50:46]
Yeah. Totally. Thank you all for listening.

**Alessio** [1:50:50]
Yeah. Thank you for listening.

**Swyx** [1:50:51]
See you in the new year.

**Alessio** [1:50:52]
Bye-bye.

---

This library is powered by PodHood (https://podhood.com), the podcast website platform.
